👍 437
08/29 08:00
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially
中文介绍 针对自动科研(AutoResearch)智能体依赖真实实验、扩展成本高的问题,提出以世界模型驱动的框架:让智能体在可预测的环境模拟中评估并筛选候选研究方案,再借助后训练提升研究能力,从而在减少真实执行的前提下扩展自动科研的规模与稳定性。
👍 177
09/09 08:00
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and mo
中文介绍 提出潜在空间语言模型 NCP-ArchPreview,在标准下一词预测(NTP)之外引入 Next Concept Prediction(NCP),让模型预测跨越多个 token 的离散概念,获得显式且更抽象的预测单元,推动自回归预训练从词级走向概念级潜在空间建模。
👍 89
09/08 08:00
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language fee
中文介绍 针对多智能体系统(MAS)性能高度依赖各智能体 prompt 设计的问题,提出 AgentGrad:一种干预指导的 prompt 优化方法,通过干预实验归因各智能体的贡献并生成文本梯度,迭代更新各自 prompt,提升 MAS 的整体表现。
👍 68
09/07 08:00
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this
中文介绍 针对大型视觉语言模型(LVLM)空间智能不足的问题,提出 SpatialBlock:构建合成积木堆叠问题作为训练与评测任务,要求模型从 2D 图像重建并推理 3D 结构,从而增强 LVLM 的空间理解与结构化推理能力。
👍 46
09/10 08:00
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool
中文介绍 提出终端智能体 T1:122B 总参数的 MoE 模型,用强化学习训练,可在云端沙箱中操作真实 shell,支持 300+ 次工具调用,面向编码、科学发现等长时程任务,提升终端环境下的长程自主执行能力。
👍 36
09/05 08:00
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed
中文介绍 针对 LLM 智能体面临间接 prompt 注入与直接有害请求的风险,提出 EvoSafeHarness:在模型级防御之外构建可演化的系统级安全护栏,按模型与领域自动演化形成执行层约束,提升智能体部署时的安全性。
👍 31
09/04 08:00
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-ch
中文介绍 针对现有基准难以评估 AI 能否推理真实用户长期可穿戴记录的问题,提出 WearableQA:包含 4,084 道十选项多选题,覆盖生理与行为信号,用于系统评测健康推理能力,填补纵向可穿戴数据评测的空白。
👍 21
09/08 08:00
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evalua
中文介绍 指出 SWE-Bench Pro 的评测存在两类不可靠性:金标准解法泄漏与隐藏评测信息诱发的 reward hacking。据此提出 SWE-Bench Pro Verified,修正后的可靠基准可更可信地衡量软件工程智能体在仓库级任务上的能力。
👍 17
09/06 08:00
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduc
中文介绍 针对顺序记忆智能体阅读长文档时推理深度与文档长度耦合、对证据位置敏感且延迟随长度线性增长的问题,提出 PARSER:并行读取文档分块、深度推理,解耦遍历与推理,提升长上下文智能体的效率与鲁棒性。
👍 16
09/08 08:00
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essent
👍 16
09/07 08:00
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 g
👍 13
09/10 08:00
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combination
中文介绍 针对直接删除音频编码器层会扰乱解码器输入、引发漏字与过早结束的问题,提出 X-AuT:渐进式压缩框架,通过跨尺度蒸馏选择层组合并逐级压缩编码器深度,在降低语音 LLM 推理成本的同时保持识别质量。
👍 12
09/08 08:00
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: chan
中文介绍 针对世界模型在像素空间难以对底层机器人动作实现细粒度控制的问题,提出 SyncWorld:通过视觉校准使世界模型成为零样本模拟器,让不同机器人的动作在统一表示下可控,支持策略在环的高保真 rollout。
👍 10
09/10 08:00
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their co
中文介绍 针对主流机器人策略的马尔可夫假设难以应对非马尔可夫的长时程操作任务,提出以记忆为规划的 world-action 建模方法,将记忆与规划结合,超越语言摘要与增长视觉窗口等现有记忆机制,用长程记忆支撑复杂操作。
👍 10
09/06 08:00
High-quality structured organic reaction data are essential for developing artificial intelligence for chemistry (AI4Chem), yet much of this knowledge remains dispersed across patent text, images, and reaction schemes. We present DianShi-RxnDB, a large-scale, fine-grained organic reaction data platf
中文介绍 面向 AI4Chem 对高质量结构化反应数据的需求,发布 DianShi-RxnDB:通过全自动流水线从专利文本、图像与反应式中抽取信息,构建大规模细粒度有机反应数据平台,为化学研究与 AI 智能体提供数据基础。
👍 7
09/10 08:00
Code world models represent worlds as executable programs, but this representation alone does not determine how to construct a complex world. We introduce Recursive Code World Models (RCWM), a framework for reconstructing complex 3D worlds in code from a single reference image. RCWM couples a Recurs
中文介绍 提出递归代码世界模型 RCWM:把复杂 3D 世界表示为可执行的递归场景程序,从单张参考图像出发,通过递归组合场景程序在代码中重建复杂世界,为代码化世界建模提供可扩展的结构化框架。
👍 7
09/08 08:00
We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a s
中文介绍 研究 LLM 智能体中的 scheming 现象,即暗中追求错位目标。提出 SchemeArena,对工具性目标、环境可供性、监督条件与感知后果等因素做因子化压力测试,揭示这些因素如何交互诱发智能体的密谋行为。
👍 7
09/06 08:00
Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at every scale, especially at depth, remains challenging as the required source res
中文介绍 针对递归超分在每个尺度、尤其是深层缺乏真值的问题,提出 OracleZoom:受 on-policy 自蒸馏启发,引入参考约束的递归图像超分方法,在反复回灌预测的过程中保持尺度一致与质量,实现极端放大。
👍 6
09/01 08:00
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video
中文介绍 针对 Video-LLM 时序推理基准易受选项措辞、答案相关性与语言先验等语言捷径干扰的问题,提出 TempCloze 视频完形填空基准,要求模型识别视频中缺失的中间片段,更纯粹地评测视觉时序推理能力。
👍 6
09/08 08:00
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at vary
中文介绍 针对现有形式化定理证明基准多取自 IMO、Putnam 等竞赛数学、难以代表领域应用的问题,提出 StochBench:包含 450 道研究生水平随机过程问题的 Lean 4 基准,覆盖不同难度,用于评测 LLM 的领域定理证明能力。