👍 444
08/29 08:00
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially
中文介绍 针对自动科研(AutoResearch)智能体难以在长程实验中有效学习的问题,提出用世界模型扩展其能力:让智能体在预测环境中预演实验过程并结合后训练,提升独立实现方案与从执行结果中学习的能力,推动经验研究自动化。
👍 231
09/09 08:00
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and mo
中文介绍 提出潜空间语言模型 NCP-ArchPreview,在标准下一词预测(NTP)之外引入 Next Concept Prediction(NCP),预测跨越多个 token 的离散概念,获得显式、可解释的中间表示,把自回归预训练从 token 层面推进到概念层面。
👍 92
09/08 08:00
Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language fee
中文介绍 LLM 多智能体系统(MAS)的性能高度依赖各智能体的提示设计。AgentGrad 针对文本梯度方法难以定位问题来源的缺陷,以干预方式定位关键智能体并引导提示更新,实现多智能体系统的自动化提示优化。
👍 91
09/07 08:00
Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this
中文介绍 大型视觉语言模型(LVLM)从 2D 图像重建与推理 3D 场景结构的空间智能仍然有限。论文提出合成积木堆叠任务 SpatialBlock,以可控生成的堆叠问题提供训练与评测信号,弥补现有方法在空间推理上的不足。
👍 53
09/10 08:00
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool
中文介绍 提出 T1,一个总参数 122B 的 MoE 终端智能体,通过强化学习训练,可在云端沙箱的真实 shell 中执行 300 余次工具调用,面向编码、科学发现等长程任务,验证了 RL 训练的大规模专家混合模型在终端长程任务上的可行性。
👍 47
09/05 08:00
Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed
中文介绍 针对 LLM 智能体面临的间接提示注入与直接有害请求,提出 EvoSafeHarness,自动演化出模型与领域特定的系统级安全防护层。相比通用 harness,它在模型级防御之外提供定制化的执行期约束,增强智能体实际行为的安全性。
👍 37
09/04 08:00
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-ch
中文介绍 提出 WearableQA 基准,包含 4084 道十选一选择题,用于评测 AI 系统对真实用户长期可穿戴记录的健康推理能力。现有基准很少检验模型能否对纵向生理与行为信号进行推理,该工作填补了这一评测空白。
👍 30
09/10 08:00
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combination
中文介绍 提出渐进式音频编码器压缩框架 X-AuT,通过跨尺度蒸馏选择层组合,而非直接删除整块,以降低语音 LLM 的推理成本。该方法缓解块删除造成的嵌入扰动、内容丢失与过早结束序列问题,在压缩深度的同时保持解码质量。
👍 26
09/10 08:00
Mainstream robotic policies often adopt a Markovian formulation, but many complex real-world manipulation tasks are inherently non-Markovian, requiring long-horizon memory beyond the current observation. Existing memory mechanisms often rely on language summaries, growing visual windows, or their co
中文介绍 机器人策略多采用马尔可夫假设,但真实操作任务常具非马尔可夫性质,需要超出当前观测的长程记忆。论文提出以记忆为规划的世界-动作建模,用记忆锚定规划,替代语言摘要或不断扩张的视觉窗口等现有记忆机制。
👍 23
09/08 08:00
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evalua
中文介绍 分析发现 SWE-Bench Pro 的评测受两类不可靠因素影响:金标准答案或隐藏评测信息泄漏导致的 reward hacking,以及其它评测缺陷。为此提出 SWE-Bench Pro Verified,构建更可靠的基准来衡量软件工程智能体的仓库级能力。
👍 20
09/06 08:00
Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduc
中文介绍 顺序记忆智能体逐块读取长文档,把文档遍历与推理深度耦合,导致对证据位置敏感、延迟随文本长度线性增长。PARSER 提出并行读取、深度推理的范式,解耦遍历与推理,提升长上下文智能体的效率与鲁棒性。
👍 19
09/09 08:00
We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In
中文介绍 提出 MetroLLM-Bench,包含 955 个用例,覆盖六个真实地铁系统(37 至 414 座车站)以及路径规划、票价计算、运营中断、无障碍与对抗输入等十一类任务,用于评测语言模型作为地铁自助终端策略层的实际表现。
👍 19
09/08 08:00
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essent
中文介绍 递归自我改进(RSI)研究多聚焦训练流程自动化,缺少事后监测与审计。论文提出 SAEScientist-Bench,评测 AI 智能体能否自主开展 SAE 可解释性研究,为理解模型学到了什么提供可验证的自动化研究评估。
👍 18
09/01 08:00
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video
中文介绍 Video-LLM 的时序推理基准常经由语言中介,易受选项措辞、答案相关性或语言先验等捷径影响。TempCloze 提出视频完形填空式基准,要求模型识别被挖去的中间片段,从而更纯粹地评测视觉时序推理能力。
👍 18
09/10 08:00
Code world models represent worlds as executable programs, but this representation alone does not determine how to construct a complex world. We introduce Recursive Code World Models (RCWM), a framework for reconstructing complex 3D worlds in code from a single reference image. RCWM couples a Recurs
中文介绍 代码世界模型以可执行程序表示世界,但如何构造复杂世界尚无定论。RCWM 提出从单张参考图像出发,用递归场景程序重建复杂 3D 世界的框架,将递归程序结构与场景生成结合,实现复杂世界的代码化建模。
👍 17
09/10 08:00
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the
中文介绍 在线自蒸馏(OPSD)让 LLM 以自身为教师、借助标准答案等特权信息自我提升,但已有研究发现其会严重损害模型能力。Negative Self-Distillation 转而让模型识别并规避自身推理中的缺陷,以避开错误的方式学习推理。
👍 17
08/28 08:00
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques
中文介绍 注意力模块在极低位宽量化下误差显著,导致 LLM 性能下降,现有方法主要依赖平滑技术缓解。HyQuant 提出面向注意力的混合精度量化方案,在不同部分分配不同位宽精度,在降低量化误差的同时保持训练与推理效率。
👍 17
09/07 08:00
AI research agents combine prior knowledge, public sources, and experimental feedback to produce useful results. The Discovery Certification Protocol (DCP) turns claims about these results into executable recovery and feedback tests. Gate 1 validates useful improvement on sealed evaluation. Gate 2 g
中文介绍 AI 科研智能体的成果声明往往难以验证,仅凭分数不足以证明发现。论文提出 Discovery Certification Protocol(DCP),把声明转化为可执行的复现与反馈测试:Gate 1 在密封评测上验证有效改进,Gate 2 进一步检验其可恢复性。
👍 16
09/09 08:00
Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This is inaccessible for non-English-speaking users, risks losing
中文介绍 推理语言模型的能力高度偏英语中心:无论提示使用何种语言,模型多数仍用英语推理,损害非英语用户的可及性。该工作研究数据混合策略,通过多语言数据配比提升模型以目标语言进行推理的泛化能力。
👍 15
09/10 08:00
Neurosymbolic systems rely on mathematical solvers to guarantee reasoning correctness, yet solvers are fundamentally blind to whether a formal translation maintains strict reference-equivalence to a designated formalization. We formalize this vulnerability as Verdict-Preserving-Unfaithfulness (VPU):
中文介绍 神经符号系统依赖数学求解器保证推理正确性,但求解器无法判断形式化翻译是否与指定形式化严格等价。论文将该漏洞形式化为 Verdict-Preserving-Unfaithfulness(VPU),并提出生成式奖励模型用于自动形式化的质量评估。