👍 95
09/17 08:00
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transf
中文介绍 面向长时程 agent 的输入密集型负载,prefill 仍然昂贵,KV cache 持续挤占 HBM 与 SSD 容量及带宽。DeepSeek-V4.1-Flash 进一步推进 KV cache 压缩极限,缓解长上下文推理的显存与存储压力。
👍 93
09/16 08:00
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unif
中文介绍 评估统一建模文本、图像、视频与音频的 Omni-Modal 生成模型 MiniMax-H3 能否真正推理物理世界。该模型在共享潜空间中结合多模态上下文理解与音视频联合生成,论文围绕其物理世界推理能力展开系统性评测。
👍 75
09/17 08:00
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama,
中文介绍 研究 on-policy distillation(OPD)中的回答长度膨胀现象:学生输出过长,甚至耗尽生成预算。作者发现基座学生与后训练教师之间的终止 token(EOS)不匹配是重要成因,并在 Qwen3、Llama 等模型上验证。
👍 74
09/16 08:00
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the sci
中文介绍 科学代码库以可执行模型、方法与工具承载数十年知识,但工具链碎片化、领域约定隐式、正确性标准专门化,难以转化为可靠的学习经验。ScienceIDE 将全球科学代码库改造为 agent 可学习的环境,为科学智能体提供训练与评测基础。
👍 69
09/17 08:00
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an
中文介绍 针对 coding agent 从有监督补全走向全天候无人探索、工作轨迹不断变长带来的 token 效率问题,提出 SoL-Pi,通过递归扩展 auto-research 循环提升 agent harness 效率,支撑递归式自我改进的可扩展性。
👍 55
09/17 08:00
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparison
中文介绍 现有工作多把 coding harness 当作整体系统评估,难以判断各组件各自的贡献。本文对编码 agent 的 harness 设计开展实证研究,实现组件级对比,厘清各模块对长时程软件工程性能的实际作用。
👍 54
09/15 08:00
Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: th
中文介绍 可靠部署语言模型需要校准的置信度估计,以决定输出是否上线、升级或重试。已有估计器共享同一设计前提,忽视了模型在交互中积累的经验。本文提出经验性置信度估计,覆盖从推理到 agent 的场景。
👍 48
09/16 08:00
Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating codi
中文介绍 提出 ProgramDistill 基准,评估编码 agent 能否从可运行的 web 应用中推断行为,并在不完整的应用里将其实现,从而把交互式网页转化为可验证、参考引导的 SWE 任务,贴近真实 web 开发场景。
👍 43
09/16 08:00
Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: resear
中文介绍 AutoResearch 等自主研究循环能让单个编码 agent 无人值守地改进训练设置,但多个循环并行时各自从头开始,agent 越多反而重复搜索越多。Agora 以 Git 作为共享记忆,使研究进展在 agent 之间累积复用。
👍 40
09/17 08:00
World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic fram
中文介绍 世界模型让智能体预判后果、指导干预并从交互中学习,但现有预测模型多为特定领域定制。JEPA-Anything 提出领域无关框架,尝试以共同的学习原则在差异巨大的系统之间训练预测模型,实现跨领域世界建模。
👍 33
09/15 08:00
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others'
中文介绍 LLM 助手常被用于日常社交建议,但其社会推理难以评估:助手需从主观的用户叙述中理解情境,还涉及他人隐私等社会属性。本文提出可验证的社会推理评测方案,用于检验此类咨询场景下的助手表现。
👍 32
09/15 08:00
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signal
中文介绍 LLM 交易 agent 可整合行情、新闻与可执行分析,但其工具使用策略多为部署前写死的静态规则,难以自适应地收集证据、调用工具与验证信号。EvolveTrade 以经验驱动的策略精炼,让交易 agent 自我演化。
👍 29
09/15 08:00
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text
中文介绍 平台滥用活动借助 emoji、谐音、拆字与冗余符号隐藏跳转指令,再经伪装链接导流至色情、诈骗、赌博等非法服务。RiskChainBench 针对混淆平台消息还原与基于证据的 web 调查构建基准,弥补现有混淆文本评测的不足。
👍 26
09/17 08:00
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge containe
中文介绍 检索质量取决于索引键能否充分暴露文档所载知识。本文提出自演化检索索引(Self-Evolving Search Index),让索引表示随检索需求迭代更新,以更好支撑 LLM agent 处理信息需求多样的复杂任务。
👍 26
09/17 08:00
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This rec
中文介绍 多轮 agent 的 RL 训练每条轨迹仅得一个标量奖励,self on-policy distillation 借具备特权任务技能的教师提供 token 级密集监督,使无技能学生内化这些能力。RetireOPD 提出自退休的 on-policy distillation 用于 agentic RL。
👍 26
09/14 08:00
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affec
中文介绍 科学 agent 通过整合证据、评估提案来发现假设,近期系统将其与演化搜索结合,经批评、比较与修订迭代。HypoEvolve 用遗传算法组织多 agent LLM 协作,考察不同协作形式对科学假设发现效果的影响。
👍 25
09/15 08:00
GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before execution. Recent agent-skill frameworks encapsulate reusable procedural knowledge to mitigate this, yet existing skill designs are
中文介绍 GUI agent 在动态界面上执行长时程任务,弹窗、延迟加载与控件移位常使预先制定的计划失效。Reflect, Revise, Reuse 提出免训练的技能演化方法,通过反思、修订与复用持续更新可复用的程序性知识。
👍 23
09/17 08:00
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address
中文介绍 文档解析需把文档图像转为结构化内容,并在多样版式与采集条件下保持可靠;但训练语料偏向常见文档类型与干净数字页,仅扩大覆盖并不足够。WeVisDoc 从覆盖率转向能力建设,提升端到端文档解析鲁棒性。
👍 22
09/16 08:00
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margin
中文介绍 直接偏好对齐方法因计算与内存高效而被广泛用于 LLM 与人类偏好对齐,但似然位移问题促使研究者寻找在似然差距较小时仍能提取偏好信息的新途径。本文提出面向 LLM 偏好对齐的零阶优化范式。
👍 20
09/17 08:00
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose
中文介绍 空间智能不止于描述物体位置:在观测不完整时,模型需识别并获取缺失证据、在统一空间框架中解读并据此行动。VABench 通过视觉演示、主动感知与度量控制,评测「观测—推理—行动—修正」完整闭环。