👍 135
09/17 08:00
The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transf
中文介绍 面向长时程 agent 带来的输入密集型负载,论文针对 prefill 计算开销高、大 KV cache 持续挤占 HBM 与 SSD 容量和带宽的问题,在既有长上下文降本工作基础上进一步压缩 KV cache,提出 DeepSeek-V4.1-Flash,以降低长上下文推理与数据搬运成本。
👍 114
09/16 08:00
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unif
中文介绍 评估 Omni-Modal 生成模型是否真正理解物理世界:以在共享隐空间中统一多模态上下文理解与音视频联合生成的 MiniMax-H3 为对象,考察这类统一建模文本、图像、视频、音频的模型在物理常识与因果推理上的能力与局限。
👍 94
09/17 08:00
We study length inflation in on-policy distillation (OPD), where student responses can become excessively long and even exhaust the generation budget. We identify termination-token mismatch between base students and post-trained teachers as an important source of this behavior. Across Qwen3, Llama,
中文介绍 研究 on-policy distillation(OPD)中的长度膨胀:学生回复过长甚至耗尽生成预算。作者在 Qwen3、Llama 等模型上发现,base student 与 post-trained teacher 之间的终止 token(EOS)不匹配是重要成因,为诊断和抑制该现象提供依据。
👍 94
09/17 08:00
As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback. Token efficiency therefore becomes important for scaling recursive self-improvement. We take an
中文介绍 编码 agent 正从监督式代码补全转向无人值守的持续探索,长轨迹中的 token 效率成为递归自我改进能否扩展的关键。SoL-Pi 通过递归扩展自动研究循环来构建更高效的 agent harness,以在同等 token 预算下获得更多改进。
👍 79
09/16 08:00
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the sci
中文介绍 科学代码库以可执行模型、方法与工具的形式承载数十年人类知识,但工具链碎片化、领域约定隐式、正确性判据专门,难以转化为可靠学习经验。ScienceIDE 将全球科学代码库改造成 agent 可学习的交互环境,以弥合这一科学代码鸿沟。
👍 71
09/17 08:00
Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparison
中文介绍 coding harness 决定自主编码 agent 能否把模型能力转化为长时程软件工程表现,但既有工作多把 harness 当作整体评估,各组件作用不明。本文面向组件级对比开展实证研究,拆解并比较 harness 各设计要素对 agent 性能的影响。
👍 56
09/17 08:00
World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic fram
中文介绍 现有预测模型多局限于单一领域。JEPA-Anything 提出领域无关的世界建模框架,试图用统一的学习原则,让预测模型跨越物理系统、视觉场景等差异极大的「世界」进行建模,从而具备预测后果、指导干预并从交互中学习的能力。
👍 56
09/15 08:00
Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what to retry. Existing confidence estimators, however, share one design premise: th
中文介绍 可靠部署 LLM 需要校准的置信度估计,以决定输出是上线、升级处理还是重试。现有置信度估计器多共享同一设计前提(单次推理)。本文提出经验驱动的置信度估计,让模型从推理到 agent 的多轮经验中判断自身输出的正确概率。
👍 50
09/16 08:00
Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating codi
中文介绍 ProgramDistill 是一个新基准:让 coding agent 从可运行的 web 应用中推断其行为,并在功能不完整的应用里复现该行为,把交互式网页转化为可验证、有参考实现的 SWE 任务,弥补以往仅凭 issue 或指令描述目标行为的评测局限。
👍 48
09/15 08:00
LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others'
中文介绍 LLM 助手常被用于日常社交建议,但其社交推理难以评估:既需要助手从主观的用户叙述中理解社交情境,又涉及他人感受等难以验证的社交属性。本文提出面向 LLM 助手的可验证社交推理,使这类咨询场景下的能力可被检验。
👍 45
09/17 08:00
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge containe
中文介绍 检索质量高度依赖索引键能否有效暴露文档所含知识,而传统索引一经构建便固定不变。本文提出自演化搜索索引(Self-Evolving Search Index),让索引随 LLM agent 的多样化信息需求持续更新键表示,提升复杂任务下的检索效果。
👍 45
09/16 08:00
Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratch, so more agents tend to mean more duplicated search rather than more discovery. Agora is a shared memory for such agents: resear
中文介绍 AutoResearch 类自主研究循环可让单个编码 agent 无人值守地改进训练配置,但多个循环各自从零开始,agent 越多重复搜索越多。Agora 以 Git 作为共享内存,让多个研究 agent 共享代码、结果与发现,把并行搜索转化为累积式发现。
👍 42
09/15 08:00
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text
中文介绍 平台滥用者用 emoji、谐音、拆字与冗余符号混淆跳转指令,再经伪装链接把用户导流至色情、诈骗、赌博等非法服务。RiskChainBench 面向混淆平台消息还原与证据溯源的网页调查,弥补现有基准只评测混淆文本还原的不足。
👍 39
09/17 08:00
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This rec
中文介绍 多轮 agent 的 RL 仅提供每条轨迹一个标量奖励,自 on-policy 蒸馏(OPD)用具备特权技能的 self-teacher 补充 token 级稠密监督。RetireOPD 引入自退役机制:学生内化技能后教师自动退出,避免长期依赖特权教师。
👍 37
09/15 08:00
GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before execution. Recent agent-skill frameworks encapsulate reusable procedural knowledge to mitigate this, yet existing skill designs are
中文介绍 GUI agent 在动态界面上执行长时程任务,弹窗、延迟加载与控件移位常使执行前固定的计划失效。现有技能框架封装可复用过程知识,但设计僵化。本文提出免训练的技能演化范式——反思、修正、复用,让技能在运行中自适应更新。
👍 35
09/17 08:00
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address
中文介绍 文档解析需在各种版式与拍摄条件下稳定输出结构化内容,但训练语料偏向常见文档类型与干净数字页面,单纯扩大覆盖并不能说明能力提升。WeVisDoc 从覆盖转向能力,提升端到端文档解析的鲁棒性。
👍 34
09/15 08:00
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signal
中文介绍 LLM 交易 agent 能整合行情、新闻与可执行分析,但其行为常由部署前写死的工具使用策略控制,难以调整取证与验证方式。EvolveTrade 提出经验驱动的策略精炼,让交易 agent 依据交易经验自我演化工具调用与决策策略。
👍 33
09/17 08:00
Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose
中文介绍 空间智能不止于描述物体位置:观测不完整时,模型还需识别并主动获取缺失证据、在统一空间框架中解读并据此行动。VABench 通过视觉演示、主动感知与度量控制,评测「观察—推理—行动—修正」的完整闭环。
👍 32
09/17 08:00
Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on e
中文介绍 大型推理模型常系统性地低效:简单题过度思考、难题思考不足。统一的长度惩罚或刚性路由会带来效率税,牺牲简单题上的节省或难题上的性能。When2Think 提出难度感知的长度控制,让混合推理模型按题目难度动态分配思考预算。
👍 30
09/17 08:00
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution,
中文介绍 on-policy self-distillation(OPSD)让语言模型从能看到答案或解题过程的自身副本中学习,但这份特权信息究竟带来多少增益尚不明确。本文将其贡献与蒸馏本身分离,量化特权信息对 OPSD 的净增益及其作用条件。