👍 202
09/28 08:00
Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to superv
中文介绍 潜在视觉推理(LVR)让多模态大模型在连续 latent token 中完成中间计算,但推理过程不可观测,难以为其提供监督。论文重新审视该范式,主张将潜在推理锚定在可观测的视觉证据上,以提升可监督性与可信度。
👍 186
09/30 08:00
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward imp
中文介绍 自演化搜索 agent 通过同时优化出题者与解题者来构建训练课程,却产生「co-cheating」失效模式:两者逐渐在共享错误上达成一致,使内部奖励信号失真。论文诊断该现象并提出缓解方法,以恢复自演化训练的有效性。
👍 98
09/30 08:00
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better a
中文介绍 终端 agent 依赖随机生成的动作,但能生成有用命令并不保证其可靠执行,错误的包安装等操作会改变环境、阻碍后续进展。Mid-Harness 在模型与执行 harness 之间扩展动作空间,以提升终端任务执行的可靠性。
👍 63
09/29 08:00
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when prediction
中文介绍 LLM agent 在多步决策中难以迁移到未见环境,基于世界模型的方法虽可缓解,却需额外训练且预测误差会累积。EVOKE 通过从 LLM 自身引出世界知识来指导决策,无需额外训练即可提升迁移能力。
👍 31
09/29 08:00
Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt t
中文介绍 长篇幅字幕翻译需跨集甚至整季推理话语与文化语境,并保持术语和风格一致;现有单 LLM 方法多停留在句子级,多智能体系统则依赖静态工作流。Breaking Babel 提出自演化多智能体系统,实现自适应的长字幕翻译。
👍 30
09/29 08:00
Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and becau
中文介绍 肌肉骨骼建模与强化学习已能让肌肉驱动的智能体复现复杂人体动作,但受限于平地场景,原因之一是运动数据缺少对齐的地形几何。TERRA 提出地形感知的重建、重定向与控制框架,将肌肉骨骼运动能力扩展到复杂地形。
👍 24
09/29 08:00
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-traini
中文介绍 LoopVL 探索能否将 Loop Transformer 有效扩展到视觉-语言模型:通过 Module-Loop 与 Model-Loop 计算,用共享模块迭代更新统一的视觉-语言状态,并从头经语言预训练等阶段完成训练,验证循环结构在 VLM 中的可行性。
👍 24
09/30 08:00
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchma
中文介绍 科学软件对基于视觉语言模型的计算机使用 agent 构成严峻考验:需解读专业界面、操作科学对象并产出可验证结果。OSWorld-Science 提出相应基准,用于评估 agent 学习与使用科学软件的端到端能力。
👍 16
09/26 08:00
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods ei
中文介绍 奖励模型常偏好长度、自信等表面属性,使 LLM 产出评分更高但未必更正确的回答,现有缓解方法各有限制。BiasReducer 提出自适应偏差缓解方案,动态抑制奖励模型中的表层偏差,提升偏好信号的可靠性。
👍 15
09/29 08:00
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated m
中文介绍 现代 agent 系统由 AI 模型与控制执行的 harness 共同组成,harness 设计显著影响长程表现,但其组合搜索空间需大量人力且随模型更替反复进行。MILO 通过编排式多智能体演化,实现 harness 的自动发现。
👍 15
09/30 08:00
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In t
中文介绍 循环 Transformer 以递归增加固定参数下的计算深度,MoE 以稀疏性扩大固定激活算力下的容量,但已有缩放律只单独建模其一。论文提出循环 MoE 的联合缩放律,刻画递归与稀疏性共同作用下的效率缩放规律。
👍 12
09/26 08:00
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studi
中文介绍 软件开发不仅是编辑代码,还需反复运行程序、与界面交互并视觉检查行为,而现有编码 agent 与计算机使用 agent 对此研究不足。CUA-SWE 将计算机使用 agent 引入视觉软件工程场景,构建相应任务与评测。
👍 7
09/29 08:00
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive ref
中文介绍 基于自回归语义建模的 TTS 零样本音色克隆表现优异,但顺序解码延迟高;非自回归方案更快,却依赖更受限的参考条件。Tacit-TTS 从自回归解码转向掩码预测,实现无需转录文本的高效语音克隆。
👍 6
09/30 08:00
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the
中文介绍 潜在通信让多智能体系统直接在内部表示空间交换信息,降低 token、计算与延迟开销,具体做法是用轻量可训练连接映射发送方表示。论文研究这种潜在通信的安全性,分析其可能引入的风险与防护方向。
👍 6
09/30 08:00
Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dyna
中文介绍 自注意力为 LLM 提供细粒度、查询相关的上下文访问,但密集 token 交互带来二次预填充开销与随上下文增长的 KV cache。该综述梳理注意力机制演进,涵盖显式记忆压缩、稀疏访问、循环状态与结构化状态动力学等方向及其权衡。
👍 5
09/30 08:00
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a
中文介绍 研究如何从自然语言反馈训练 LLM 评委,尤其针对评判标准取舍与加权高度影响结果的主观任务。主流结果监督 RL(如 GRPO)对 rollout 中所有 token 统一分配信用,本文提出位置选择性自蒸馏以更精准利用反馈。
👍 5
09/28 08:00
Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representati
中文介绍 LLM 路由旨在为每个查询选择最合适的模型。现有路由器多直接从查询嵌入、模型表示、偏好数据或相似样本聚类中学习决策,表征能力有限。SeLMRoute 提出基于概率语义证据的路由方法。
👍 4
09/28 08:00
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so
中文介绍 AI agent 有望自动化生物图像分析,但尚无基准评估其能否端到端完成真实分析任务——2D 图像、3D 体数据与延时序列往往过大,难以直接读入上下文。BIABench 提出面向真实生物图像分析任务的 agent 评测基准。
👍 4
09/30 08:00
We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length w
中文介绍 论文发现非侵入式脑信号解码单词的主要性能提升,在很大程度上不使用任何脑数据也能复现:d'Ascoli 等(2025)将脑活动时间序列切分为固定长度窗口,引入了可利用的时序捷径。移除该捷径后,脑到文本解码结果更为可信。
👍 4
09/29 08:00
Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this
中文介绍 技能为 LLM agent 提供专业知识以完成长程复杂任务,但如何合成可靠训练数据、如何训练 agent 使用技能仍待探索。SkillGym 通过自动生成可验证环境,为技能使用型 agent 提供可扩展的训练方案。