👍 102
09/26 08:00
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and
中文介绍 代码智能体 RL 常用可执行测试给出二值奖励,GRPO 会对同一 rollout 组内通过测试的轨迹分配相同优势,忽略实现质量差异。本文提出分组式智能体评分与优势重分配方法,按实现质量细分轨迹并重新分配优势,使训练信号更细粒度。
👍 55
09/28 08:00
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systema
中文介绍 多数多模态大模型依赖预训练视觉编码器提供视觉先验,无编码器 MLLM 则直接从原始像素学习表示,架构更统一。本文系统研究其缩放行为,给出无编码器多模态预训练的缩放规律,并衡量与有编码器方案的差距。
👍 30
09/27 08:00
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventi
中文介绍 VLA 模型具备广泛的操作能力,但在需要诊断失败并调整行为的执行场景中仍显不足。本文提出跨智能体的递归 harness 蒸馏,将强 agent 探索到的有效干预提炼进 VLA 策略,提升任务与环境变化下的机器人操作适应性。
👍 27
09/27 08:00
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental ch
中文介绍 针对单次执行可达数小时、数百次模型-环境交互、近百万 token 的超长程智能体任务,在线 RL 在采样与训练上存在根本挑战。QwenGyre 提出弹性 RL 框架,动态调度资源与并行配置以支撑 xlong 任务的训练。
👍 25
09/28 08:00
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult
中文介绍 胸片图像与放射报告配对预训练颇具潜力,但报告冗长、临床信息密集,现有方法仍依赖任务特定微调。SentZero 提出以句子为中心的视觉-语言预训练框架,通过强化句子级对齐提升多任务零样本胸片分析能力。
👍 17
09/26 08:00
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multipli
中文介绍 快速矩阵乘法在固定乘积下寻找更低成本的计算方式,本文反过来让 Transformer 的学习投影采用另一种更廉价的乘积:基于结合代数构造替换普通矩阵乘法,在不改变参数量的前提下更换层的计算形式。
👍 16
09/27 08:00
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We i
中文介绍 现有编码智能体评测多局限于单一代码库,而真实软件生态中不少功能与修复需要跨仓库协同修改。WideSWE 构建跨仓库协同变更的评测基准,考察编码智能体在多个代码库之间协调实施修改的能力。
👍 15
09/28 08:00
Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We
中文介绍 通用前沿系统正从视觉理解扩展到传统上由专用 CV 模型承担的能力。本文以 GPT-6 Astra 为对象系统评测其计算机视觉表现,界定哪些任务已被通用模型覆盖、哪些仍然困难,为 CV 社区提供能力边界参考。
👍 15
09/28 08:00
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a
中文介绍 基于工具的图像修饰通常由自回归 MLLM 依次生成推理、工具选择与参数值。FlowTool 将该任务重构为生成式建模问题,用 flow matching 直接预测并控制工具参数,替代逐步自回归解码流程。
👍 11
09/28 08:00
As language models take a growing role in AI development, a natural aspiration is for them to reflect on their own learning process, as humans do, and use that reflection to improve themselves. At the same time, these models have an advantage that human learners lack, since training leaves parameter
中文介绍 训练会在参数中留下更新痕迹。本文提出从权重更新中读出信息,并进一步据此实施行为干预,使语言模型能够像人类一样反思自身学习过程并用于自我改进,连接了内部表征读取与下游行为调整。
👍 10
09/28 08:00
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that
中文介绍 本文从强化学习视角研究 on-policy distillation,建立 OPD 中 reverse-KL 目标与 KL 正则化策略优化的联系,并据此提出 Least-Square Policy Distillation(LSPD)框架,提升 LLM 推理训练的样本效率。
👍 9
09/28 08:00
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a c
中文介绍 LLM 的 RL 后训练需在 GPU 集群上协调生成、推理与训练多个模型,运行中资源可用性、序列长度、内存压力与阶段瓶颈均会变化。Nereus 提出自适应并行方案,动态调整并行策略以改善后训练吞吐。
👍 8
09/26 08:00
Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable pro
中文介绍 从经验中学习是构建自演化 LLM 智能体的关键范式,技能合成可将积累经验转化为可复用能力。ExpVoyager 提出直接经验导航机制,支持动态的智能体技能合成,使能力持续扩展而非静态复用。
👍 6
09/27 08:00
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to
中文介绍 验证正确的解对后续 RL 的价值并不相同。本文系统研究 SFT 数据的路径多样性,即推理步骤序列的差异,并提出轻量规则化指纹来筛选多样化 SFT 轨迹,实验显示可提升 RL 之后的泛化性能。
👍 5
09/28 08:00
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of
中文介绍 提出 OLIVE(OnLine InterVEntion):每轮由演化中的学生策略生成前缀,教师自回归续写该前缀,再用教师生成 token 上的交叉熵更新学生,从而在学生自身状态分布上获得教师指导,缓解分布偏移。
👍 5
09/26 08:00
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the l
中文介绍 多智能体协作常出现冲突,例如一个编码智能体修改接口后,另一个仍在旧版本上开发导致测试失效。Relic 研究如何把一次性对话中的协作沉淀为持久的组织级能力,使参与者更替后经验仍可延续。
👍 4
09/28 08:00
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contributio
中文介绍 RLVR 提供可靠的轨迹级信用,OPSD 提供稠密的 token 级监督,但二者在确定步级信用的方向与幅度时相互耦合。本文提出解耦信用方向与幅度的自蒸馏方法,使每步获得可靠的方向与贡献分配。
👍 4
09/28 08:00
Multi-vector retrievers built on vision-language models lead visual document retrieval (VDR), but they run a multi-billion-parameter query encoder on every search. Distilling this encoder into a small student that queries the teacher's existing index would remove the bottleneck. The standard recipe,
中文介绍 多向量视觉文档检索依赖数十亿参数的查询编码器,成为每次检索的瓶颈。ColNanoVDR 利用最优传输把查询编码器蒸馏为小型学生模型,可直接查询教师已有索引,实现无文档的查询蒸馏并显著降低开销。
👍 4
09/28 08:00
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Rep
中文介绍 可靠的安全保障既需要降低不安全行为的控制机制,也需要发现风险的监控机制。本文考察表示工程在 LLM 安全中的作用,分析模型内部表示何时能有效辅助安全控制与监控,并界定其适用条件与局限。
👍 3
09/27 08:00
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-inte
中文介绍 CTC 天然支持语句级监督的离线与流式语音识别,但常规实现需在内存中保存帧与词表的完整激活,使用原生 LLM 词表训练时内存开销过高。本文提出剪枝 CTC,显著降低大词表 ASR 训练的内存占用。