👍 203
09/23 08:00
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video
中文介绍 探究视频生成模型作为世界模型是否具备「物体永久性」与「固体性」等人类认知先验,并提出训练方法让世界模型习得该能力。作者将物体永久性视为可通过训练获得的认知能力,为构建具备类人物理智能的视频世界模型提供新思路。
👍 84
09/22 08:00
Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change o
中文介绍 多方对话的长期记忆不仅要检索相关内容,还须区分谁说、涉及谁、个体间认知与群体共享信息及状态变化。为此提出 SpeakerMem-R1,以说话人为中心构建双轨记忆,用于多主体场景下的长期记忆建模与推理。
👍 75
09/24 08:00
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the S
中文介绍 发现 LLM 虽由高度非线性组件构成,却表现出基础线性性:将不同文本流的输入线性组合后,模型的 next-token 分布恰为各分布之叠加,作者称之为线性叠加现象,为理解与操控 LLM 内部表征提供了新视角。
👍 48
09/19 08:00
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an upd
中文介绍 针对 VLM 在动态环境中空间推理的不足,提出 Spatial-Interactor,让模型通过与可观测物理世界交互学习空间推理。方法要求模型感知物体运动与视角变化引起的局部状态转移,并在长轨迹上整合以维持更新后的场景表征。
👍 43
09/21 08:00
Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agent
中文介绍 提出 HappyWorld-Bench,一个面向世界模型的综合基准,不仅评估生成世界的质量,还检验其在探索、交互与修改下的一致性与响应性,即生成世界作为智能体环境时是否依然可靠,为交互式能力评估提供统一标准。
👍 39
09/23 08:00
Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generat
中文介绍 自回归视频生成通过因果展开延长视觉序列,但随生成推进出现记忆瓶颈,难以维持长时一致性。该工作提出以「过去」框定未来的记忆机制,缓解长时程生成与交互式世界建模中的误差累积问题。
👍 35
09/23 08:00
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by sim
中文介绍 现有 agent 记忆系统多在写入时固化:任务结束后把轨迹蒸馏为反思、工作流或技能等固定产物,检索时缺乏灵活性。论文提出 Just-in-Time Memory,让智能体在推理时按任务需求实时筛选与整合适配记忆。
👍 22
09/20 08:00
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address th
中文介绍 人类可轻松定位声源方向并与视觉线索融合推理,具身智能体却难以做到。论文提出 OmniEcho,旨在评估与建模具身场景下的空间音频理解,推动声音定位与视觉信息的联合推理能力。
👍 19
09/23 08:00
Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool res
中文介绍 现有语言世界模型通常预测环境观测,难以重构高熵、依赖执行细节的工具返回结果。论文提出 Agent-Editing World Model,重新思考 LLM 智能体中的世界建模方式,以编辑而非完整生成的形式刻画环境状态变化。
👍 17
09/24 08:00
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure
中文介绍 Rufus-Air 是开放可复现的 LLM 后训练配方,基于 GLM-4.5-Air-Base(106B-A12B)构建八阶段串行流水线:SFT、推理 RL、代码 RL、指令遵循 RL、通用/编码/搜索 Agent 及 RLHF,并公开数据、奖励与基础设施细节。
👍 16
09/22 08:00
Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completen
中文介绍 RL 已成为 LLM 后训练核心,但 token 级 credit assignment 缺乏公认的数学定义,与常用训练信号的关系不清。论文提出完整性等三条正则性条件,并据此提出 PACT,将信用分配问题转化为 critic 对齐。
👍 15
09/22 08:00
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externaliz
中文介绍 前沿模型能力提升常被归因于推理增强,但闭源系统的原始 CoT 不可见。作者借助标准 API 注册自定义工具,诱导前沿模型外化隐藏的思维链,进而提取并刻画其推理过程,揭示模型「有能力却精简」的特性。
👍 12
09/23 08:00
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an enviro
中文介绍 语言模型智能体面临状态演化、决策耦合与延迟反馈的长时程任务,扩大训练需要多样环境与可靠结果信号。论文提出可验证隐藏动态玩法,从已求解的机制出发自动生成智能体 RL 环境,降低环境扩展成本。
👍 11
09/23 08:00
Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and
中文介绍 LLM 已能解决诸多高难数学问题,甚至包括悬置数十年的开放问题,但能否自动发现「有趣」的数学仍待探索。论文研究让模型学习提出有价值的数学猜想与问题,以期大规模扩展数学知识。
👍 11
09/24 08:00
Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem insta
中文介绍 任务与运动规划(TAMP)即便在完全可观测、物体中心状态下依然困难,因为离散决策与几何、运动学及动力学约束紧密耦合。论文利用跨问题实例的规律性,用编码智能体求解广义 TAMP 问题。
👍 11
09/24 08:00
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI
中文介绍 随着 LLM 从被动内容生成走向工程与科学发现,论文提出 Qwen-Planner-Agent,一个闭环 AI-for-AI 框架,让 AI 既是被开发对象又参与构建下一代移动 planner 智能体,面向真实场景的规划任务。
👍 11
09/20 08:00
Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned throu
中文介绍 机器人装箱需长时程序贯决策,每次放置都会改变后续可用空间,现有方法多依赖手工几何启发式或强化学习策略。论文提出 PackLab,一个用于多模态大模型在机器人装箱任务中开发、训练与评估的综合框架。
👍 11
09/23 08:00
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on
中文介绍 VLA 模型为通用机器人控制提供了坚实基础,但多数策略仅依赖当前观测,不保留整段 episode 的历史信息,在依赖历史操作的任务中受限。论文提出 MemBodied,以循环联想记忆为 VLA 引入情节级记忆。
👍 10
09/23 08:00
We present Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. The model is pretrained on a rig
中文介绍 Hunyuan-A13B 是开源 MoE 大语言模型,总参数 80B、推理时仅激活 13B,在模型能力、计算效率与部署成本之间取得平衡。技术报告公开其预训练数据与架构设计等细节,并给出完整评测结果。
👍 9
09/24 08:00
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing
中文介绍 深度搜索要求智能体分解复杂查询、检索证据并合成有依据的答案,但现有 ReAct 式智能体存在角色耦合与上下文累积两大问题。论文提出 IterSynth,通过角色解耦的迭代合成重构深度搜索智能体。