👍 80
09/24 08:00
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the S
中文介绍 研究发现,尽管 LLM 由高度非线性组件构成,却展现出基本线性性质:将来自不同文本流的输入线性组合后,模型输出恰为各自 next-token 分布的叠加,作者称之为 linear superposition,为理解 LLM 的表示与并行处理能力提供证据。
👍 45
09/22 08:00
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to
中文介绍 Prefill 与 decode 对量化的需求不同:低精度算术可加速 prompt 处理,紧凑权重则降低生成阶段的内存流量。本文提出 disaggregated quantization(DQ),为两个阶段分别特化计算格式、权重与存储位置,同时优化预填充与解码效率。
👍 21
09/24 08:00
Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure
中文介绍 Rufus-Air 是在 GLM-4.5-Air-Base(106B-A12B)上开源、可复现的后训练配方,由八个串行阶段组成:SFT、推理 RL、代码 RL、指令遵循 RL、通用 Agent、代码 Agent、搜索 Agent 与 RLHF,并公开数据、奖励设计与基础设施细节。
👍 20
09/25 08:00
Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but selecting the retained blocks remains a bottleneck. Conventional block selection requires scoring all query-block pairs and therefore remains quadr
中文介绍 长上下文自注意力的二次开销可由块稀疏注意力缓解,但传统块选择需为所有 query-block 对打分,仍为二次复杂度并成为瓶颈。本文提出对数线性复杂度的块稀疏注意力方法,降低选块代价,使长上下文建模更具可扩展性。
👍 13
09/23 08:00
Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and
中文介绍 LLM 已能求解许多悬置数十年的数学难题,使数学知识的大规模扩展成为可能,但模型能否提出有价值的新猜想仍不清楚。本文研究如何让模型「学习发现有趣的数学」,以生成与筛选更具研究价值的数学内容。
👍 11
09/24 08:00
Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem insta
中文介绍 任务与运动规划(TAMP)中离散决策与几何、运动学、动力学约束紧耦合,即使完全可观测、状态以物体为中心也依然困难;广义 TAMP 试图利用跨问题实例的规律性。本文提出用 coding agents 求解广义 TAMP 问题。
👍 11
09/24 08:00
We introduce PUBG Ally, an embodied agent for PUBG: BATTLEGROUNDS that can reason, act autonomously, and play alongside players as a voice-enabled teammate. Building such a teammate requires combining two difficult capabilities: it must perceive and respond to a constantly changing game world under
中文介绍 提出 PUBG Ally,一个面向《绝地求生》的具身智能体,能以语音队友身份与玩家并肩作战,自主推理并行动。其难点在于同时具备两项能力:在持续变化的游戏世界中感知与响应,以及与人类进行实时语音交互。
👍 10
09/24 08:00
Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained wit
中文介绍 现有对齐失败检测器多为生成式评判器,每条标准都要消耗一次解码,或如 Llama Guard 每次只输出固定标签。本文提出 Jev,用强化学习训练校准决策模型,以零样本方式检测对齐失败,支持选项、二值判断与打分输出,成本更低。
👍 9
09/24 08:00
Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing
中文介绍 深度搜索要求 LLM agent 分解复杂查询、检索证据并综合有依据的答案,但现有 ReAct 式 agent 存在角色耦合(单一策略兼顾规划、证据使用与合成)与上下文持续累积两大问题。本文提出角色解耦的迭代合成方法 IterSynth 加以解决。
👍 8
09/24 08:00
Jev is a fast, low-cost decision model that answers natural-language questions with choices, binary judgments, and scores. As its public ecosystem grows rapidly, it remains unclear how Jev is used across applications and how public attention relates to project distribution. To answer these questions
中文介绍 Jev 是一个快速、低成本的决策模型,以选项、二值判断和打分形式回答自然语言问题。随着其公开生态迅速扩张,其实际使用方式以及公众关注度与项目分布之间的关系仍不清晰。本文对 Jev 的功能、应用与生态系统进行数据驱动分析。
👍 8
09/24 08:00
In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it
中文介绍 提出 RGBD20K,一个用于推动更鲁棒、更通用的 RGB-D 语义分割的大规模数据集,具备类别丰富、标注质量高等特点,并显著扩展语义空间,为该任务的泛化性与鲁棒性研究提供新基准。
👍 8
09/19 08:00
Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure: two architectures with identical parameter budgets but different depth-width, head, or
中文介绍 Transformer 设计与压缩本质是在预算下分配容量,但常用的 #Params 与 #FLOPs 只反映规模与计算,无法刻画架构结构:参数预算相同而深度宽度、head 配置不同的架构难以区分。本文提出 Neural Spectral Capacity,仅从网络规格衡量并指导架构设计。
👍 7
09/23 08:00
Time series agents answer analytical questions by calling external tools, and which tools they carry is decided by people before the agent runs. However, we identify two failures in this setup. Human-Agent Tool Misalignment: a library of 21 expert-curated tools helps on some tasks and hurts on other
中文介绍 时序 agent 依靠调用外部工具回答分析问题,但工具集由人工预先设定,作者指出两类失败,其中人-代理工具错配表现为 21 个专家精选工具在部分任务上有帮助、在另一些任务上反而有害。TimeEvo 通过失败驱动的自演化让 agent 自行调整所需工具。
👍 7
09/25 08:00
Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-an
中文介绍 现有多智能体基准多聚焦竞争场景、20 步以内的短时程交互,或仅聚合个体表现,难以隔离并凸显 LLM agent 的真正协作能力。本文提出 AgentWorld,一个包含 100 项人工设计任务、面向长时程协作的多智能体基准。
👍 7
09/24 08:00
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a h
中文介绍 工具调用 agent 的输出高度异构,结构化工具调用与面向用户的自然语言总结交错出现;GRPO 等标准 on-policy RL 会把同一优势广播到所有 token,造成跨段信用误分配。本文提出 SLCA-GRPO 以解决这一问题。
👍 7
08/29 08:00
Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an attack surface in which malicious payloads can enter through model inputs and caus
中文介绍 现代 AI agent 频繁跨越信任边界:摄入不可信内容、与特权指令混合、把中间信念写入长期记忆并调用特权工具,形成可被恶意载荷利用的攻击面。本文提出 AgentKernel,一个以信任为原生的 agent 操作系统,在系统层面提供隔离与管控。
👍 6
09/24 08:00
Long-horizon reasoning remains a central challenge for large language models (LLMs) under sparse-reward regimes. We argue that this brittleness arises from two biases induced by complex reasoning spaces: an exploration bias, where models are drawn toward locally plausible but structurally unstable b
中文介绍 长时程推理在稀疏奖励下仍是 LLM 的核心难题。作者认为其脆弱性源于复杂推理空间带来的两种偏差:探索偏差使模型偏向局部看似合理但结构不稳定的分支。SAGE 通过拓扑引导缓解这些偏差,提升长时程推理的稳定性。
👍 6
09/23 08:00
World-action models (WAMs) transfer visual and motion priors from pretrained video generators to robot control by jointly modeling visual dynamics and actions. Existing WAMs, however, predict dense future frames during training, repeatedly modeling largely unchanged content and coupling action-condi
中文介绍 World-action model 通过联合建模视觉动态与动作,把预训练视频生成器的视觉与运动先验迁移到机器人控制;但现有方法训练时预测稠密未来帧,反复建模几乎不变的内容并使动作条件相互耦合。本文提出 DeltaWAM,面向双臂操作只建模帧间变化。
👍 4
09/24 08:00
Banking assistants must use account-specific information to answer requests and, in many cases, take actions through tools. Evaluating only the final response misses important errors. An assistant may ask for information it already has, rely on stale context, select the wrong account, or write an in
中文介绍 银行助手须结合账户专属信息作答,并常需通过工具执行操作,仅评测最终回复会漏掉关键错误,例如索要已知信息、依赖过期上下文、选错账户或写出不当内容。本文提出 IndicBankBench,评测印度零售银行场景下语言模型助手的安全性与可靠性。
👍 4
09/20 08:00
Adapting large language models to individual users remains challenging due to the tension between fine-grained personalization and scalable deployment. We present CARD, a hierarchical framework that achieves effective personalization through progressive refinement. CARD first clusters users accordin
中文介绍 将大模型适配到单个用户面临细粒度个性化与可扩展部署之间的张力。CARD 采用层次化框架,先按用户聚类,再通过奖励引导解码逐步细化,实现高效且可扩展的个性化文本生成。