👍 173
08/24 08:00
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this working capability: sustained, veri
中文介绍 提出 Apodex 1.1,聚焦大模型在复杂工作中的“工作能力”——对文件、信息源和可执行代码的持续交互,以及状态维护、失败恢复和可验证交付,使模型不仅能推理,还能完成多步骤实际任务。
👍 55
08/21 08:00
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scen
中文介绍 提出 TLive-Omni,面向电商直播的全模态理解模型,融合语音、视频帧、商品图像、叠加文本和用户查询,在嘈杂且时序延长的直播流中抽取商品事实,服务实时导购与问答场景。
👍 34
08/24 08:00
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. Yet existing benchmarks fall into two camps, each with a critical blind spot: GUI-centric benchmarks test surface-level scre
中文介绍 提出 MobilePA-Bench,面向移动端规划智能体在复杂真实任务上的能力评测。现有基准分 GUI 导向与任务导向两类,各有盲区;该基准强调对系统状态、应用内操作和长期目标的综合规划。
👍 33
08/17 19:24
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token con
中文介绍 提出 ParaTempo,用时间置信度调度并行推理路径,替代依赖最终答案共识或局部 token 置信度的路径管理方法,在保留大推理模型多路径探索收益的同时,降低随推理深度和分支数增长的计算成本。
👍 32
08/24 08:00
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows. A persistent IPython REPL follows the Recursive Lang
中文介绍 Prime Agent 是一个开源的长周期智能体评测与编码工作流 harness,通过持久化 IPython REPL 并遵循递归语言模型机制扩展上下文,支持自我改进,适合需要持续交互、外部计算和可靠交付的智能体任务。
👍 28
08/21 08:00
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, us
中文介绍 提出 OmniAssistBench,面向 omni-LLM 的助手式交互基准。与被动视频理解不同,它评测模型在实时视频场景中持续感知环境,结合视觉状态和用户指令主动引导用户完成目标的能力。
👍 27
08/20 08:00
While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging. Existing text-to-3D methods either decode discrete shape tokens autoregressively or iteratively refine global 3D representations with diffusion or flow models. However, auto
中文介绍 提出 Block3D,通过分块扩散实现文本到 3D 生成,避免自回归离散 token 或全局迭代扩散的高推理成本,在保持几何保真度的同时提升生成效率。
👍 15
08/24 08:00
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dro
中文介绍 针对 LLM 策略优化中稳定性和探索的两难,提出环境正则化方法:不再单靠 action-side Policy-KL 约束响应行为,而是对环境侧进行正则,释放被占用的探索预算,同时维持策略稳定性。
👍 15
08/13 08:00
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be
中文介绍 提出 ARC,面向开放式真实交互中多种合法行为并存的情况,解决 group-based RL 因组内 rollout 不可直接比较而失效的问题,实现更公平的相对优势比较,提升策略优化信号质量。
👍 15
08/17 08:00
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a control
中文介绍 通过对照实验揭示 on-policy distillation 的泛化双面性:模型可能获得向训练分布外迁移的正向泛化,也可能因学生轨迹分布偏差产生负面泛化;强调 OPD 评估需覆盖多域和分布外基准。
👍 13
08/18 08:00
Game world models have recently demonstrated promising capabilities in generating visually coherent and action-controllable gameplay videos. However, non-player character (NPC) behavior in existing models is either implicitly entangled with video generation or explicitly prescribed through external
中文介绍 提出 WorldMind,将游戏世界模型与 NPC 行为解耦,使 NPC 行为由显式世界状态驱动,而非隐式纠缠在视频生成或外部脚本中,从而实现状态感知、动作可控的游戏角色行为。
👍 12
08/15 08:00
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing f
中文介绍 提出 LongRCA Bench,用于诊断长周期智能体执行失败:在结果级评估之外,自动定位最早的决定性错误步骤,并判定责任角色,减少开发者逐段检查完整轨迹的成本。
👍 12
08/22 08:00
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifa
中文介绍 提出 GameXpert-Bench,评测 coding agents 从自然语言需求构建完整游戏的程度,覆盖程序逻辑、视觉音频内容、界面、交互与可玩性的整体集成,衡量其接近专家级游戏开发的水平。
👍 9
08/21 08:00
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, s
中文介绍 提出量化感知修复方法,恢复经结构压缩和 4-bit 量化后的 LLM 在推理、数学、代码和长上下文任务上的性能下降,为低成本模型部署提供实用修复流程。
👍 9
08/20 08:00
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing inform
中文介绍 提出 Thinkingbox,一个有状态业务工作流智能体的沙箱与评测基准。它强调可靠完成多步骤任务不仅需要生成文本或工具调用,还要主动收集缺失信息、维护状态,并以多次成功而非单次成功衡量可靠性。
👍 9
08/21 08:00
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows wh
中文介绍 提出 AgentMercury,让智能体大规模合成可验证的业务场景环境,突破人工构建或围绕预定义任务生成的 task-centric 局限,为训练和评测适应现实演化工作流的智能体提供可扩展环境。
👍 8
08/23 08:00
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both
中文介绍 提出 AutoResearch 两阶段系统,将想法生成与想法执行相连,确保自主研究系统在长流程自动化执行中保持科学 grounding,减少幻觉,实现从洞察到可验证研究产出的闭环。
👍 8
08/19 08:00
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progre
中文介绍 提出按推理进展过滤 on-policy distillation 训练样本的方法,不再把教师奖励直接视为推理进步信号,而是基于学生轨迹中的实际推理进展筛选监督,减少对教师行为的简单模仿。
👍 7
08/24 08:00
We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing app
中文介绍 提出 Task-CoEvolve,通过自适应验证任务选择优化 LLM harness 代码。它迭代改写 harness 而不更新模型权重,并智能挑选验证任务,以更少评估成本取得更大的性能提升。
👍 7
08/20 08:00
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it. The result keeps full attention in only 6 of its 18 bloc
中文介绍 Daedalus-150M 是为 CPU 推理设计的 150M 参数卷积-注意力混合模型:以单用户逐 token、4-bit 权重和普通 CPU 为约束,18 个 block 中仅 6 个保留完整注意力,兼顾效率与质量。