👍 234
10/01 08:00
A companion that talks with a person for months should come to understand them. It should remember what they said, infer who they are, and know when the past bears on the message in front of it. Testing this requires a real person's record, and such records are private, so benchmarks generate the pe
中文介绍 针对陪伴型 AI 需长期理解用户的能力评测,提出 RealCompanion 基准:以真实人的长期对话记录为基础(因隐私限制采用生成式替代),考察记忆、人物推断与上下文关联等推理能力,用可核验答案衡量模型对用户的理解程度。
👍 109
09/30 08:00
Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can lea
中文介绍 论文追问蛋白质折叠的学习能否泛化到更广泛的推理能力:以蛋白质折叠为测试床,一个解析出的结构可生成数千条可精确校验的空间与拓扑陈述,据此训练并评估 LLM 的结构化推理,检验其迁移效果。
👍 85
09/29 08:00
Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs
中文介绍 针对 VLA 模型零样本泛化差、依赖专用机器人数据训练的问题,MotorMind 提出以通用 VLM 为骨架、补充运动相关模块的框架,使通用视觉语言模型可直接用于零样本机器人操作,提升新任务与新环境中的表现。
👍 69
09/29 08:00
The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for impr
中文介绍 论文研究 on-policy 后训练泛化能力强的成因,指出以往工作仅把参数更新行为视为副产品,转而将其作为优化原则,分析 on-policy 参数更新方向对泛化的决定性作用,并据此设计改进训练的方法。
👍 44
10/02 08:00
We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, pub
中文介绍 提出 HyperBrowseComp,面向网页浏览智能体的多语言、多模态压力测试基准,含 423 道由母语者编写并人工校验的题目,覆盖 13 种语言,每题要求给出简短且可公开核验的答案,难度极高,用以暴露现有 agent 的短板。
👍 42
10/02 08:00
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and sh
中文介绍 针对 masked diffusion 语言模型在去噪中少数位置确定后不确定性骤降、信用分配困难的问题,Pivot-SD 提出高效自蒸馏方法,聚焦关键决策位置施加监督,在复杂推理任务上提升生成质量并降低计算开销。
👍 42
09/28 08:00
Physical trajectories contain more than snapshots of a system: they also reveal how its states evolve under governing conditions. However, representation learning for parametric partial differential equations (PDEs) has largely relied on reconstruction-based objectives that emphasize recovering obse
中文介绍 针对参数化 PDE 的表示学习多依赖重建目标、只关注单帧观测的局限,PDE-JEPA 提出预测式表示学习框架,在隐空间建模状态随控制条件的演化动力学,从而学到更有利于下游预测的表征。
👍 39
09/30 08:00
Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce L
中文介绍 针对法律语言模型奖励信号粗糙、缺乏领域特异性与可解释性的问题,LexReward 提出基于分类体系(taxonomy)的奖励框架,从答案正确性与多维法律回答质量两方面评分,为法律领域后训练提供更细粒度的奖励。
👍 39
09/30 08:00
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the
中文介绍 针对 AI 生成科学论文中难以被现有 token 级检测器识别的「学术垃圾」,论文构建相应基准并刻画其模式:各部分看似合理但整体缺乏实质,同时给出检测与缓解方法,用于评估和改进 AI 生成论文的质量。
👍 38
10/01 08:00
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrate
中文介绍 现有多教师 on-policy 蒸馏仅在输出分布层面迁移专家知识,Latent-MOPD 提出首个面向 LLM 的表征级多教师 on-policy 蒸馏方法,将多个专家模型的隐层表征融合进学生训练,以整合各自专长。
👍 37
10/01 08:00
Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation task
中文介绍 现有 Simulink 基准只看模型能否编译、运行或与参考模型相似,无法判断是否满足工程需求。SimuVerity 构建 101 个文本到可执行 Simulink 模型的生成任务,按工程需求而非表面相似度评估智能体的建模能力。
👍 36
10/01 08:00
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields
中文介绍 针对长程任务中 LLM agent 输出难以验证、测试时又无参考答案与评分标准的问题,VeriHarness 研究如何用固定基座模型强化验证能力,通过重复采样与 agent 化验证流程提升验证的规模与可靠性。
👍 34
10/01 08:00
Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, man
中文介绍 针对长视频生成与世界模型中记忆序列不断增长、结构日益复杂而难以管理的问题,该工作提出 Spatial Memory Intelligence,为世界模型赋予理解驱动的长期空间记忆,使其在用户动作条件下生成更一致的未来观测。
👍 33
10/02 08:00
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluat
中文介绍 论文构建多语言 GSM-Symbolic 数据集,系统考察某一语言中获得的能力如何迁移到其他语言,并联合分析影响迁移的决定因素,避免对每种语言逐一穷举评测,从而揭示跨语言能力迁移的关键变量。
👍 30
10/01 08:00
Scientific progress emerges from a longitudinal ecosystem in which researchers, institutions, funding agencies, collaboration networks, and the scientific literature co-evolve. As AI becomes increasingly involved throughout the scientific research cycle, understanding these interconnected and evolvi
中文介绍 将科研视为研究者、机构、资助方、合作网络与文献共同演化的长期生态系统,论文提出基于 LLM 的闭环模拟框架,用于研究 AI 深度参与科研流程后,各要素相互作用对科学进展的影响。
👍 20
10/02 08:00
Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate s
中文介绍 自回归视频模型依赖下一片段预测,只能短期、反应式生成,难以通过有效的中间步骤达成目标。ProAR 提出前瞻式推理方法,让 AR 视频模型能够规划并生成通向目标结果的中间过程,提升推理导向的生成能力。
👍 18
10/01 08:00
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this
中文介绍 论文将大预训练模型的灾难性遗忘视为每个权重矩阵输入空间中的几何问题,由此导出一个自然的保留目标,并指出梯度优化器产生的更新在该目标下并非最优,据此提出 Local Support Learning 缓解遗忘。
👍 17
09/29 08:00
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed eva
中文介绍 平均回报相近的策略在罕见失败上可能差异巨大,而准确估计下尾 CVaR 往往需要大量昂贵采样。论文提出 Tail-Influence Sampling,在可分别查询随机流程各条件分量时,将固定评估预算分配给对尾部风险影响最大的分量。
👍 17
10/02 08:00
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen,
中文介绍 现有国际象棋引擎棋力超人类却从不解释着法,而语言模型能生成看似合理的解释但棋力偏弱。论文提出 Queen,使语言模型既下出较强着法又能说明理由,兼顾棋力与可解释性。
👍 17
10/02 08:00
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization shou
中文介绍 针对 AlphaEvolve 等 LLM 引导演化方法只优化固定迭代次数下的性能、忽视算力成本的问题,FrugalEvo 提出成本感知的程序演化框架,在圆堆积等优化任务中显式权衡性能增益与调用开销。