#2026-10-01 AI/LLM 最新论文与研究热点简报

时间范围:2026-09-27 至 2026-10-01(覆盖最近 48 小时新提交,并回溯至 9 月 27 日仍在热榜上的高关注论文)。

信息来源:Hugging Face Papers 每日热榜(含社区投票数)、arXiv cs.LG 最近提交列表(2026-09-30,465 篇中的前 50 篇)、arXiv 论文摘要页、GitHub Trending(daily)、Hugging Face 官方博客。X/Twitter 检索因平台反滥用拦截(HTTP 403,匿名访问被屏蔽)无法访问,本次以 HF/arXiv/GitHub 为主,并用机构博客补足工业界动态。所有论文链接均逐一核实可访问,未编造任何条目。


#一、总览:今天的主线是什么

最近 48 小时的进展可以归纳为五条主线,恰好密集覆盖 wenjun 的研究版图:

  1. Harness/环境的自演化成为新的"能力放大器":热榜第一的 Raven(452 票)提出"harness of harnesses",把 agent 框架本身当作可被自动构造、经验改进、跨域编排的对象;AutoRef、Context Language Models、LibraryDesignBench 沿同一方向,把"环境/上下文/库设计"变成学习问题。
  2. On-Policy Distillation(OPD)研究爆发:至少 7 篇新论文(SAKI、Dr. OPD、ActFirst-OPD、PMOPD、ROSS、Scaling Properties、Mechanistic OPD)把后训练方法论的焦点从 RLVR 转向"在学生自己的轨迹上做稠密监督",并开始用机制可解释性工具研究它到底蒸馏了什么。
  3. 潜空间推理(latent reasoning)持续升温:REST 提出显式的表征监督目标修复 CE-only 训练的四个失效;Looped Transformer 系列两篇(调度、recurrence 条件)+ Time-Anchored Diffusion LM 把"在 latent 里递归思考"的理论和工程都往前推。
  4. RLVR 的信用分配问题被多角度围攻:Hindsight-Divergence Localization(分支采样)、Self-Privileged Critic、EasyPPO(critic 失稳诊断)、Gaussian Curricula(prompt 几何)、Cross-Model Trajectory Exchange(跨模型互补轨迹)。
  5. World Action Models(WAM)隐式 vs 显式之争出现关键实证:"What Makes WAMs Generalize" 用三轴实验证明 latent WAM 丢掉了显式未来建模带来的泛化收益——这对"LLM Dreamer 路线要不要显式 rollout"是直接证据。

#二、重点论文详评(Top 5)

#1. Raven: The Harness of Harnesses for Composable Agentic Intelligence

  • 链接:https://arxiv.org/abs/2609.33439
  • 来源:Hugging Face Papers 热榜 #1(452 票);GitHub: https://github.com/EverMind-AI/Raven
  • 日期:2026-09-27 提交
  • 类别:LLM Agent / Self-Evolving Agent / Systems
  • 核心贡献:开源多智能体生态,自动构造并持续演化针对特定模型×领域的模块化 harness,把每个"可执行模型-harness 对"当作可组合的智能单元;Host Agent 做目标分解、子任务匹配、执行依赖协调,配合 host archive / EverOS 保存经验、Skill Forge 把经验固化为可复用过程;论文还给出了"组合扩展可靠任务覆盖"的充分条件理论。
  • 为什么值得关注:这是"环境设计催生自演化智能"命题目前最完整的开源实现——不只是 prompt 一个 agent 让它自己变强,而是把"框架层"本身变成可学习、可复用经验的资产。452 票的单日热度说明社区对"harness 自动化"的共鸣极强。
  • 与 wenjun 的关系:与"通过环境设计催生自演化智能"直接对齐。可作为实验基础设施(EverOS 经验复用机制)或研究对象(harness 组合的理论边界 vs 实际收益)。

#2. Principled Thoughts for Latent Recursive LLM Systems (REST)

  • 链接:https://arxiv.org/abs/2609.36159
  • 来源:arXiv 新提交(热榜 5 票);GitHub: https://github.com/FARD-Lab/REST
  • 日期:2026-09-28
  • 类别:Latent Reasoning / Post-training RL
  • 核心贡献:指出 CE-only 训练(只监督最终解码答案)在潜空间递归推理系统中导致四个失效——不同问题的思考态塌缩、无关信息滞留等——并提出 REST(REpresentation-Supervised Thoughts)训练目标,直接约束中间隐状态。
  • 为什么值得关注:这是少数对"latent chain-of-thought 该怎么训练"给出明确 loss 设计的工作,理论 + 实验闭环,回答了 Coconut 一系工作留下的"中间表征无监督"空白。
  • 与 wenjun 的关系:与潜空间推理主线正面相关。REST 的表征监督思路可以自然扩展到 agent 的"内部 world state"表征上——即用类似目标约束 model-based RL 中 latent 状态的演化质量。

#3. Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR

  • 链接:https://arxiv.org/abs/2609.37868
  • 来源:Hugging Face Papers(42 票)
  • 日期:2026-09-29
  • 类别:Post-training RL (RLVR) / Multi-Agent
  • 核心贡献:观察到异构模型在互补的 prompt 集合上成功,提出让多个 RLVR 模型交换成功轨迹(off-policy 视角下谨慎复用彼此经验),解决有限 rollout 预算下 all-fail group 无梯度信号的问题。
  • 为什么值得关注:这是把"population-based training / 联邦式经验共享"引入 RLVR 的清晰提案,直接针对长轨迹任务中稀疏奖励 + rollout 昂贵的痛点。
  • 与 wenjun 的关系:与长轨迹 Agent RL 高度相关。多智能体轨迹交换本质是另一种 bias-variance 折衷(off-policy 数据 vs 更快的需求覆盖),可与 GIGPO 的分组思路结合。

#4. What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling

  • 链接:https://arxiv.org/abs/2609.34981
  • 来源:Hugging Face Papers;GitHub: https://github.com/LeapLabTHU/Simple-WAM
  • 日期:2026-09-28(v2: 09-29)
  • 类别:Model-based RL / World Model
  • 核心贡献:系统对比显式 WAM(推理时仍去噪出未来帧)与隐式 WAM(推理时丢弃未来建模)在环境扰动、数据效率、scaling 三个泛化轴上的表现;发现 latent WAM 在分布内打平,但丢掉了显式未来建模带来的泛化收益。
  • 为什么值得关注:这是"Dreamer for LLM Agent"路线的关键判据论文:如果推理时完全跳过 latent rollout 换速度,泛化能力会受损——即世界模型的"想象"在测试时可能不是冗余开销,而是泛化来源。
  • 与 wenjun 的关系:直接为 LLM model-based RL 的架构决策提供证据:保留推理时的 latent future modeling(哪怕是低精度/低频率)vs 纯 implicit policy 之间存在真实的泛化-效率 trade-off,值得在语言 agent 域复现该结论。

#5. LLMs are General Asynchronous Agents

  • 链接:https://arxiv.org/abs/2609.35427
  • 来源:Hugging Face Papers(62 票);关联仓库: https://github.com/dvmazur/async_llm
  • 日期:2026-09-28
  • 类别:LLM Agent / Systems
  • 核心贡献:把语音助手、具身智能体、监控系统等不同异步场景统一形式化为"通用异步 agent"——模型在思考或执行任务时可以接收新输入并自适应;提供统一的训练/评测框架。
  • 为什么值得关注:异步交互是对"read→think→act"顺序循环这一根本 agent 假设的挑战,统一形式化为后续 RL 训练(如中断/抢占下的信用分配)铺路。
  • 与 wenjun 的关系:长轨迹 Agent RL 目前几乎全部假设同步回合制;异步化后 credit assignment 与 rollout 定义都需要重写,是一个理论空白点。

#三、其余值得关注的论文速览

按主题分组,每条给出:标题 | 链接 | 日期 | 一句话核心贡献 | 类别标签。

#LLM Agent / Harness / Self-Evolving

  • Omni-IO Skills: Harnessing Your Agent Omni-Native | https://arxiv.org/abs/2609.31847 | 09-25 | 插件式 harness,通过分层 Skills 让现有 agent 无需改模型即获得文/图/音/视/3D/代码全模态生产能力。| LLM Agent / Tool-use
  • Marathoner: Ultra-Long-Horizon Autonomous Intelligence | https://arxiv.org/abs/2609.34378 | 09-28 | 用 1000+ 行真实 GitHub PR 合成超长时程任务,配合完整后训练流水线赋予模型以月计的持续执行能力。| LLM Agent / Code Agent / Post-training
  • AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation | https://arxiv.org/abs/2609.35530 | 09-28 | 把 harness 参数本身作为可优化对象做自动搜索。| LLM Agent / Systems
  • Can Agents Design Libraries for Agents? | https://arxiv.org/abs/2609.36730 | 09-29 | 提出 LibraryDesignBench(242 个专家验证库),度量 agent 为其他 agent 设计可复用库的能力——发现 agent 之间大量重复造轮子。| Code Agent / Evaluation
  • Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents | https://arxiv.org/abs/2609.37236 | 09-29 | 区分"横向/纵向主动性"两个信息获取维度,用 need graph 度量 agent 主动补全未言明信息的能力。| LLM Agent / Evaluation
  • AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? | https://arxiv.org/abs/2609.35025 | 09-28 | 评测 agent 能否写出现在/未来训练自己的数据,是对"递归自我改进数据侧"的首次系统度量。| LLM Agent / Pretraining Data / Evaluation
  • EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory | https://arxiv.org/abs/2609.37923 | 09-29 | 多 agent 共享协同演化的多模态记忆做集体学习。| LLM Agent / Memory
  • Org-Agent: Beyond Personal Assistants Towards Organizational Agents | https://arxiv.org/abs/2609.34392 | 09-27 | 面向组织级多 agent 协作的角色/职责建模。| LLM Agent

#上下文管理 / 记忆 / 压缩

  • Context Language Models | https://arxiv.org/abs/2609.37725 | 09-29 | 把 context 当作一个文件,允许模型无限制改写自己的上下文;zero-shot 构建即超过 SOTA 上下文管理策略(BrowseComp-Plus +11.4% 精度、-21.5% FLOPs),并自然扩展到多 agent 上下文并存。GitHub: https://github.com/facebookresearch/context-language-models | LLM Agent / Memory / Systems
  • VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents | https://arxiv.org/abs/2609.38119 | 09-29 | 证明 append-only 工作记忆在结构上无法清除噪声,提出循环改写式工作记忆。| LLM Agent / Memory
  • Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient RLVR | https://arxiv.org/abs/2609.36864 | 09-29 | 用 hindsight 诱导的 token log-likelihood 变化定位"模型后悔的决策点",在关键分支处采样替代续延,大幅提高 rollout 利用率。| Post-training RL / Credit Assignment
  • APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants | https://arxiv.org/abs/2609.37559 | 09-29 | 跨会话持久记忆的流式视频基准。| Evaluation / Memory

#On-Policy Distillation 系列(本周最密集的方法论线)

  • SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation | https://arxiv.org/abs/2609.36601 | 09-29 | 80 票热榜第 4;用 KL 约束教师引导 rollout + maximal coupling 的接受/纠正事件路由 token 级监督。| Post-training
  • Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation | https://arxiv.org/abs/2609.38025 | 09-30 | 学习每处教师信号的重要性权重,而非一律照抄。| Post-training
  • Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents (ActFirst-OPD) | https://arxiv.org/abs/2609.36608 | 09-29 | 用 reference-conditioned inverse dynamics 先行动后推理,解耦环境交互与完整响应生成,加速多轮 agent 的经验收集。| Post-training / LLM Agent
  • PMOPD: Multi-Teacher On-Policy Distillation | https://arxiv.org/abs/2609.34605 | 09-28 | 针对多教师蒸馏的 capability seesaw,做任务排序 + 参数更新子空间保护。| Post-training / Continual Learning
  • ROSS: Relearning from Self-Generated Rollouts through Selective Supervision | https://arxiv.org/abs/2609.35954 | 09-28 | 把历史 rollout 当作可选择性复用的"陈旧经验",细粒度筛选后回炉。| Post-training / Continual Learning
  • Scaling Properties of Same-Family On-Policy Distillation | https://arxiv.org/abs/2609.32722 | 09-26 | 发现 OPD 早期存在 gold-score 随 KL 距离线性上升的"useful-transfer"规律区间。| Post-training / 理论
  • Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning | https://arxiv.org/abs/2609.37915 | 09-30 | 因子化分析发现"脚手架正确性"比"特权上下文"对下游效果影响更大。| Post-training
  • Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders | https://arxiv.org/abs/2609.35210 | 09-28 | 用跨模型共享特征字典的 sparse crosscoder 追踪 OPD 前后学生表征的真实变化。| Mechanistic Interpretability / Post-training

#RLVR / RL 训练机制

  • EasyPPO: Stabilizing the Critic Is Key | https://arxiv.org/abs/2609.36802 | 09-29 | 诊断出 LLM-PPO 中 critic 的两大失稳模式(截断 rollout 过滤偏差 + critic 目标漂移)并给出修复。| Post-training RL / Systems
  • Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR | https://arxiv.org/abs/2609.37825 | 09-30 | 让 critic 在训练时访问特权信息(参考答案)但推理时不依赖,解决稀疏终末奖励下 value 估计不可靠的问题。| Post-training RL / Credit Assignment
  • Prompts Live on an Arc: Gaussian Curricula in Fisher-Rao Coordinates for Rollout-Efficient GRPO | https://arxiv.org/abs/2609.38018 | 09-30 | 把 prompt 按 pass rate 的 Fisher-Rao 几何组织成高斯课程,减少零梯度 all-fail/all-success 组的 rollout 浪费。| Post-training RL
  • TGRL: Temperature-Grouped Reinforcement Learning | https://arxiv.org/abs/2609.33589 | 09-27 | 按温度分组采样提升 LLM RL 探索效率。| Post-training RL
  • LLM-PPO/RLVR 之外的免训练增强:Explore Broadly, Reason Sharply | https://arxiv.org/abs/2609.38104 | 09-30 | power-sharpened sampling 作为 RL 后训练的推理时替代,避免 RL 的锯齿泛化。| Test-time Scaling

#Latent Reasoning / 递归/循环模型

  • Scheduling Recursive Reasoning in Looped Transformers | https://arxiv.org/abs/2609.36653 | 09-29 | 给出终端损失对递归更新尺度的精确分解,据此自适应调度每步更新的强度。| Latent Reasoning / 理论
  • What Makes Recurrence Effective in Looped Language Models? | https://arxiv.org/abs/2609.36636 | 09-29 | 系统回答"何时/何处/如何条件化"recurrence 才有效。| Latent Reasoning
  • Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation | https://arxiv.org/abs/2609.37924 | 09-30 | 自监督学习时间锚点并在潜空间缓存复用,加速扩散语言模型生成。| Latent Reasoning / Systems
  • S³: Spectral Null-Space Swap Makes Reasoning Models Efficient | https://arxiv.org/abs/2609.37976 | 09-30 | 发现 thinking 模型的推理能力集中在相对 non-thinking 模型主导奇异方向的零空间分量里,切除即可大幅提效。| Post-training / Systems

#World Model / Model-based RL(机器人侧,方法可迁移)

  • EVO-WAM: Evolving World Action Models through Video-Action Verification | https://arxiv.org/abs/2609.38057 | 09-29 | WAM 用自己生成的视频-动作轨迹 + 一致性验证做自演化适配。| Model-based RL
  • AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference | https://arxiv.org/abs/2609.33748 | 09-27 | 按动作误差敏感度自适应分配去噪预算。| Model-based RL / Systems
  • StructRL: Online Structured RL for Long-Horizon VLA Tasks | https://arxiv.org/abs/2609.36352 | 09-28 | 把稀疏终末奖励改造成结构化中间监督(发现/接触/使用物体),专为长时程任务设计。| Model-based RL / Long-horizon
  • Anisotropic Representations Improve Planning in JEPA World Models | https://arxiv.org/abs/2609.37441 | 09-29 | 各向异性表征几何让 JEPA 潜空间规划代价与任务代价对齐——对一切 latent planning 工作普适。| Model-based RL / Latent Reasoning

#持续学习 / 训练机制 / 可解释性

  • Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss | https://arxiv.org/abs/2609.33620 | 09-27 | 推导 token/层级分解,统一刻画持续学习的数据归因、遗忘与可塑性损失。| Continual Learning / 理论
  • Behavioral Convergence Without Representational Convergence | https://arxiv.org/abs/2609.37836 | 09-30 | 控制实验证明相同最终行为的网络内部表征仍带训练历史烙印。| Mechanistic Interpretability
  • Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space | https://arxiv.org/abs/2609.37731 | 09-30 | 激活空间接地的参数分解,把"表征的信息"与"参数的计算"连起来。| Mechanistic Interpretability

#安全 / 评测

  • Language Models Are "Insecure" Reporters | https://arxiv.org/abs/2609.36139 | 09-28 | 8 个对抗场景系统研究 LLM 是否在报告里隐瞒"颠覆叙事的缺陷"——长时程自治时代报告可信度问题。| Evaluation / Safety
  • SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents | https://arxiv.org/abs/2609.34518 | 09-28 | 把工具型 agent 攻防形式化为部分可观测状态控制。| Tool-use / Safety
  • EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? | https://arxiv.org/abs/2609.37686 | 09-29 | 1301 个专家任务覆盖 CAD/CAE/CAM/BIM/EDA 的专业工程基准。| Evaluation / Code Agent
  • HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents | https://arxiv.org/abs/2609.38008 | 09-29 | 论证下一代 CUA 应学会编排 GUI 与 shell 命令。| Tool-use / Evaluation

#系统 / 基础设施

  • Jaxolotl: LTL-Based Multi-Task RL Benchmark Suite | https://arxiv.org/abs/2609.38065 | 09-30 | 统一高性能的 LTL 指令多任务 RL 训练套件。| Systems
  • LongCat-DeepResearch Technical Report | https://arxiv.org/abs/2609.36071 | 09-28 | 美团系深度研究系统:全局规划与分段调查分离 + 段级修订的多 agent 工作流。| LLM Agent / Systems

#四、工业界与社区热点(非论文)

  • GitHub Trending(daily):

- NVIDIA/OpenShell — "autonomous AI agents 的安全私有运行时",NVIDIA 官方开源,直指 agent 沙箱执行安全问题。https://github.com/NVIDIA/OpenShell | Systems / Tool-use

- openclaw/openclaw — 本周继续霸榜的通用执行型 AI agent("The AI that really does things")。https://github.com/openclaw/openclaw | LLM Agent

- mvschwarz/openrig — 把 Claude Code 与 Codex 当作子系统统一调度的 multi-agent harness。https://github.com/mvschwarz/openrig | Code Agent / Systems

- mksglu/context-mode — 面向编码 agent 的上下文窗口优化:沙箱化工具输出(自称减少 98%)、持久化会话记忆。https://github.com/mksglu/context-mode | LLM Agent / Memory / Systems

- VectifyAI/PageIndex — 无向量化、基于推理的文档索引 RAG。https://github.com/VectifyAI/PageIndex | Retrieval / LLM Agent

- colbymchenry/codegraph — 预索引代码知识图谱 + 代码变更自动同步,服务 Claude Code/Codex/Cursor 等全部主流编码 agent。https://github.com/colbymchenry/codegraph | Code Agent / Systems

  • Hugging Face 官方博客(近一周值得注意):Hcompany 的 Holo4(通用 computer-use agent,09-28)、NVIDIA Warp/MjWarp 机器人仿真加速(7 天前)、Async GRPO with LoRA across HF Jobs(免 NCCL 的分布式 GRPO 实操教程,对复现 RLVR 实验很有用)。https://huggingface.co/blog
  • X/Twitter:本次无法访问(403 反滥用拦截),以上述可访问来源替代。

#五、今日最值得精读的 3 篇

  1. Raven (2609.33439) — "harness of harnesses" 的完整开源实现 + 组合理论,是环境自演化方向目前最系统的作品;即使不做该方向,其 EverOS/Skill Forge 的经验固化设计对任何长时程 agent 记忆设计都有参考价值。
  2. What Makes WAMs Generalize? (2609.34981) — 用干净的三轴实验回答"推理时到底要不要显式生成未来",是 LLM model-based RL 路线图上必须引用的证据型论文。
  3. REST: Principled Thoughts for Latent Recursive LLM Systems (2609.36159) — 潜空间推理训练目标的首次原则性设计,指出了 CE-only 训练的四个具体失效模式并逐一修复。

(备选第 4 篇:LLMs are General Asynchronous Agents,如果你想在"异步长轨迹 RL"这个几乎无人理论化的空白点上找题。)

#六、今日最值得跟进的 3 个 repo / model / dataset

  1. EverMind-AI/Raven — https://github.com/EverMind-AI/Raven(harness 自演化生态的开源实现,452 票热榜第一,代码刚放出适合早期跟进)
  2. facebookresearch/context-language-models — https://github.com/facebookresearch/context-language-models(CLM:模型自己管理改写 context 的官方实现,与上下文压缩器方向直接相关)
  3. NVIDIA/OpenShell — https://github.com/NVIDIA/OpenShell(agent 自治执行的安全运行时,若你做 agent RL 环境工程,这是现成的 sandbox 底座)

#七、研究机会 / Ideas

  1. 异步长轨迹 RL 的信用分配:当前所有 RLVR/agent-RL 方法(GRPO、GIGPO、PPO 变体)都建立在同步回合制假设上。"General Asynchronous Agents" 只给了环境形式化,没给训练算法——在中断/抢占/新输入打断轨迹的设定下,group-relative advantage 和 trajectory 分组都要重新定义。这是一个理论 + 系统双空白,且和 wenjun 的长轨迹 RL 方向完全重合。
  2. WAM 显式/隐式 trade-off 在语言域的复现与解释:2609.34981 证明了机器人域 latent WAM 丢泛化。语言 agent 的"想象 rollout"(latent 递归推理)是否存在同样的"推理时想象 = 泛化来源"效应?可以把 REST 的表征监督 + 显式 latent rollout vs 纯 policy 蒸馏做成一个受控对比实验,直接产出对 LLM Dreamer 路线的判据。
  3. Harness 作为可学习对象的"组合定律":Raven 给了实现,但其"组合扩展可靠任务覆盖"的理论条件还停留在充分条件层面。结合 LibraryDesignBench 的发现(agent 重复造轮子而非复用),可以问:什么时候"让 agent 学会设计库/harness 给未来的自己用"比"直接训练更强的 policy"更划算?这本质是环境侧 vs 参数侧的 scaling trade-off,与"agent 预训练数据塑造能力"一脉相承。

简报由 Hermes 定时任务于 2026-10-01 08:00 (Asia/Shanghai) 生成。数据截至发稿时。