Self-Improvements in Modern Agentic Systems: A Survey

Schmidhuber 也在作者里 —— 他是自指学习(self-referential learning, 1987)和 Gödel Machine(2003)的发明人。这篇 survey 几乎是”现代 RSI 领域的纲领性文件”。

和 harness-engineering-self-improvement 的关系:两篇同一时期(差 10 天),同一主题 —— 但视角互补:这篇学术综述(97 页)vs Lilian 的工业博客。术语差异很重要:survey 用 scaffold,Lilian 用 harness,同一回事(survey 第 1 节明确说”recently often also referred to as an agent harness”)。

摘要(核心论点)

自改进自治 agent 从研究原型走向部署系统。核心目标是可控进化 —— 在最少甚至无人介入下从经验中自适应。

现代 self-improving agent 被建模为 adaptive systems,把 experience 转化为 accumulated capability gains。

形式化框架

Survey 提出一个统一的系统级框架:

现代 agent = FM(基础模型) ⊕ Scaffold(操作层)
         = θ + Σ
其中 θ = 模型参数(核心智能)
    Σ = {prompt, memory, tool, control logic}(操作层)

Self-improvement = self-induced update operator U 作用于 θ 或 Σ:

θ_{t+1} = U_θ(θ_{1:t}, S_t)   ← 改模型(慢、稳、长期)
Σ_{t+1} = U_Σ(Σ_t, S_t)       ← 改 scaffold(快、可逆)

S_t 是 agent 自己产生的学习信号(trajectory / critique / preference / artifact)。

两大路径

1️⃣ Foundation Model Improvement(改 θ)

慢但稳定、长期 consolidation。三类驱动信号:

子类说明例子
Intrinsic Generative DemonstrationsFM 自己生成训练样本(input-output pair)self-instruct, Magpie
Intrinsic Evaluative FeedbackFM 自己评估生成质量(scalar reward / preference / critique)RLAIF, Self-Rewarding LM
Extrinsic Exploratory Experience跟真实/仿真环境交互获得 trajectoryRL in games, robotics

2️⃣ Scaffolding Improvement(改 Σ)

快、可逆、改结构。四个子类:

子类优化对象方法
Promptprompt pScalar-Feedback / Qualitative-Feedback / Population-Based Evolution / Textual Gradient
Memorymemory object + structure + processingRAG, episodic / semantic / procedural memory
Tooltool set TDynamic Tool Routing / Iterative Refinement / Autonomous Tool Creation
Full Scaffolding整体 Σ端到端改 control logic

术语映射(很重要):完整对照表 → scaffold-vs-harness

  • Prompt 的 textual gradient optimization 跟 Lilian 说的 “self-improving harness” 直接对应
  • Memory = Lilian 的 “file system as persistent memory”
  • Tool = Lilian 的 “sub-agent and backend jobs” 的工程化版本
  • Full Scaffolding = Lilian 的 “evolutionary search over the entire harness”

历史脉络(1790s-2026)

Survey 给了一个五阶段时间轴:

  1. Foundational Concepts (1790s-1960s):Babbage → Turing → Good 1965 提出 “ultraintelligent machine”
  2. Symbolism & Heuristic Self-Modification (1960s-1980s):Gödel Machine 的前身
  3. Connectionism & Meta-Learning (1980s-2000s):Schmidhuber 1987 自指学习、1992 fast weights、1993 self-modifying NN、2003 Gödel Machine
  4. Formal & Architecture-Level (2000s-2020s):AutoML, Neural Architecture Search
  5. Scalable FMs & Agentic Systems (2020s-Now):ChatGPT / Claude / GPT-4 / coding agents 让 self-improvement 变成工程现实

为什么 2026 年才爆发:FMs 把”自修改的搜索空间”从汇编码/raw weights 收缩到自然语言,搜索效率质变。

应用领域(Chapter 7)

领域代表
Software EngineeringClaude Code / Codex 这类 coding agents
Web Navigationbrowser-use 类 agent
Games & Strategic ReasoningAlphaGo / AlphaProof 类
Scientific DiscoveryAI Scientist / Sakana AI Scientist
Embodied AI & RoboticsRT-2 / π0 类
General Computer ControlComputer-Use 类

Future Directions(六个方向,Chapter 9.2)

Theme A: Algorithmic Paradigms for Lifelong Adaptation

  1. Test-Time Continual Adaptation — 训练/部署二分法的终结,让 model 在 deployment 中动态改 retrieval/routing/memory。关键风险:局部更新悄悄腐蚀全局性能。
  2. Active Exploration & Curiosity — agent 主动寻找有价值的经验(稀疏反馈下尤其重要)。Schmidhuber 1991 提出的”compression progress”作为 intrinsic reward 再次被引用。
  3. Parametric Distillation & Joint Optimization — 把 scaffold 的 System-2 能力”蒸馏”进 model 的 System-1 weights。同时 θ 和 Σ 联合优化(credit assignment 难题:失败时该改 prompt?改 tool?改 weights?)

Theme B: Complexity, Constraints, Open-World Robustness

  1. Resource-Constrained Improvement Dynamics — 自我改进不能烧光 token / compute 预算;评估从”峰值性能”转向”提升效率”。
  2. Multi-Agent Cooperative Co-Evolution — 多个 specialist agent 共享 regression tests / 成功的 patches / 改进的 tool wrappers。类 GitHub for agents。
  3. Surviving Open-World Distribution Drift — 抛弃 static leaderboard,部署在持续 drift 的环境里强迫 self-improvement 抵抗 catastrophic forgetting。

安全论点(Chapter 9 之前)

Survey 在 Section 9 之前提出一个重要的安全论点:

“a self-improving agent should be conceptualized as untrusted code executing in a protected runtime environment”

Security 不能只靠 FM 的初始 alignment。系统需要 layered gating + 严格 self-modification 权限。每次 Σ 或 θ 更新前必须过 verifier 检查:功能正确性、tool 权限边界、对随机扰动的鲁棒性。

改进必须显式限定在持续审计的安全边界内。

—— 跟 harness-engineering-self-improvement 里 Lilian 的”sub-agent + permission control”思路直接呼应。

Conclusion(一句话总结)

我们应该从”每次交互后重置的 stateless AI tools”过渡到”能持续自我改进的系统”。这需要 robust feedback、安全的 self-modification 架构、把 evaluation 重塑为持续的整合过程(而非静态 benchmark)。

我自己的 takeaway(3 件事)

  1. 术语统一:scaffold = harness = agent system 的”非参数部分”。我之后写笔记应该统一用 scaffold(survey 用得多),但 cross-link 到 Lilian 的 harness 概念。

  2. RSI 的两大轴:θ 轴(FM 改进,慢、稳) vs Σ 轴(scaffold 改进,快、可逆)。工业界现在主战场是 Σ 轴(Claude Code / Codex 都是),θ 轴需要 RL 基础设施所以慢。这是为什么 Lilian 那篇集中讲 harness。

  3. Schmidhuber 的影子无处不在:从 1987 自指学习到 2026 这个 survey,作者列表里有他 = 这条线从 40 年前就一脉相承。这不是新概念,只是终于有了工程基础。

待办

  • 读 Chapter 5.1(Intrinsic Generative Demonstrations)详细,对照 self-instruct / Magpie
  • 读 Chapter 6.1.4(Textual Gradient Optimization)—— 这跟 Lilian 的 “self-improving harness” 直接对应
  • 写一份跟 harness-engineering-self-improvement 的对比表(术语映射 + 视角差异)
  • 跟进 Schmidhuber 在 2026 年的其他工作
  • 跟踪 Computer-Use 类 agent 的 self-improvement 论文(应用领域里最热的)

引用格式(BibTeX)

@article{ren2026selfimprovements,
  title={Self-Improvements in Modern Agentic Systems: A Survey},
  author={Ren, Zhe and Chen, Yimeng and Guo, Dandan and Rong, Guowei and Li, Tonghui and Xiong, R.B. and Lan, Qingfeng and Wang, Wenyi and Li, Nanbo and Yang, Yibo and Zhuge, Mingchen and Schmidhuber, J{\"u}rgen},
  journal={arXiv preprint arXiv:2607.13104},
  year={2026}
}

元信息

  • 总页数:97
  • 参考文献:数百条(References 章节占了 30+ 页)
  • GitHub 仓库:Self-Improving-Agents —— survey 提到”for convenience, we track technical updates on this GitHub page”