# 外部材料学习队列 [说明与维护](./README.md) · [其余候选](./CANDIDATES.md) 排名是建议阅读顺序,不代表已读或同时开工承诺。按材料 ID 保留重复来源;当前队列不含历史归档。 ## Top 30 | 顺序 | 材料 | 档位 | 读取范围与备注 | | --- | --- | --- | --- | | 1 | [Manus Context Engineering:缓存稳定性、外置记忆、目标复述与失败证据]() | S | 沿用既有阅读记录;比较缓存稳定前缀、文件外置记忆、目标复述与失败证据各自解决什么问题。 | | 2 | [How we built Grok Bot in a month:产品迭代、onboarding 与持久 Agent]() | S | 沿用既有阅读记录;关注原型到上线的产品取舍、手动 onboarding 与功能减法;区分采用证据和发布叙事。 | | 3 | [Agents Should Be Durable, Not Long-Lived:durable state + leased executor 的 agent 运行时抽象(Ju Lin)]() | S | 沿用既有阅读记录;画出 durable state 与 leased executor 的边界,解释 worker 重启和任务恢复的区别。 | | 4 | [Co-ReAct:把 Rubric 从结果评分器前移为 Step-Level Action Spec]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() · [来源 7]() · [来源 8]() | S | 沿用既有阅读记录;追踪 rubric 如何进入逐步行动选择,区分训练奖励、推理约束与结果评分。 | | 5 | [长程智能体自我检查与早停行为:Reward Seeking 现象及其缓解措施]() · [来源 2]() | S | 本次全文审查;已读青稞公开转载全文;知乎原页未读,图片未解释。区分 reward seeking 与 reward hacking,关注 grader 猜测、浅层自检导致早停和交付失真。已核对部分相关一手文本;LangChain 的 52.8→66.5 是整体 harness 收益,不能归因于单一中间件。建议读 §1、§3、§5,结合外部完成契约、独立验证及预算/无进展停止条件。 | | 6 | [Anthropic Frontier Red Team:多 Agent 行为失败模式与协调机制]() | S | 沿用既有阅读记录;按共享目标、信息隔离与激励结构归类多 Agent 失败,思考哪些需要协议而非提示词。 | | 7 | [pstack 作者原文 Pt.1:验证 CLI、Feature Map 与云 Agent]() | S | 沿用既有阅读记录;选一个工程动作设计验证 CLI,列出输入、可观察结果和失败判定。 | | 8 | [Designing Grok Bot:为 persistent agents 设计五个产品原语(Bot roster/avatar presence/Bot 自有 computer/capability-context 分界/Routine 触发)]() | S | 沿用既有阅读记录;比较 Bot 身份、存在感、自有计算机、能力与上下文、Routine 触发五个产品原语。 | | 9 | [Open Muse:跨设备个人 Agent 的 MA 客户端、未确认写入恢复与权限边界]() · [来源 2]() | S | 沿用既有阅读记录;已读固定 commit 的 README、关键设计文档和部分源码,未运行测试或应用。关注 MA 云端循环与设备操作分界、未确认会话恢复、记忆写后读回、Agent 版本和跨设备 claim。优先读 conversations/identity 的恢复路径及 verification 未覆盖项。跨设备编辑没有原子 CAS;历史验证记录不等于当前版本实测。 | | 10 | [Google Antigravity Teamwork:多 agent 研究编排框架,patterns 解耦 + 自适应 agent 数量,Long Proof 竞争式策略搜索解决开放问题]() | S | 沿用既有阅读记录;读协作 patterns 与动态 Agent 数量;区分竞争式搜索、并行分工和结果集成。 | | 11 | [Server-side Agent 的 Harness tradeoff(高策):sandbox 单用户 → 多租户服务端 / sandbox=无状态执行环境 / session 与 context 工程 / eval parity 与轻量 JS runtime 探索]() | S | 沿用既有阅读记录;比较本地单用户与服务端多租户的 sandbox、session、资源隔离与 eval 约束。 | | 12 | [DeepSeek DSec:Agent 训练沙箱的平台约束与 RL 协同(jerrychaox 线索)]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录;已选读 DSec v1 摘要、引言及 §2、§4–6、§8;X 仅截断预览,配图与后续未读。关注环境分层、按需加载、超卖 QoS、trainer/worker/sandbox 状态归属、结果复用和 pause/resume。性能与规模是作者报告,未复现;§8 不验证 RL 集成。优先读 §4、§6。 | | 13 | [异步续延与 Effect 架构的第一性原理:要虚拟化时间/确定性调度/中断就必须持有续延,用户态解释器(Effect)vs 运行时下沉(PocketJS)]() | S | 沿用既有阅读记录;用续延解释中断、虚拟时间与确定性调度;区分用户态解释器和运行时支持。 | | 14 | [Mastra Observational Memory:Observer/Reflector 后台把原始消息压缩成有界 observation log,prompt-cache 友好并保留 retrieval 还原]() | S | 沿用既有阅读记录;画出 Observer/Reflector 与检索回查链路,观察压缩对 token 成本和信息保真的影响。 | | 15 | [Tencent WorkBuddy Bench:真实分布驱动的多域 Agent Benchmark 与 Harness 敏感性]() · [来源 2]() | S | 沿用既有阅读记录;检查真实任务分布、双 harness 对照、评分边界和预算记录。 | | 16 | [WorkBuddy:从非目标用户的越界使用发现真实需求,并封装成高价值结果工作流]() | A | 沿用既有阅读记录;关注越界使用怎样暴露需求,以及怎样把通用 Agent 包装为可验收的结果工作流。 | | 17 | [The AI-native SDLC playbook:6 阶段 + plays + intent.md 产物链,AI 嵌入每个环节但人保留在 gate 之上]() | S | 沿用既有阅读记录;对照六阶段产物链,说明模型自动化、人工 gate 与持续评估分别落在哪里。 | | 18 | [Knowledge-Centric Self-Improvement:让共享知识而非 Agent 实现成为持续改进对象]() · [来源 2]() | S | 沿用既有阅读记录;区分共享知识更新和 Agent 实现更新;追踪知识如何验证、复用及失效。 | | 19 | [SkillOpt-Lite 与 HarnessOpt:把 rollout 调试、skill 演化和执行框架演化统一成受控闭环]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录;比较 rollout 调试、skill 更新和 harness 修改三个优化对象及其验证边界。 | | 20 | [Pi AgentHarness v3:durable runtime 实现规范(官方 spec)]() | S | 沿用既有阅读记录;读 durable runtime 的状态、恢复和接口契约;规范声明需与具体实现区分。 | | 21 | [DeepSeek-V4.1-Flash:CED、CSA2、FP4 KV cache 与 bounded replay]() · [来源 2]() | S | 沿用既有阅读记录;拆解 prefill/decode 计算与 KV 存储账本;核对压缩收益成立的模型和负载条件。 | | 22 | [火山引擎张鑫:企业 Agent 落地,我们之前忽视了经营问题]() | A | 沿用既有阅读记录;把企业 Agent 采用拆为业务结果、运营责任、留存与成本,避免只看单次演示。 | | 23 | [跨太平洋的 AI 棋局:公开译文中的开源商业化与 RSI 观点]() | A | 沿用既有阅读记录;公开译文的产业观点;对照开源成本、云厂商价值捕获与 RSI 判断,事实需追一手证据。 | | 24 | [DeepSeek Harness(dsh):一切皆插件的 agent harness 全栈]() · [来源 2]() | S | 沿用既有阅读记录;追踪插件、Session Log、权限和工具执行的所有权边界。 | | 25 | [Argus:长期研究 Agent 的证据闭环与验证门禁]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录;关注四角色的证据交接、验证门禁与 Core/Vertical 分层,而不只看运行时长。 | | 26 | [Cordis:时空可组合性的编程范式(design paper + 框架)]() · [来源 2]() | S | 沿用既有阅读记录;用一个插件加载与卸载例子理解服务注入、资源生命周期和可回收副作用。 | | 27 | [Pi Coding Agent:minimal core + self-extensible agent harness]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() · [来源 7]() | S | 沿用既有阅读记录;从最小 agent loop 追到 session 与扩展点,区分核心责任和可替换策略。 | | 28 | [AI4AI at Test-Time:strong builder 用 5% validation 迭代构造 harness,把 weak target 精度从 0.49 提到 0.91]() | S | 沿用既有阅读记录;核对 builder/target 分工、验证集预算与基线;避免把 test-time 搜索收益混为模型增益。 | | 29 | [Fast Weight Attention for Continual Learning:把 fast-weight 记忆统一到在线持续学习框架(Falcon-1/2/3 家族)]() | A | 沿用既有阅读记录;从在线学习看 fast weights,区分参数记忆、上下文记忆和持续学习假设。 | | 30 | [Dream-RSI:通过已验证发现树回放,优化探索策略]() · [来源 2]() | S | 沿用既有阅读记录;手算发现树上的一次策略回放,检查候选预算、零执行评估与收益估计的条件。 | ## Ranked backlog | 顺序 | 材料 | 档位 | 读取范围与备注 | | --- | --- | --- | --- | | 31 | [Linear Sync Engine 逆向研究:事务、delta 顺序、离线队列与权限订阅]() | S | 沿用既有阅读记录;按本地事务、delta 顺序、离线重试和权限订阅追一条同步路径。 | | 32 | [Dream-RSI 第三方复现:独立实现、回放契约与证据边界]() | A | 沿用既有阅读记录;对照论文与第三方实现的回放和预算契约;独立复现不等于作者官方代码。 | | 33 | [数据库与分布式系统补课讲义:从 FP 与状态机到事务、幂等、fencing 与恢复](<./distributed-systems-for-loopx/补课讲义.md>) | S | 本次全文审查;17 节连接类型、纯决策、提交时验证、并发历史、日志与恢复、共识、存储与排队;用公开 commit 固定的 LoopX 代码作案例。 | | 34 | [数据库与分布式系统:故障推演、交叉讨论与四周学习路线](<./distributed-systems-for-loopx/练习与讨论.md>) | A | 本次全文审查;8 道故障题、10 组讨论与四周安排;区分会复述、会推演、能实现和能提出反例。 | | 35 | [Recovery Lab:幂等扣减、旧执行者、乱序投影与 write skew 教学实验](<./distributed-systems-for-loopx/recovery_lab.py>) | A | 本次运行通过;SQLite 演示盲重试余额 80、原子去重余额 90;覆盖参数冲突、提交前回滚、epoch 条件写和 snapshot 版本;write skew 是抽象内存模型。 | | 36 | [Temporal 架构:History、Matching、内部任务队列、Outbox 与持久化]() | A | 沿用既有阅读记录 | | 37 | [TeamAI:通过 Git 分发团队 skills、rules、MCP 与记忆]() | S | 沿用既有阅读记录 | | 38 | [Koishi 官方文档:可逆的插件系统(Cordis 的可回收副作用与资源安全)]() | B | 沿用既有阅读记录 | | 39 | [Apache Maka(Incubating):Log Is the Runtime——append-only RuntimeEvent Log 作为 agent 状态事实基座,UI/模型上下文/恢复都是投影]() · [来源 2]() | S | 沿用既有阅读记录 | | 40 | [π-agent book:pi-agent-core 源码逐行架构解读(在线连载)]() · [来源 2]() | S | 沿用既有阅读记录 | | 41 | [SoL-Pi:基于 Pi 的 harness 自动研究与机制筛选]() | S | 沿用既有阅读记录 | | 42 | [Rust Book Ch.6:Enums 与 Pattern Matching——sum type + match 穷尽检查 + if let 语法糖]() | S | 沿用既有阅读记录 | | 43 | [引用透明性:表达式替换、纯函数与程序优化的条件]() | S | 沿用既有阅读记录 | | 44 | [Huxley–Gödel Machine:Clade-Metaproductivity、Thompson Sampling 与自修改搜索]() · [来源 2]() | S | 沿用既有阅读记录 | | 45 | [Grok Bot for Engineering:工程 Agent 的工种、验收与复盘]() | S | 沿用既有阅读记录 | | 46 | [Crouzeix 猜想候选证明仓库:long-horizon 多 Agent 科研编排 prompt + 公开 trace + Lean 公理审计]() | S | 沿用既有阅读记录 | | 47 | [Cloudflare:Build your own vulnerability harness]() | A | 沿用既有阅读记录 | | 48 | [LongHorizonOS:时间片、上下文切换、重启门禁与版本化进展图]() | S | 沿用既有阅读记录 | | 49 | [Omarchy:Agent launcher、崩溃交接与跨 harness skills]() · [来源 2]() | S | 沿用既有阅读记录 | | 50 | [Scala 与 Algebraic data type:sealed trait + case class + match 穷尽性,TypeScript/Rust 之外的第三种 ADT 实现]() | S | 沿用既有阅读记录 | | 51 | [Kimi K2.5 / Agent Swarm]() | A | 沿用既有阅读记录 | | 52 | [Hermes Mixture-of-Agents:多模型协同与聚合]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() | A | 沿用既有阅读记录 | | 53 | [JIT-Agent:训练模型即时生成受协议约束的 harness]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录 | | 54 | [Agentic Harness Engineering:把 harness 自迭代从 craft 变成可观测工程系统]() · [来源 2]() | A | 沿用既有阅读记录 | | 55 | [Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 56 | [EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning]() | A | 沿用既有阅读记录 | | 57 | [EnvHarness:把 agent harness 思想镜像到环境侧,用 Setup/Rules/Link 插件层重塑冻结环境并保留原 verifier,LLM 设计器 EnvRigger 自动诊断并共进化]() · [来源 2]() | S | 沿用既有阅读记录 | | 58 | [EvoMap / Evolver:Gene、Capsule 与进化资产共享协议]() · [来源 2]() | S | 沿用既有阅读记录 | | 59 | [陈宇森 / MuleRun Agent Builder:runtime + skills + conversational marketplace]() | A | 沿用既有阅读记录 | | 60 | [Palantir 的“本体论骗局”:企业 AI 平台叙事解构与采购判断框架]() | A | 沿用既有阅读记录 | | 61 | [Running a Software Factory Efficiently at Uber Scale:软件工厂成本工程(四层 agent、六项成本方程、模型路由、Code-Mode、context graph、16 类 session 反模式)]() | S | 沿用既有阅读记录 | | 62 | [Dan Shipper / After Automation:异步委派 Agent、共享工作界面与必须有人照料的自动化]() | A | 沿用既有阅读记录 | | 63 | [DSH + AWiki:Agent 原生身份与外部消息/邮箱接入的实现样本]() | A | 沿用既有阅读记录 | | 64 | [Chat 不是终局:长程任务的状态卡片、监督与信任设计]() | S | 沿用既有阅读记录 | | 65 | [pstack 指南解读:验证基础设施、Agent CLI 与 Feature Map]() · [来源 2]() | S | 沿用既有阅读记录 | | 66 | [机器之心:DeepSeek Harness 震撼开源(二手报道 / 产品采用信号)]() | B | 沿用既有阅读记录 | | 67 | [Cordis 在做什么:从 DeepSeek Harness 看——服务/inject/effect/Loader 的机制级代码映射]() | A | 沿用既有阅读记录 | | 68 | [万字长文:DeepSeek Harness 一文全看懂——Profile/Bundle/Patch、Cordis、Session Log、安全管线、四种模式与动态造工具]() | A | 沿用既有阅读记录 | | 69 | [Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() | S | 沿用既有阅读记录 | | 70 | [OpenAI Codex Persistent mode:reasoning effort 新档让 Codex “继续工作直到被休眠”,并主动创建后续任务(WIRED 一手报道)]() | A | 沿用既有阅读记录 | | 71 | [Learn Claude Code:从 0 到 1 构建 nano Claude Code 风格 agent harness(20 个渐进 session,102→1708 LOC)]() | S | 沿用既有阅读记录 | | 72 | [Prime Agent:自我改进的 RLM coding/research harness(持久 IPython + Continual Harness)]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录 | | 73 | [Multica:human + agent teams 的开源 managed agents platform]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录 | | 74 | [FDE Observatory:Claude Managed Agents 深度调研 / Anthropic 的第三条路]() | A | 沿用既有阅读记录 | | 75 | [Paperclip:AI-agent company control plane]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() | A | 沿用既有阅读记录 | | 76 | [Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents]() | A | 沿用既有阅读记录 | | 77 | [Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses]() · [来源 2]() | A | 沿用既有阅读记录 | | 78 | [Cloudflare Workflows saga rollbacks:durable step compensation for long-running workflows]() | A | 沿用既有阅读记录 | | 79 | [repo-harness:把 Claude/Codex 工程会话变成 repo-local 的可恢复工作流]() | A | 沿用既有阅读记录 | | 80 | [Terminal-Bench:命令行环境中的 hard realistic tasks]() | A | 沿用既有阅读记录 | | 81 | [SkillsBench:benchmarking how well agent skills work across diverse tasks]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 82 | [Agents' Last Exam:economically valuable long-horizon agent benchmark]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 83 | [Orca:worktree-native 的多 Agent 开发工作台与早期 orchestration control plane]() | A | 沿用既有阅读记录 | | 84 | [Measuring AI Agent Autonomy:用 code inspection 评估 autonomy / oversight / observability]() | A | 沿用既有阅读记录 | | 85 | [Harnessing Embodied Agents:runtime governance 作为独立执行层]() | A | 沿用既有阅读记录 | | 86 | [MI9:goal-aware authorization + telemetry + conformance + containment]() · [来源 2]() | A | 沿用既有阅读记录 | | 87 | [Position: AI Agents Need Authenticated Delegation:agent 委托、scope 与审计链]() · [来源 2]() | A | 沿用既有阅读记录 | | 88 | [Unfireable Safety Kernel:execution-time authorization kernel for escapable AI systems]() | A | 沿用既有阅读记录 | | 89 | [OpenRSI / OpenMLE:可执行 AI4AI 元进化全栈(Frontis-MA1)]() · [来源 2]() | S | 沿用既有阅读记录 | | 90 | [Evolvent AI:长期 Agent、权限、skills 与 RSI 评测研究目录]() | S | 沿用既有阅读记录 | | 91 | [MetaRSI-v1 / RSI-Harness:Data、Harness、Model 三类自改进算子]() · [来源 2]() | A | 沿用既有阅读记录 | | 92 | [RSIAgent:curriculum、actor、verifier 与可迁移经验]() · [来源 2]() | A | 沿用既有阅读记录 | | 93 | [递归自我改进现场观点汇编:验证、学习过程与生态分工]() | A | 沿用既有阅读记录;公开转载或观点整理;不复制非公开原始录音、转写或内部文档,观点不作为已验证事实。 | | 94 | [红杉 RSI 讨论的公开图文整理:搜索、验证成本与自改进层级]() | A | 沿用既有阅读记录;公开转载或观点整理;不复制非公开原始录音、转写或内部文档,观点不作为已验证事实。 | | 95 | [Lilian Weng: Harness Engineering for Self-Improvement]() | S | 沿用既有阅读记录 | | 96 | [Pat Grady 的公开转载观点:计算革命、能力采纳鸿沟与应用组织]() | A | 沿用既有阅读记录;公开转载或观点整理;不复制非公开原始录音、转写或内部文档,观点不作为已验证事实。 | | 97 | [OpenAI Understands Something Important and Rare:客户创造、分发与定价]() | A | 沿用既有阅读记录 | | 98 | [五源资本:Loop Engineering——被高估的循环,被低估的拓扑]() · [来源 2]() | A | 沿用既有阅读记录 | | 99 | [Karpathy autoresearch + Loop Engineering:从人工 prompt 到 autonomous experiment loop]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() | A | 沿用既有阅读记录 | | 100 | [AMIE (Video):实时 Agent 不能只有一个 Loop——Talker/Planner/Perception 按时间尺度异步编排]() | S | 沿用既有阅读记录 | | 101 | [Reef:把推理服务器变成持续自改进 Agent 基础设施(stateful inference + experience stream + model/harness 联合演化)]() · [来源 2]() | S | 沿用既有阅读记录 | | 102 | [Microsoft ASSERT + Agent Control Specification:spec-driven agent eval / open trust stack]() | A | 沿用既有阅读记录 | | 103 | [TypeSafe System One:类型化决策、概率校准与路由]() | A | 沿用既有阅读记录 | | 104 | [Building a Harness with Jev:模型路由与工具执行前的风险门控]() | A | 沿用既有阅读记录 | | 105 | [3SPO:state-score-supervised policy optimization for long-horizon LLM agents]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 106 | [HIPIF / Context-Folding:planning + information folding for long-horizon agent context control]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 107 | [DeepSeek V4 × J-Space 能力释放报告:capability-realization loss、思维链二极管与长程状态账本]() | A | 沿用既有阅读记录 | | 108 | [ktx: open-source executable context layer for data agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 109 | [Agent Client Protocol / acpx / Paperclip ACP fit:structured session protocol vs control-plane boundary]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 110 | [stdrc:agent becomes the interface over everything]() | A | 沿用既有阅读记录 | | 111 | [Raft source-available:release mirror、prompt request 与贡献模式]() | A | 沿用既有阅读记录 | | 112 | [阿里技术:重新思考研发基础设施,当 Agent 成为第一公民]() | A | 沿用既有阅读记录 | | 113 | [平行记陆:AI Infra:跳出数据库,看到更大的 Agent Data Runtime]() | A | 沿用既有阅读记录 | | 114 | [MemPoison / Hijacking Agent Memory:绕过 selective extraction 的 long-term memory poisoning]() · [来源 2]() | A | 沿用既有阅读记录 | | 115 | [Defeating Prompt Injections by Design / CaMeL:capability-based agent security]() | A | 沿用既有阅读记录 | | 116 | [Overeager Coding Agents:Benchmarking Horizontal Privilege Escalation in Software Engineering Agents]() | A | 沿用既有阅读记录 | | 117 | [Skill-Pro / ProcMEM:从 episodic trace 学 reusable executable procedural memory]() · [来源 2]() | A | 沿用既有阅读记录 | | 118 | [SWE-Skills-Bench:skills 是否真的帮助软件工程 agent]() | A | 沿用既有阅读记录 | | 119 | [SkillLens / From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills]() · [来源 2]() | A | 沿用既有阅读记录 | | 120 | [SkillOpt: Executive Strategy for Self-Evolving Agent Skills]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 121 | [Bayesian-Agent:用 posterior-guided evidence loop 做 skill / SOP 演化]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 122 | [SkillAxe + SkillVetBench:agent skill 从 prompt 库升级成可评测、可审计、可上架的对象]() · [来源 2]() | A | 沿用既有阅读记录 | | 123 | [vLLM Hybrid SSM Disaggregated Serving:agentic serving 的 state transfer 复杂度参照]() | A | 沿用既有阅读记录 | | 124 | [SGLang HiCache / PD Disaggregation:hierarchical KV cache and prefill-decode split]() | A | 沿用既有阅读记录 | | 125 | [LMCache / vLLM APC:prefill once, reuse wherever possible]() | A | 沿用既有阅读记录 | | 126 | [dots3-note Preview:挑战 IMO 的 LLM 选手(Proof/Verify/Refine harness、TEMPO macro-step RL 与 40-50 小时 agent 负载的 serving/cache 账本)]() | S | 沿用既有阅读记录 | | 127 | [NVIDIA Dynamo agentic inference:harness -> orchestrator 的 agent hints interface]() | A | 沿用既有阅读记录 | | 128 | [Sail Research:agent inference throughput / long-running sandbox cost]() | A | 沿用既有阅读记录 | | 129 | [GRU-Mem / When to Memorize and When to Stop:长上下文 recurrent memory 的 update gate 与 exit gate]() · [来源 2]() | A | 沿用既有阅读记录 | | 130 | [OpenAI Dreaming: Better memory for a more helpful ChatGPT]() | A | 沿用既有阅读记录 | | 131 | [TaskMem:把 agent 记什么建模成 task-focused memorization policy]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 132 | [TRUSTMEM:trustworthy memory consolidation verifier]() | A | 沿用既有阅读记录 | | 133 | [AMA-Bench:real agent trajectories 上的 long-horizon memory eval]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 134 | [AgentMemoryBench: Benchmarking Continual Agent Memory for Online Learning, Transfer, and Forgetting]() · [来源 2]() | A | 沿用既有阅读记录 | | 135 | [MRAgent:Memory is Reconstructed, Not Retrieved]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 136 | [STITCH / Grounding Agent Memory in Contextual Intent]() · [来源 2]() | A | 沿用既有阅读记录 | | 137 | [Memoir / OpenViking:versioned and filesystem-like memory substrate]() · [来源 2]() | B | 沿用既有阅读记录 | | 138 | [主流 Agent Harness 实现对比:Memory 篇]() | A | 沿用既有阅读记录 | | 139 | [CodeGraph:为 Agent 预计算语义代码图与 surgical context]() | A | 沿用既有阅读记录 | | 140 | [Patronus AI digital worlds:agent simulation / stress-test environments]() | A | 沿用既有阅读记录 | | 141 | [Heuresis:quality-diversity search for autonomous AI research agents]() | A | 沿用既有阅读记录 | | 142 | [Autodata:agentic data scientist for synthetic training/eval data]() | A | 沿用既有阅读记录 | | 143 | [真实工作流与训练数据:数据市场、hillclimbability 与 verifier 约束]() | B | 沿用既有阅读记录 | | 144 | [DeepSWE:long-horizon coding-agent benchmark]() · [来源 2]() | A | 沿用既有阅读记录 | | 145 | [唐杰 / 姚顺宇 long-horizon task 观点 + ProgramBench / RepoZero]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() | A | 沿用既有阅读记录 | | 146 | [DeNovoSWE:whole-repository generation as verifiable long-horizon SWE environment]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 147 | [Tasteful Agent / Taste-Bench:长程轨迹中决策岔口的测量与训练]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录 | | 148 | [Fiona Fung / Claude Code-Cowork:What happens after coding is solved?]() | A | 沿用既有阅读记录 | | 149 | [AEvo 上游谱系精读包:DGM -> GEPA -> ADAS]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() | A | 沿用既有阅读记录 | | 150 | [Autogenesis / AGP:自进化 Agent 的协议化资源治理]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 151 | [CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 152 | [Uni-Agent:veRL 通用 Agent 构建、运行、训练统一框架]() · [来源 2]() | A | 沿用既有阅读记录 | | 153 | [XiaomiMiMo/verl 与 MiMo-V2.6 §7:多 harness rollout、轨迹结构与信用分配]() | S | 沿用既有阅读记录 | | 154 | [SearchAgent-Zero:从零训练多轮 Search Agent 的 verl RL 框架]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 155 | [Agent Lightning:Train ANY AI Agents with Reinforcement Learning]() · [来源 2]() | A | 沿用既有阅读记录 | | 156 | [UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems]() · [来源 2]() | A | 沿用既有阅读记录 | | 157 | [JitRL / Just-In-Time Reinforcement Learning:用经验记忆做 test-time policy optimization]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 158 | [Progress Advantage for LLM Agents:annotation-free step-level scoring from RL post-training]() | A | 沿用既有阅读记录 | | 159 | [Agentic Reinforced Policy Optimization / ARPO]() · [来源 2]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 160 | [Agentic Entropy-Balanced Policy Optimization / AEPO]() · [来源 2]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 161 | [Let It Flow / ROLL / iFlow-ROME:Agentic RL 的三个苦涩教训]() · [来源 2]() | A | 沿用既有阅读记录 | | 162 | [EMPO²: Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization]() | A | 沿用既有阅读记录 | | 163 | [Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() | A | 沿用既有阅读记录 | | 164 | [POM special issue + OM human-AI interaction:把 Goal Harness catalog 映射到运营管理语言]() · [来源 2]() | A | 沿用既有阅读记录 | | 165 | [Cursor Composer 2.5:targeted RL with textual feedback]() | A | 沿用既有阅读记录 | | 166 | [Codex 正在重塑传统推理框架开发流程:SGLang Diffusion、benchmark/profile、custom op 与 torch.compile graph break]() · [来源 2]() | A | 沿用既有阅读记录 | | 167 | [AI-Infra-Auto-Driven-SKILLS:SGLang/vLLM SOTA Humanize Loop 与推理框架 agent workflow]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 168 | [Kernel Design Agents:让 Agent 自己优化 CUDA kernel,并在 MLSys 2026 FlashInfer Full-Agent Track 拿下前三]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 169 | [AI agents for science & the human-value debate(Zesen Huang @ AstroAI)]() | B | 沿用既有阅读记录 | | 170 | [The State Monad:纯函数状态转换与 LoopX 状态机重构参考]() | B | 沿用既有阅读记录 | | 171 | [商业传记与通识书目](<./booklists/business-biography-and-general-knowledge.md>) | B | 仅元信息;《沃顿商学院时间管理课(修订版)》《小米创业思考》《埃隆·马斯克传》《冲向火星》《金钱博弈:重振韩国第一银行内幕》《进化论50讲》;正文未读。 | | 172 | [未核实的交易研究海报:论文核验与 backtest overfitting 方法]() | B | 沿用既有阅读记录;来源真实性未确认;保留为核验案例,不作为论文结论。 | | 173 | [量化交易教程的商业化样本:harness 结构、营销主张与证据边界]() | B | 沿用既有阅读记录 |