# 外部候选材料 [返回有排名的队列](./QUEUE.md) · [说明](./README.md) 以下记录尚未进入正式排名。序号仅便于浏览,保留原 catalog 的记录顺序,不表示优先级。 | 顺序 | 材料 | 档位 | 读取范围与备注 | | --- | --- | --- | --- | | 1 | [SWE-Bench Pro Public Dataset / Scale AI]() | A | 沿用既有阅读记录 | | 2 | [SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution]() | A | 沿用既有阅读记录 | | 3 | [LongCat-2.0]() · [来源 2]() | A | 沿用既有阅读记录 | | 4 | [AReaL 2.0 / Next-Generation Agentic RL Systems]() · [来源 2]() | A | 沿用既有阅读记录 | | 5 | [AutoMem: Automated Learning of Memory as a Cognitive Skill]() | A | 沿用既有阅读记录 | | 6 | [Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity]() | A | 沿用既有阅读记录 | | 7 | [NVIDIA ASPIRE: Agentic Skills Discovery for Robotics]() | A | 沿用既有阅读记录 | | 8 | [Safari MCP server for web developers]() | A | 沿用既有阅读记录 | | 9 | [Manufact MCP Cloud / mcp-use production lifecycle]() | A | 沿用既有阅读记录 | | 10 | [Cloudflare AI traffic options: Search / Agent / Training crawler split]() | A | 沿用既有阅读记录 | | 11 | [OpenSquilla 0.4.0 coding mode and verification workflow]() · [来源 2]() | A | 沿用既有阅读记录 | | 12 | [Raft Blog:Agent-native workspace / AX / human-agent team product language]() | A | 沿用既有阅读记录 | | 13 | [AutoPass:Evidence-Guided LLM Agents for Compiler Performance Tuning]() | A | 沿用既有阅读记录 | | 14 | [PowerAgentBench-Dyn:A Benchmark for Agentic AI in Power System Dynamic Studies]() | A | 沿用既有阅读记录 | | 15 | [RetailBench:Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments]() | A | 沿用既有阅读记录 | | 16 | [Multi-LCB:Extending LiveCodeBench to Multiple Programming Languages]() | A | 沿用既有阅读记录 | | 17 | [Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference]() | A | 沿用既有阅读记录 | | 18 | [Apodex-1.0:verification-centric deep-research agent team / AgentOS / AgentHarness]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 19 | [zartbot:用 Agentic Workflow 做大模型全栈研究 / AI-infra auto-research]() | A | 沿用既有阅读记录 | | 20 | [Oriol Vinyals / Gemini Co-Lead:World Models、Agent Scaffolding、Memory 与 Post-Training RL 路线访谈]() | A | 沿用既有阅读记录 | | 21 | [MemDecoder:把 memory composition 做成 autoregressive index decoding]() | A | 沿用既有阅读记录 | | 22 | [MemEvolve:把 agent memory architecture 本身纳入 meta-evolution]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 23 | [SimpleMem:semantic lossless compression + adaptive retrieval 的 lifelong memory substrate]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 24 | [MemOCR:把 memory 渲染成 layout-aware visual context 来分配信息密度]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 25 | [Darwinian Memory:GUI agent 的 utility-driven natural selection memory]() · [来源 2]() | A | 沿用既有阅读记录 | | 26 | [EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective]() · [来源 2]() | A | 沿用既有阅读记录 | | 27 | [MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems]() | A | 沿用既有阅读记录 | | 28 | [Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents]() | A | 沿用既有阅读记录 | | 29 | [SemiAnalysis / 腾讯科技:AI Dark Output 与 GDP 统计黑洞]() | A | 沿用既有阅读记录 | | 30 | [ACE: Agentic Context Engineering for Self-Improving Language Models]() | A | 沿用既有阅读记录 | | 31 | [ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory]() | A | 沿用既有阅读记录 | | 32 | [XENON: Experience-based Knowledge Correction for Robust Planning in Minecraft]() | A | 沿用既有阅读记录 | | 33 | [FlowSearcher: Synthesizing Memory-Guided Agentic Workflows for Web Information Seeking]() | A | 沿用既有阅读记录 | | 34 | [AgentFlow: In-the-Flow Agentic System Optimization for Effective Planning and Tool Use]() | A | 沿用既有阅读记录 | | 35 | [MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent]() | A | 沿用既有阅读记录 | | 36 | [REMem: Reasoning with Episodic Memory in Language Agent]() | A | 沿用既有阅读记录 | | 37 | [ReMemR1: Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents]() | A | 沿用既有阅读记录 | | 38 | [GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs]() | A | 沿用既有阅读记录 | | 39 | [GLoW: Dual-Scale World Memory for LLM Agents towards Hard-Exploration Problems]() | A | 沿用既有阅读记录 | | 40 | [Improving Code Localization with Repository Memory]() | A | 沿用既有阅读记录 | | 41 | [TokMem: One-Token Procedural Memory for Large Language Models]() | A | 沿用既有阅读记录 | | 42 | [Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management]() | A | 沿用既有阅读记录 | | 43 | [DataMind: Scaling Generalist Data-Analytic Agents]() | A | 沿用既有阅读记录 | | 44 | [MemGen: Weaving Generative Latent Memory for Self-Evolving Agents]() | A | 沿用既有阅读记录 | | 45 | [Agent-World:可扩展真实环境合成与自进化 Agent 训练]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 46 | [LangMem / LangGraph long-term memory docs]() | B | 沿用既有阅读记录 | | 47 | [Microsoft GraphRAG / graph-based retrieval baseline]() | B | 沿用既有阅读记录 | | 48 | [Hands-On Modern RL]() | A | 沿用既有阅读记录 | | 49 | [AgentLongBench:用 environment rollout 测 long-context agent,而不是静态 retrieval]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 50 | [PlugMem:把 raw experience 压成 knowledge-centric memory graph]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 51 | [SGLang Disaggregated Serving:P/D split 与 KV transfer 的另一套工程参照]() | B | 沿用既有阅读记录 | | 52 | [Continuum:multi-turn agent scheduling with KV cache TTL]() · [来源 2]() | A | 沿用既有阅读记录 | | 53 | [Agentic KV-cache radar:PrefillShare / TokenDance / CONCUR]() · [来源 2]() · [来源 3]() | B | 沿用既有阅读记录 | | 54 | [ThunderAgent:program-aware agentic inference system]() · [来源 2]() | A | 沿用既有阅读记录 | | 55 | [Cloudflare Code Mode MCP:把工具面压成可执行代码接口]() · [来源 2]() | A | 沿用既有阅读记录 | | 56 | [OpenAI Responses API WebSocket:agent loop state reuse as API/runtime optimization]() | A | 沿用既有阅读记录 | | 57 | [Don't Break the Cache:long-horizon agent prompt caching 的策略边界]() | A | 沿用既有阅读记录 | | 58 | [Datadog State of AI Engineering:生产 agent 的 prompt/cache/observability 统计锚点]() | A | 沿用既有阅读记录 | | 59 | [NVIDIA Extreme Co-Design for Agentic Systems:agent token economics / sub-agent / compaction 公开案例]() | A | 沿用既有阅读记录 | | 60 | [LangSmith / Langfuse observability model:trace / run / session / feedback / observation]() | A | 沿用既有阅读记录 | | 61 | [AgentTrace / AgentSight:agent observability beyond LLM tracing]() · [来源 2]() · [来源 3]() | B | 沿用既有阅读记录 | | 62 | [Lilian Weng: Why We Think]() | A | 沿用既有阅读记录 | | 63 | [ezyang: OSS code review, in the era of LLMs]() | A | 沿用既有阅读记录 | | 64 | [Workflow-aware serving radar:Helium / ForkKV / ToolCacheAgent]() · [来源 2]() · [来源 3]() | B | 沿用既有阅读记录 | | 65 | [GPU MODE Resource Stream / Reference Kernels]() · [来源 2]() · [来源 3]() | B | 沿用既有阅读记录 | | 66 | [PyTorch Dev Discuss: performance / deployment categories]() | B | 沿用既有阅读记录 | | 67 | [Claw-Eval-Live:用 live workflow demand signal 校准 agent benchmark]() · [来源 2]() | A | 沿用既有阅读记录 | | 68 | [Claw-Eval:Pass@k / Pass^k 与 trajectory-aware grading 的 agent eval 锚点]() · [来源 2]() | A | 沿用既有阅读记录 | | 69 | [AgentSwing:parallel branch + lookahead routing 的 context management 机制线索]() | A | 沿用既有阅读记录 | | 70 | [SkillFlow:lifelong skill discovery / repair / library evolution]() | A | 沿用既有阅读记录 | | 71 | [Live-Evo:online self-evolving memory from continuous feedback]() | A | 沿用既有阅读记录 | | 72 | [MemSkill:learnable and evolvable memory skills]() | A | 沿用既有阅读记录 | | 73 | [MemoryCD:cross-domain user memory benchmark from real behavior]() | A | 沿用既有阅读记录 | | 74 | [A-Evolve:agent evolution as infrastructure]() · [来源 2]() | B | 沿用既有阅读记录 | | 75 | [agent-eval:Inspect-format eval suite and leaderboard tooling]() | B | 沿用既有阅读记录 | | 76 | [Anthropic: Harness design for long-running application development]() | S | 沿用既有阅读记录 | | 77 | [10 篇论文拆解 Skill + 自进化的技术路线]() | S | 沿用既有阅读记录 | | 78 | [火山养“龙虾”日志:Self-Improving Skill 让 AI 学会“自我进化”]() | S | 沿用既有阅读记录 | | 79 | [Context Engineering 2.0: The Context of Context Engineering]() · [来源 2]() | A | 沿用既有阅读记录 | | 80 | [Anthropic: Effective context engineering for AI agents]() | A | 沿用既有阅读记录 | | 81 | [Prompting Guide: Context Engineering Guide]() | B | 沿用既有阅读记录 | | 82 | [微信导读:超越 Prompt 和 RAG,「上下文工程」成了 Agent 核心胜负手]() | B | 沿用既有阅读记录 | | 83 | [Memex(RL):stable index + full-fidelity evidence 的 indexed experience memory]() | S | 沿用既有阅读记录 | | 84 | [AgentRR:Record & Replay 范式、multi-level experience 与 check function]() | A | 沿用既有阅读记录 | | 85 | [A-MEM:Agentic Memory 的动态索引、链接与 memory evolution]() · [来源 2]() | S | 沿用既有阅读记录 | | 86 | [Mem0:生产级 agent memory 的 hybrid search、entity linking 与 token/latency 约束]() · [来源 2]() | A | 沿用既有阅读记录 | | 87 | [ProRAG:RAG 中 process reward 与 step-level credit assignment]() | A | 沿用既有阅读记录 | | 88 | [AgeMem:把长期/短期记忆管理整合进 agent policy]() | A | 沿用既有阅读记录 | | 89 | [Mem-T:Memory Operation Tree 与 MoT-GRPO 的 reward densification]() | A | 沿用既有阅读记录 | | 90 | [UMA / Learning to Remember:end-to-end memory agent、CRUD 与 Ledger-QA]() | A | 沿用既有阅读记录 | | 91 | [Memory-R1:用 RL 学习 ADD/UPDATE/DELETE/NOOP 与 memory utilization]() | A | 沿用既有阅读记录 | | 92 | [OCR-Memory:用视觉锚点保存长程 agent 轨迹并按需转录原文]() | A | 沿用既有阅读记录 | | 93 | [MemOS:把 memory 当作可管理系统资源的 Memory OS]() · [来源 2]() | A | 沿用既有阅读记录 | | 94 | [Engram:DeepSeek 条件记忆,把静态查表从动态推理中拆出来]() · [来源 2]() | B | 沿用既有阅读记录 | | 95 | [LLMs-augmented Contextual Bandit:LLM encoder + bandit 的基础概念锚点]() | B | 沿用既有阅读记录 | | 96 | [MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks]() · [来源 2]() | S | 沿用既有阅读记录 | | 97 | [AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling]() | S | 沿用既有阅读记录 | | 98 | [ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents]() | S | 沿用既有阅读记录 | | 99 | [SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees]() | S | 沿用既有阅读记录 | | 100 | [Stabilizing Reinforcement Learning with LLMs: Formulation and Practices]() | S | 沿用既有阅读记录 | | 101 | [Welcome to the Era of Experience:Sutton / Silver 的经验时代与 Agentic RL 主张]() | A | 沿用既有阅读记录 | | 102 | [DeepResearcher:真实 Web 环境中的 deep research agent RL]() · [来源 2]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 103 | [Rethinking Memory in LLM-based Agents:AI 记忆表示、六大原子操作与 Memory Compass]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 104 | [WebXSkill + Browserbase Skills:可执行 Skill 作为 Web Agent substrate]() · [来源 2]() | S | 沿用既有阅读记录 | | 105 | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment]() | A | 沿用既有阅读记录 | | 106 | [agentevals:基于 OpenTelemetry trace 的本地离线 Agent 评测]() | A | 沿用既有阅读记录 | | 107 | [BudgetMem: Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory]() | A | 沿用既有阅读记录 | | 108 | [MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents]() | A | 沿用既有阅读记录 | | 109 | [ParetoBandit: Budget-Paced Adaptive Routing for Non-Stationary LLM Serving]() | A | 沿用既有阅读记录 | | 110 | [AgentTrace: A Structured Logging Framework for Agent System Observability]() | A | 沿用既有阅读记录 | | 111 | [o11y-bench:AI Agent 的 Observability 任务 benchmark]() | A | 沿用既有阅读记录 | | 112 | [ATBench:trajectory-level agent safety benchmark]() | A | 沿用既有阅读记录 | | 113 | [mem9:OpenClaw 云端永续记忆、Memory Space 与 ContextEngine 生命周期接口]() · [来源 2]() | A | 沿用既有阅读记录 | | 114 | [ArkClaw 十大热门 Skills:技能生态、加载优先级、多 Agent 绑定与 APM 观测]() | A | 沿用既有阅读记录 | | 115 | [How To Be A World-Class Agentic Engineer:保持简单、上下文管理与 contract-driven session]() | A | 沿用既有阅读记录 | | 116 | [Agent Skills Prompt Injection:Skill 文件带来的现实攻击面]() | A | 沿用既有阅读记录 | | 117 | [Claw Code / Claude Code 开源 Rust agent harness 解读]() | A | 沿用既有阅读记录 | | 118 | [Multi-Agent 是伪技术路线?OpenClaw subagents 与多 Agent 适用边界]() · [来源 2]() | A | 沿用既有阅读记录 | | 119 | [Why Do Multi-Agent LLM Systems Fail?:MAST 多智能体失败分类与 trace 数据集]() · [来源 2]() | A | 沿用既有阅读记录 | | 120 | [Claude Code Superpowers:结构化软件工程技能框架]() · [来源 2]() | A | 沿用既有阅读记录 | | 121 | [EvoMap / Evolver:GEP 能力进化资产与跨 Agent 经验共享]() | A | 沿用既有阅读记录 | | 122 | [Agent Harness Engineering: A Survey]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录;旧记录含已读备注但受管生命周期仍为 candidate;保留状态冲突,待单独核对。 | | 123 | [Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents]() | A | 沿用既有阅读记录 | | 124 | [SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization]() · [来源 2]() | A | 沿用既有阅读记录 | | 125 | [SkillOS: Learning Skill Curation for Self-Evolving Agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 126 | [Voyager: An Open-Ended Embodied Agent with Large Language Models]() · [来源 2]() | A | 沿用既有阅读记录 | | 127 | [Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers]() | A | 沿用既有阅读记录 | | 128 | [AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems]() | A | 沿用既有阅读记录 | | 129 | [Coding Agents are Effective Long-Context Processors]() | A | 沿用既有阅读记录 | | 130 | [On Data Engineering for Scaling LLM Terminal Capabilities:Terminal Agent 数据工程]() · [来源 2]() | A | 沿用既有阅读记录 | | 131 | [A Guide to Gen AI / LLM Vibecoding for Expert Programmers]() | A | 沿用既有阅读记录 | | 132 | [Exgentic Open Agent Leaderboard:跨 benchmark 的统一 Agent 评测协议]() | A | 沿用既有阅读记录 | | 133 | [AgencyBench:1M-token long-horizon autonomous agent benchmark]() | A | 沿用既有阅读记录 | | 134 | [Agent Replay:local-first desktop evals / observability / memory]() | B | 沿用既有阅读记录 | | 135 | [OpenAI: Speeding up agentic workflows with WebSockets in the Responses API]() | A | 沿用既有阅读记录 | | 136 | [Glean Waldo:专门的 agentic search model 做检索规划]() | A | 沿用既有阅读记录 | | 137 | [CC Pocket]() | A | 沿用既有阅读记录 | | 138 | [ArkClaw 动手实验上新:财经、评论洞察、财报、视频剪辑、播客生成]() | A | 沿用既有阅读记录 | | 139 | [火山联网搜索 Skill:ArkClaw / OpenClaw 的 Agent 原生搜索能力]() | A | 沿用既有阅读记录 | | 140 | [被 OpenClaw 推上风口的飞书:企业 Agent 落地、工作空间容器与组织协作变迁]() | A | 沿用既有阅读记录 | | 141 | [碎片化编程与手机访问 Claude Code / OpenClaw]() | A | 沿用既有阅读记录 | | 142 | [Agent-Reach:开源本地互联网工具脚手架]() | A | 沿用既有阅读记录 | | 143 | [Amazon SynerGen:搜推召排一体化的生成式推荐框架]() · [来源 2]() | A | 沿用既有阅读记录 | | 144 | [快手 TagCF:用户角色、行为逻辑图与推荐系统显式破茧]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 145 | [DualGR:快手长短期兴趣建模的生成式召回实践]() · [来源 2]() | A | 沿用既有阅读记录 | | 146 | [LEMUR:大规模端到端多模态推荐]() | A | 沿用既有阅读记录 | | 147 | [Don’t Waste It:用结构化 Human Priors 指导生成式推荐多头解码]() · [来源 2]() | A | 沿用既有阅读记录 | | 148 | [SSCTL:多领域推荐中的数据不均衡与半监督迁移]() | A | 沿用既有阅读记录 | | 149 | [GNOLR:多隐式反馈的有序偏好统一嵌入与推荐召回简化]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 150 | [搜索推荐推理合一 / LLM 混合 ID 生成:NEO at Spotify]() · [来源 2]() | A | 沿用既有阅读记录 | | 151 | [Google STATIC:向量化 Trie,加速生成式召回约束解码]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 152 | [CollectiveKV:序列推荐中的跨用户 KV cache 共享与压缩]() · [来源 2]() | A | 沿用既有阅读记录 | | 153 | [Hiformer:Google Play 推荐排序中的异构特征交互 Transformer]() · [来源 2]() | A | 沿用既有阅读记录 | | 154 | [HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction]() | A | 沿用既有阅读记录 | | 155 | [X / Twitter For You 推荐算法:Grok-based Transformer、两塔召回与 Candidate Isolation]() · [来源 2]() | A | 沿用既有阅读记录 | | 156 | [Meta SilverTorch:召回与粗排全链路 GPU 算子化]() · [来源 2]() | A | 沿用既有阅读记录 | | 157 | [LORE:阿里广告搜索相关性大模型与垂直 RL 经验]() · [来源 2]() | A | 沿用既有阅读记录 | | 158 | [谈谈 RL Infra / FlashRL:训练、推理、调度、权重同步与 observability]() · [来源 2]() | A | 沿用既有阅读记录 | | 159 | [Ilya Sutskever:An Observation on Generalization / 用压缩视角理解无监督学习]() | A | 沿用既有阅读记录 | | 160 | [Muon 优化器:矩阵正交化更新、LLM 训练可扩展性与 Kimi K2 的 MuonClip]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 161 | [DeepSeek-V4 Technical Report:Million-Token Context、CSA/HCA、mHC、Muon 与长程 Agent 能力]() · [来源 2]() | A | 沿用既有阅读记录 | | 162 | [PyGraph:让 PyTorch 的 CUDA Graph 优化更高效]() · [来源 2]() | A | 沿用既有阅读记录 | | 163 | [PyTorch FlexAttention:用 score_mod / mask_mod 表达 Attention 变体并编译成高性能 Triton kernel]() · [来源 2]() | A | 沿用既有阅读记录 | | 164 | [Intel / 龙蜥:xFasterTransformer、至强 AMX/HBM/CXL 与 CPU LLM 推理背景]() | A | 沿用既有阅读记录 | | 165 | [Look Ma, No Bubbles!:Llama-1B 低延迟 Megakernel 与 memory pipeline bubbles]() · [来源 2]() | A | 沿用既有阅读记录 | | 166 | [CUDA Agent:基于真实性能反馈的 agentic RL CUDA 内核生成]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 167 | [AutoResearch 自研高性能 GPU 算子 Flash Attention:MFU 42%、Opus 4.6、8 小时 25 轮迭代]() | A | 沿用既有阅读记录 | | 168 | [DualPath:Agentic LLM Inference 的 KV-Cache 存储带宽瓶颈]() | A | 沿用既有阅读记录 | | 169 | [MetaShuffling:Llama 4 MoE 推理的 shuffle-based Fused MoE kernel 与 Padding 避免]() | A | 沿用既有阅读记录 | | 170 | [Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?]() | A | 沿用既有阅读记录 | | 171 | [少用 sense 挑战 math:如何把 post train 做好]() | A | 沿用既有阅读记录 | | 172 | [TTT-E2E:End-to-End Test-Time Training for Long Context]() · [来源 2]() | A | 沿用既有阅读记录 | | 173 | [Nested Learning / Hope:多频率持续学习、优化器即联想记忆]() · [来源 2]() | A | 沿用既有阅读记录 | | 174 | [林俊旸详解 Qwen 模型设计中的每一次取舍]() · [来源 2]() | A | 沿用既有阅读记录 | | 175 | [mHC: Manifold-Constrained Hyper-Connections]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 176 | [MiniMax VTP:Towards Scalable Pre-training of Visual Tokenizers for Generation]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 177 | [Gated Attention for Large Language Models:Non-linearity, Sparsity, and Attention-Sink-Free]() | A | 公开摘要页已核验;本次核验公开论文页面;技术细节沿用既有阅读记录。 | | 178 | [Over-Tokenized Transformer:扩大输入词表作为稀疏 scaling 维度]() | A | 沿用既有阅读记录 | | 179 | [MiniMax M2:为什么最终选择 full attention]() | A | 沿用既有阅读记录 | | 180 | [王云鹤:Harness 是复杂优化问题,Agent = Models + Harness]() | A | 沿用既有阅读记录 | | 181 | [田渊栋访谈:Meta 裁员后,对 AI 研究、RL 与职业选择的思考]() | A | 沿用既有阅读记录 | | 182 | [WhynotTV Podcast #4:翁家翌谈 OpenAI、post-training、RL infra、工程能力与 impact]() | S | 沿用既有阅读记录 | | 183 | [田渊栋:告别 OpenAI 小作文(二)未来会是什么样子]() | S | 沿用既有阅读记录 | | 184 | [Manus季逸超访谈解读(1):Benchmark是所有AI公司的唯一护城河]() | S | 沿用既有阅读记录 | | 185 | [小红书:从 L3 到 L8 分别需要怎么 vibe]() | A | 沿用既有阅读记录 | | 186 | [小红书:张咋啦 follow-builders skill / 关注 builders 而不是 KOL]() | A | 沿用既有阅读记录 | | 187 | [Aha:AI 产品真正杠杆在后端,而不是对话框]() | A | 沿用既有阅读记录 | | 188 | [如何做出好产品:张一鸣产品定律、推荐、AB 测试与务实文化]() | A | 沿用既有阅读记录 | | 189 | [对话腾讯 ima 产品团队:有价值的产品,不需要告诉用户「这是智能体」]() | A | 沿用既有阅读记录 | | 190 | [锦秋基金臧天宇:我们看 AI 应用的思考脉络]() | A | 沿用既有阅读记录 | | 191 | [Agent 经济学:一人公司、交易成本、Evals 与管理成本转移]() | B | 沿用既有阅读记录 | | 192 | [Anthropic CEO Dario Amodei 访谈:Scaling Laws、动态数据、应用层机会与 AI 时代职业判断]() | B | 沿用既有阅读记录 | | 193 | [DeepMind AlphaGo 十年复盘:游戏化、搜索、验证器与科学发现]() | B | 沿用既有阅读记录 | | 194 | [Werner Vogels final keynote:Renaissance Developer 与 AI 时代工程师价值]() | B | 沿用既有阅读记录 | | 195 | [SaaS + Agent 十人谈:盖雅工场的垂直 Agent 与结果付费判断]() | B | 沿用既有阅读记录 | | 196 | [AI Agent 产品落地五个坑:记忆、工具幻觉、调度、成本与执行精度]() | B | 沿用既有阅读记录 | | 197 | [Stripe Sessions 2026:Agentic commerce 与智能体支付基础设施]() | B | 沿用既有阅读记录 | | 198 | [胡渊鸣:生成式 AI 这门生意(上)]() | B | 沿用既有阅读记录 | | 199 | [Nicholas Wilt 推荐的计算机书单]() | B | 沿用既有阅读记录 | | 200 | [小红书:为什么人类活不过二百岁]() | B | 沿用既有阅读记录 | | 201 | [大模型第一性原理:信息论、定向信息与 token 信道抽象]() | B | 沿用既有阅读记录 | | 202 | [Artem Kirsanov:概率背后的核心方程,熵、交叉熵与 KL 散度]() | B | 沿用既有阅读记录 | | 203 | [DataFunTalk:字节和腾讯 Agentic RL 最佳实践议题线索]() | B | 沿用既有阅读记录 | | 204 | [大模型的第一性原理(一):统计物理篇]() | B | 沿用既有阅读记录 | | 205 | [Visual Language Hypothesis:为什么自监督难以直接学到语义]() · [来源 2]() | B | 沿用既有阅读记录 | | 206 | [知乎专栏文章:标题暂不可见]() | U | 未读 | | 207 | [Textual SGD: Optimizing External Memory]() | U | 未读 | | 208 | [Static AI to Continual Enterprise Learning: A Living Dialect of Tribal Knowledge]() | S | 沿用既有阅读记录 | | 209 | [Sema Code: Decoupling AI Coding Agents from the Terminal]() | S | 沿用既有阅读记录 | | 210 | [Thinking with Visual Primitives]() | A | 沿用既有阅读记录 | | 211 | [Microsoft Agent 365 GA]() | A | 沿用既有阅读记录 | | 212 | [Introducing Cloudflare Agent Cloud]() | A | 沿用既有阅读记录 | | 213 | [Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructure]() | S | 沿用既有阅读记录 | | 214 | [MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents]() | S | 沿用既有阅读记录 | | 215 | [Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers]() | A | 沿用既有阅读记录 | | 216 | [Joel Leibo 研究脉络:multi-agent 的对象不是单个 agent,而是关系结构]() | A | 沿用既有阅读记录 | | 217 | [Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces]() | S | 沿用既有阅读记录 | | 218 | [From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills]() | S | 沿用既有阅读记录 | | 219 | [GPT-5.5 Instant: smarter, clearer, and more personalized]() | A | 沿用既有阅读记录 | | 220 | [Gemini API File Search is now multimodal]() | A | 沿用既有阅读记录 | | 221 | [How OpenAI delivers low-latency voice AI at scale]() | A | 沿用既有阅读记录 | | 222 | [SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection]() | A | 沿用既有阅读记录 | | 223 | [OpenAI MRC: Supercomputer networking to accelerate large scale AI training]() | S | 沿用既有阅读记录 | | 224 | [Agents can now create Cloudflare accounts, buy domains, and deploy]() | S | 沿用既有阅读记录 | | 225 | [Anthropic: Higher usage limits for Claude and a compute deal with SpaceX]() | A | 沿用既有阅读记录 | | 226 | [Gemini API File Search is now multimodal]() | A | 沿用既有阅读记录 | | 227 | [Gemini API Webhooks for long-running jobs]() | A | 沿用既有阅读记录 | | 228 | [CopilotKit raises $27M Series A]() | A | 沿用既有阅读记录 | | 229 | [Model-Harness-Fit:同一模型换壳性能会显著漂移]() | A | 沿用既有阅读记录 | | 230 | [Anthropic: Natural Language Autoencoders: Turning Claude's Thoughts into Text]() · [来源 2]() | S | 沿用既有阅读记录 | | 231 | [Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces]() · [来源 2]() | S | 沿用既有阅读记录 | | 232 | [Do Agent Rules Shape or Distort? Guardrails Beat Guidance in Coding Agents]() | A | 沿用既有阅读记录 | | 233 | [agent-skills-eval]() | A | 沿用既有阅读记录 | | 234 | [Accelerating Gemma 4: faster inference with multi-token prediction drafters]() | A | 沿用既有阅读记录 | | 235 | [ds4: DeepSeek 4 Flash local inference engine for Metal]() | A | 沿用既有阅读记录 | | 236 | [OpenAI: Advancing voice intelligence with new models in the API]() | A | 沿用既有阅读记录 | | 237 | [AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields]() | A | 沿用既有阅读记录 | | 238 | [NVIDIA Dynamo Agent Context and Tracing]() | A | 沿用既有阅读记录 | | 239 | [Ray A2A multi-agent + Anyscale scalable MCP on Ray Serve]() | A | 沿用既有阅读记录 | | 240 | [OpenAI Codex browser / Chrome workflow signal]() | A | 沿用既有阅读记录 | | 241 | [Claude Microsoft 365 connector and enterprise context permissions]() | A | 沿用既有阅读记录 | | 242 | [STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?]() | S | 沿用既有阅读记录 | | 243 | [LLMs Corrupt Your Documents When You Delegate]() | S | 沿用既有阅读记录 | | 244 | [Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents]() | S | 沿用既有阅读记录 | | 245 | [SysMoBench / Specula: Can LLMs model real-world systems in TLA+?]() · [来源 2]() | A | 沿用既有阅读记录 | | 246 | [re_gent: Version-Control for AI coding agents]() | A | 沿用既有阅读记录 | | 247 | [A recent experience with ChatGPT 5.5 Pro]() | A | 沿用既有阅读记录 | | 248 | [MiniMax “马嘉祺”问题:后训练 token 覆盖不足导致输出映射层退化]() | A | 沿用既有阅读记录 | | 249 | [How AI-pilled are you? / P9 AI Fluency Index]() | A | 沿用既有阅读记录 | | 250 | [Jiayi Weng:Learning Beyond Gradients / Heuristic Learning]() | S | 沿用既有阅读记录 | | 251 | [MiniMax:一个 AI 还是不够 / Agent Teams 与 Mavis]() | A | 沿用既有阅读记录 | | 252 | [甲子光年:Kimi 总裁张予彤北大实录:我们想要有抽象能力和偏执的人]() | A | 沿用既有阅读记录 | | 253 | [MemReranker:Reasoning-Aware Reranking for Agent Memory Retrieval]() | S | 沿用既有阅读记录 | | 254 | [Thinking Machines:Interaction Models]() | S | 沿用既有阅读记录 | | 255 | [OpenAI:Running Codex safely at OpenAI]() | S | 沿用既有阅读记录 | | 256 | [Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning]() | S | 沿用既有阅读记录 | | 257 | [WildClawBench:Real-World Long-Horizon Agent Evaluation]() | A | 沿用既有阅读记录 | | 258 | [Needle:26M function calling model]() | A | 沿用既有阅读记录 | | 259 | [Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks]() · [来源 2]() | S | 沿用既有阅读记录 | | 260 | [OpenAI Daybreak / Codex Security control plane]() | S | 沿用既有阅读记录 | | 261 | [OpenAI: Investigating the consequences of accidentally grading CoT during RL]() | A | 沿用既有阅读记录 | | 262 | [UiPath for Coding Agents]() | A | 沿用既有阅读记录 | | 263 | [SGLang v0.5.11 release]() | A | 沿用既有阅读记录 | | 264 | [ELF: Embedded Language Flows]() · [来源 2]() | A | 沿用既有阅读记录 | | 265 | [LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues]() | S | 沿用既有阅读记录 | | 266 | [TencentDB Agent Memory]() · [来源 2]() | S | 沿用既有阅读记录 | | 267 | [Needle: 26M function call model]() | A | 沿用既有阅读记录 | | 268 | [Notion Developer Platform]() | A | 沿用既有阅读记录 | | 269 | [Claude Agent SDK credits and agent subscription economics]() | A | 沿用既有阅读记录 | | 270 | [Alexa for Shopping: personalized agentic commerce]() | A | 沿用既有阅读记录 | | 271 | [Meta Incognito Chat with Meta AI]() | B | 沿用既有阅读记录 | | 272 | [Is Grep All You Need? How Agent Harnesses Reshape Agentic Search]() | A | 沿用既有阅读记录 | | 273 | [Qoder 1.0 / Quest Mode:AI IDE -> autonomous agent workbench]() | A | 沿用既有阅读记录 | | 274 | [ChatGPT Personal Finance]() | B | 沿用既有阅读记录 | | 275 | [Grok Build CLI]() | B | 沿用既有阅读记录 | | 276 | [MeMo: Memory as a Model]() | A | 沿用既有阅读记录 | | 277 | [APWA: A Distributed Architecture for Parallelizable Agentic Workflows]() | A | 沿用既有阅读记录 | | 278 | [Ring-2.6-1T:面向 Agent 执行的开源推理模型]() | A | 沿用既有阅读记录 | | 279 | [Anthropic: Teaching Claude why]() | A | 沿用既有阅读记录 | | 280 | [AGenUI:三端原生 A2UI Renderer]() | B | 沿用既有阅读记录 | | 281 | [TwiSTAR: Think Fast, Think Slow, Then Act]() | A | 沿用既有阅读记录 | | 282 | [小红书 paper radar:4月近两周自觉不错的 paper 合集第三期]() | A | 沿用既有阅读记录 | | 283 | [ADAS / Automated Design of Agentic Systems:用搜索自动设计 agent system]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 284 | [DGM / Darwin Gödel Machine:开放式自改进 agent]() · [来源 2]() | A | 沿用既有阅读记录 | | 285 | [AFlow:自动生成 agentic workflow]() · [来源 2]() | A | 沿用既有阅读记录 | | 286 | [SPO / Self-Supervised Prompt Optimization:无人工标签的 prompt 优化]() · [来源 2]() | A | 沿用既有阅读记录 | | 287 | [GEPA / Reflective Prompt Evolution:反思式 prompt evolution]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 288 | [COMPASS:Enhancing Agent Long-Horizon Reasoning with Evolving Context]() | A | 沿用既有阅读记录 | | 289 | [Training-Free Group Relative Policy Optimization]() | A | 沿用既有阅读记录 | | 290 | [Motive Notes:What Makes 5% of AI Agents Actually Work in Production?]() | A | 沿用既有阅读记录 | | 291 | [AgentEvolver:Towards Efficient Self-Evolving Agent System]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 292 | [To The Crazy Ones \| 致超级个体]() | A | 沿用既有阅读记录;既往已读文字正文,视频、图片与白板未逐一展开。讨论完整问题所有权、工具权限、直接用户反馈与组织认可如何支持 AI Builder。 | | 293 | [Codex / Claude Code memory 模式的收益、风险与 token 成本]() · [来源 2]() | A | 沿用既有阅读记录 | | 294 | [SkyRL:full-stack RL library for LLMs]() · [来源 2]() | A | 沿用既有阅读记录 | | 295 | [OpenRLHF:high-performance RLHF / GRPO / PPO framework]() | A | 沿用既有阅读记录 | | 296 | [Selective Rollout:Efficient Reinforcement Learning for Long-Horizon LLM Agents]() | A | 沿用既有阅读记录 | | 297 | [HiPER:State Abstraction and Value-Guided Search for Efficient Long-Horizon Agents]() | A | 沿用既有阅读记录 | | 298 | [AgentFly:Extensible and Scalable Reinforcement Learning for LM Agents]() | A | 沿用既有阅读记录 | | 299 | [Code as Agent Harness]() · [来源 2]() | S | 沿用既有阅读记录 | | 300 | [SWE-Chain:Benchmarking Coding Agents on Chained Release-Level Package Upgrades]() | S | 沿用既有阅读记录 | | 301 | [AgentTrust:Runtime Safety Evaluation and Interception for AI Agent Tool Use]() | A | 沿用既有阅读记录 | | 302 | [Google I/O 2026:Gemini 3.5 Flash / Antigravity 2.0 / Managed Agents]() | A | 沿用既有阅读记录 | | 303 | [OpenAI / Databricks:GPT-5.5 on OfficeQA Pro]() | A | 沿用既有阅读记录 | | 304 | [Anthropic acquires Stainless:SDK / CLI / MCP server tooling]() | A | 沿用既有阅读记录 | | 305 | [RRFP:A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability]() | A | 沿用既有阅读记录 | | 306 | [DashAttention:Differentiable and Adaptive Sparse Hierarchical Attention]() | A | 沿用既有阅读记录 | | 307 | [Qwen3.7-Max Agent Frontier]() | A | 沿用既有阅读记录 | | 308 | [OpenAI Guaranteed Capacity:模型 API 的 reserved compute contract]() | A | 沿用既有阅读记录 | | 309 | [TIDE:Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload]() | A | 沿用既有阅读记录 | | 310 | [KoRe:Compact Knowledge Representations for Large Language Models]() | A | 沿用既有阅读记录 | | 311 | [OpenAI model disproves Erdős unit-distance conjecture]() | A | 沿用既有阅读记录 | | 312 | [Multi-Stream LLMs:Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs]() | A | 沿用既有阅读记录 | | 313 | [vLLM x PegaFlow:Production-Grade External KV Cache]() | A | 沿用既有阅读记录 | | 314 | [AWS SageMaker OpenAI-compatible API endpoints]() | A | 沿用既有阅读记录 | | 315 | [Google AI Mode Ads:Gemini-powered ad formats in Search]() | A | 沿用既有阅读记录 | | 316 | [Tencent Hy-MT2:translation model family and IFMTBench]() | A | 沿用既有阅读记录 | | 317 | [Cohere Command A+:可私有部署的 open-source enterprise agent model]() | A | 沿用既有阅读记录 | | 318 | [Domain-Camouflaged Injection Attacks:领域伪装注入绕过 agent guard]() | A | 沿用既有阅读记录 | | 319 | [唐杰:关于 long-horizon tasks 的近期思考]() | A | 沿用既有阅读记录 | | 320 | [Runtime (YC P26):团队级沙盒化 coding agents]() | B | 沿用既有阅读记录 | | 321 | [Claude Code 源码精读:compact 每次压缩时发生了什么]() | A | 沿用既有阅读记录 | | 322 | [Neuromancer:Claude Code 上下文管理个人学习笔记]() | A | 沿用既有阅读记录 | | 323 | [MemGym:长程 agent memory 执行评测环境]() | A | 沿用既有阅读记录 | | 324 | [Anthropic Project Glasswing / Claude Mythos CVD dashboard]() | A | 沿用既有阅读记录 | | 325 | [CODA:Rewriting Transformer Blocks as GEMM-Epilogue Programs]() | A | 沿用既有阅读记录 | | 326 | [KanBots OSS:本地 Kanban 多 agent 编排]() · [来源 2]() | B | 沿用既有阅读记录 | | 327 | [ReAct:reasoning-action-observation 循环的经典起点]() | A | 沿用既有阅读记录 | | 328 | [Toolformer:模型自监督学会调用工具]() | A | 沿用既有阅读记录 | | 329 | [API-Bank:tool-augmented LLM 的 API 评测早期基准]() | A | 沿用既有阅读记录 | | 330 | [Gorilla:面向海量 API 的 tool retrieval / function calling]() · [来源 2]() | A | 沿用既有阅读记录 | | 331 | [ToolLLM / ToolBench:大规模真实 API 的工具学习与评测]() · [来源 2]() | A | 沿用既有阅读记录 | | 332 | [WebArena:真实 Web 环境中的 autonomous agent benchmark]() | A | 沿用既有阅读记录 | | 333 | [VisualWebArena:多模态 Web agent 评测]() | A | 沿用既有阅读记录 | | 334 | [OSWorld:真实桌面环境中的 computer-use agent benchmark]() · [来源 2]() | A | 沿用既有阅读记录 | | 335 | [WorkArena:企业知识工作 Web agent benchmark]() · [来源 2]() | A | 沿用既有阅读记录 | | 336 | [SWE-bench:真实 GitHub issue 到 patch 的软件工程评测]() | A | 沿用既有阅读记录 | | 337 | [OpenHands:通用软件开发 agent 平台]() · [来源 2]() | A | 沿用既有阅读记录 | | 338 | [AutoGen:multi-agent conversation framework]() · [来源 2]() | A | 沿用既有阅读记录 | | 339 | [LLMCompiler:并行 function calling / tool execution 编排]() | A | 沿用既有阅读记录 | | 340 | [AgentLens:agent 行为可视分析与 lucky pass 问题]() · [来源 2]() | A | 沿用既有阅读记录 | | 341 | [Contextual Agent Security:面向不同目的的 agent policy]() | A | 沿用既有阅读记录 | | 342 | [Generative Agents:长期记忆驱动的交互式行为模拟]() | A | 沿用既有阅读记录 | | 343 | [Lost in the Middle:长上下文位置偏置经典问题]() | A | 沿用既有阅读记录 | | 344 | [MemGPT:把 LLM memory 管理类比为操作系统]() · [来源 2]() | A | 沿用既有阅读记录 | | 345 | [OpenShell:声明式 policy 驱动的 agent sandbox runtime]() | A | 沿用既有阅读记录 | | 346 | [SWE-ReX:coding agent remote execution / sandbox infrastructure]() | A | 沿用既有阅读记录 | | 347 | [ContextForge:MCP / A2A / REST gateway with governance and observability]() | A | 沿用既有阅读记录 | | 348 | [Agent Governance Toolkit:deterministic policy / identity / sandbox / audit before actions]() | A | 沿用既有阅读记录 | | 349 | [Browser Harness:可编辑 CDP browser harness]() | A | 沿用既有阅读记录 | | 350 | [Symphony:ticket-driven orchestration layer for autonomous implementation runs]() | A | 沿用既有阅读记录 | | 351 | [R2E-Gym:从真实 repo issue 构造 executable coding-agent RL environments]() · [来源 2]() | A | 沿用既有阅读记录 | | 352 | [Prime Intellect verifiers:LLM RL environments + evals as reusable verifier library]() | A | 沿用既有阅读记录 | | 353 | [Meta-Harness:把 harness design 本身作为 automated search object]() | A | 沿用既有阅读记录 | | 354 | [Anthropic Context Management:tool result clearing and compaction]() | A | 沿用既有阅读记录 | | 355 | [Context Rot:long context 变长时的性能退化]() | A | 沿用既有阅读记录 | | 356 | [Anthropic: How we built our multi-agent research system]() | A | 沿用既有阅读记录 | | 357 | [LCGuard:Defending Against Latent Communication in Multi-Agent Systems by System-Level KV Cache Sandboxing]() | A | 沿用既有阅读记录 | | 358 | [Claude Code network sandbox bypass reports]() | A | 沿用既有阅读记录 | | 359 | [Cloudflare Agent Infrastructure Stack]() | A | 沿用既有阅读记录 | | 360 | [Reasonix:prefix-cache-aware terminal coding agent]() | A | 沿用既有阅读记录 | | 361 | [FAME:Fault-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection]() | A | 沿用既有阅读记录 | | 362 | [Epoch AI: AI chip component cost shares]() | A | 沿用既有阅读记录 | | 363 | [Kung & Robinson: On Optimistic Methods for Concurrency Control]() | A | 沿用既有阅读记录 | | 364 | [Shapiro et al.: A comprehensive study of convergent and commutative replicated data types(CRDTs)]() | A | 沿用既有阅读记录 | | 365 | [AeSlides:通过可验证奖励强化幻灯片生成]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 366 | [ACC: Compiling Agent Trajectories for Long-Context Training]() | A | 沿用既有阅读记录 | | 367 | [SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?]() · [来源 2]() | A | 沿用既有阅读记录 | | 368 | [π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows]() · [来源 2]() | A | 沿用既有阅读记录 | | 369 | [Microsoft Copilot Cowork Exfiltrates Files]() | A | 沿用既有阅读记录 | | 370 | [Constraint Decay: The Fragility of LLM Agents in Backend Code Generation]() | A | 沿用既有阅读记录 | | 371 | [CVEvolve: Autonomous Algorithm Discovery for Unstructured Scientific Data Processing]() | A | 沿用既有阅读记录 | | 372 | [Gemini app becomes more agentic, delivering proactive 24/7 help]() | B | 沿用既有阅读记录 | | 373 | [Anthropic: How we contain Claude across products]() | A | 沿用既有阅读记录 | | 374 | [Language Models Need Sleep]() | A | 沿用既有阅读记录 | | 375 | [EAGLE 3.1: Advancing Speculative Decoding Through Collaboration Between EAGLE, vLLM, and TorchSpec]() | A | 沿用既有阅读记录 | | 376 | [Robin: A multi-agent system for automating scientific discovery]() | A | 沿用既有阅读记录 | | 377 | [Alipay AI Wallet / Token Pay / Agentic Commerce Trust Protocol]() | A | 沿用既有阅读记录 | | 378 | [OpenRouter Raises $113M Series B]() | A | 沿用既有阅读记录 | | 379 | [Xiaomi MiMo-V2.5 Series Price Adjustment]() | A | 沿用既有阅读记录 | | 380 | [Minicor: managed self-healing desktop automation at scale]() | A | 沿用既有阅读记录 | | 381 | [中国企业家:6个月融25亿元,他是“字节系”最猛的AI创业者]() | A | 沿用既有阅读记录 | | 382 | [ArkClaw 漫剧虾工作流实测:从一个主题到爆款漫剧成片]() | A | 沿用既有阅读记录 | | 383 | [视频生成 agent / 短剧工具竞品池:Flova / 纳米短剧 / 巨日禄 / 万镜一刻]() | A | 沿用既有阅读记录 | | 384 | [火山引擎:Vibe Creating,让视频创作回归表达本身]() | A | 沿用既有阅读记录 | | 385 | [MUSE-Autoskill:Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 386 | [QUEST:Training Frontier Deep Research Agents with Fully Synthetic Tasks]() | A | 沿用既有阅读记录 | | 387 | [Cognition: More Devins in More Places]() | A | 沿用既有阅读记录 | | 388 | [jxnlco: Getting the most out of Codex]() | A | 沿用既有阅读记录 | | 389 | [Polar: Agentic RL on Any Harness at Scale]() · [来源 2]() | A | 沿用既有阅读记录 | | 390 | [Structured Agent Distillation for Large Language Model]() | A | 沿用既有阅读记录 | | 391 | [PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective]() | A | 沿用既有阅读记录 | | 392 | [Lenz Research: Beyond Benchmarks, Frontier LLM Disagreement on Fact-Checks]() | A | 沿用既有阅读记录 | | 393 | [DBOS: Postgres is All You Need for Durable Workflows]() | A | 沿用既有阅读记录 | | 394 | [OpenAI: How OpenAI uses Codex]() | A | 沿用既有阅读记录 | | 395 | [ClickHouse Agents + Langfuse V4:agentic data stack and observability]() | A | 沿用既有阅读记录 | | 396 | [Claude Code vs Codex scientific-computing head-to-head]() | A | 沿用既有阅读记录 | | 397 | [Coding Beyond Your Training: Claude Code and the Technological Frontier of Software Developers]() | A | 沿用既有阅读记录 | | 398 | [SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?]() | A | 沿用既有阅读记录 | | 399 | [LLMSurgeon: Diagnosing Data Mixture of Large Language Models]() | A | 沿用既有阅读记录 | | 400 | [In-Context Reward Adaptation for Robust Preference Modeling]() | A | 沿用既有阅读记录 | | 401 | [Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players]() | A | 沿用既有阅读记录 | | 402 | [WALL-WM:World Action Model at Event Boundaries]() · [来源 2]() | A | 沿用既有阅读记录 | | 403 | [ATLAS - Autoformalized Textbook Library At Scale]() | A | 沿用既有阅读记录 | | 404 | [tiny-vLLM:Build your own high performance LLM inference engine in C++ and CUDA]() | A | 沿用既有阅读记录 | | 405 | [Never Stop Learning: Continual Learning and Self-Iteration in LLMs]() | A | 沿用既有阅读记录 | | 406 | [ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents]() | A | 沿用既有阅读记录 | | 407 | [Is Agent Memory a Database? Rethinking Data Foundations for Long-Term AI Agent Memory]() | A | 沿用既有阅读记录 | | 408 | [GodeX:OpenAI Responses API 兼容网关与 provider bridge]() | A | 沿用既有阅读记录 | | 409 | [Meta AI support / Instagram account recovery exploit report]() | A | 沿用既有阅读记录 | | 410 | [Microsoft MAI / Frontier Tuning:workflow-specific RLE 与 Copilot harness model]() | A | 沿用既有阅读记录 | | 411 | [OpenAI Codex for every role / Sites / annotations]() | A | 沿用既有阅读记录 | | 412 | [Microsoft Scout / WorkIQ / Agent 365:always-on work agent control plane]() | A | 沿用既有阅读记录 | | 413 | [Bernini: Latent Semantic Planning for Video Diffusion]() | A | 沿用既有阅读记录 | | 414 | [AdaCodec: A Predictive Visual Code for Video MLLMs]() | A | 沿用既有阅读记录 | | 415 | [CLI-Anything: Towards Agent-Native Computer Use]() · [来源 2]() | A | 沿用既有阅读记录 | | 416 | [EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management]() · [来源 2]() | A | 沿用既有阅读记录 | | 417 | [Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation]() | A | 沿用既有阅读记录 | | 418 | [Bringing up DeepSeek-V4-Flash on AMD MI300X]() · [来源 2]() | B | 沿用既有阅读记录 | | 419 | [Uber AI coding budget cap / enterprise agent FinOps]() | B | 沿用既有阅读记录 | | 420 | [MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents]() | A | 沿用既有阅读记录 | | 421 | [KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks]() · [来源 2]() | A | 沿用既有阅读记录 | | 422 | [Multi-Segment Attention / AsymCache:面向 agent serving 的 KV-cache 管理]() | A | 沿用既有阅读记录 | | 423 | [Google Gemma 4 12B: unified encoder-free multimodal model]() · [来源 2]() | A | 沿用既有阅读记录 | | 424 | [Anthropic / Sakana AI recursive self-improvement signals]() | A | 沿用既有阅读记录 | | 425 | [Microsoft pg_durable: PostgreSQL in-database durable execution]() | A | 沿用既有阅读记录 | | 426 | [Anthropic Defending Code Reference Harness]() | A | 沿用既有阅读记录 | | 427 | [Alibaba Open Code Review: deterministic engineering + LLM agent code review]() | A | 沿用既有阅读记录 | | 428 | [Cloudflare AI Gateway spend limits / identity-driven budgets]() | B | 沿用既有阅读记录 | | 429 | [Lowfat: local CLI output filtering for agent token budgets]() | B | 沿用既有阅读记录 | | 430 | [AdaMEM: Test-Time Adaptive Memory for Language Agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 431 | [Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents]() | A | 沿用既有阅读记录 | | 432 | [Description-Code Inconsistency in Real-world MCP Servers]() | A | 沿用既有阅读记录 | | 433 | [MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery]() · [来源 2]() | A | 沿用既有阅读记录 | | 434 | [CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments]() | A | 沿用既有阅读记录 | | 435 | [Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill]() | A | 沿用既有阅读记录 | | 436 | [Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators]() | A | 沿用既有阅读记录 | | 437 | [TechCrunch: The token bill comes due]() | B | 沿用既有阅读记录 | | 438 | [Poke becomes the first AI agent on Apple Messages for Business]() | B | 沿用既有阅读记录 | | 439 | [知乎回答:王导是也缩略《置身钉内》与钉钉 ONE 项目复盘]() | B | 沿用既有阅读记录 | | 440 | [小红书:学习如何从 0 训练一个 SOTA LLM]() | B | 沿用既有阅读记录 | | 441 | [Motus: A Unified Latent Action World Model]() | A | 沿用既有阅读记录 | | 442 | [AKO: Agentic Kernel Optimization / AKO4ALL / AKO4X]() · [来源 2]() · [来源 3]() | S | 沿用既有阅读记录 | | 443 | [Nano World Models: minimalist world-model experiment substrate]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 444 | [OpenAI Lockdown Mode:prompt injection 数据外泄防线的产品化 capability gate]() | A | 沿用既有阅读记录 | | 445 | [Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads]() | A | 沿用既有阅读记录 | | 446 | [SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 447 | [TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration]() · [来源 2]() | A | 沿用既有阅读记录 | | 448 | [Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement]() | A | 沿用既有阅读记录 | | 449 | [Google / SpaceX AI compute deal:短期 bridge capacity 与 frontier agent demand 的市场信号]() | B | 沿用既有阅读记录 | | 450 | [Cloudflare Bot Traffic Radar:agentic web traffic 与 origin cost 进入一等指标]() | B | 沿用既有阅读记录 | | 451 | [Tencent Productivity Agent Suite / CodeBuddy / WorkBuddy / Agent Runtime / TokenHub]() | B | 沿用既有阅读记录 | | 452 | [FrontierCode:从 correctness 到 production mergeability 的 coding-agent benchmark]() | A | 沿用既有阅读记录 | | 453 | [KV cache serving exactness:Speculative KV coding + VeriCache]() · [来源 2]() | A | 沿用既有阅读记录 | | 454 | [Tokenomics / token bill:agentic software cost observability]() · [来源 2]() | A | 沿用既有阅读记录 | | 455 | [Intuned Agent:browser automation codegen + managed Playwright runtime]() | B | 沿用既有阅读记录 | | 456 | [Microsoft AI developer tooling supply-chain incident]() | B | 沿用既有阅读记录 | | 457 | [AGENTS.md / context-file tooling:agent-md-bench + context file evidence]() | B | 沿用既有阅读记录 | | 458 | [SWE-Explore:repository exploration benchmark for coding agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 459 | [End-to-End Context Compression at Scale / LCLM]() · [来源 2]() | A | 沿用既有阅读记录 | | 460 | [Anthropic biology agents:domain data infrastructure for agents]() | A | 沿用既有阅读记录 | | 461 | [Claude Fable 5 / Mythos 5:frontier capability with gated access]() | B | 沿用既有阅读记录 | | 462 | [Nemotron 3 Ultra serving stack:vLLM / SGLang / Miles day-zero path]() | B | 沿用既有阅读记录 | | 463 | [AI developer tooling security:Microsoft repo incident + SGLang RCE]() · [来源 2]() | B | 沿用既有阅读记录 | | 464 | [Asuka-Bench:underspecified intent + multi-round refinement for code agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 465 | [DiffusionGemma:parallel text diffusion for local interactive workflows]() | A | 沿用既有阅读记录 | | 466 | [GitHub Agent Apps + Copilot Code Review skills/MCP]() | A | 沿用既有阅读记录 | | 467 | [MusaCoder:native GPU kernel generation with full-stack training on Moore Threads GPU]() · [来源 2]() | A | 沿用既有阅读记录 | | 468 | [Apache Burr:state machine / telemetry / persistence for reliable AI apps]() | A | 沿用既有阅读记录 | | 469 | [Memory tools can make AI models worse:memory reliability as product risk]() · [来源 2]() | B | 沿用既有阅读记录 | | 470 | [MiMo Code:long-horizon coding agent with persistent project memory]() · [来源 2]() | A | 沿用既有阅读记录 | | 471 | [Claw Patrol:wire-level security firewall for agents]() | A | 沿用既有阅读记录 | | 472 | [Google DeepMind multi-agent AI safety research fund]() | A | 沿用既有阅读记录 | | 473 | [Coinbase for Agents / x402:agentic payments and paid resource access]() | A | 沿用既有阅读记录 | | 474 | [Alibaba Cloud Meoo CLI:local coding agent to cloud deployment bridge]() | B | 沿用既有阅读记录 | | 475 | [Can I Buy Your KV Cache?:agent-native prefill CDN / hot context cost model]() | A | 沿用既有阅读记录 | | 476 | [ReSum:self-summary as policy action for long reasoning RLVR]() | A | 沿用既有阅读记录 | | 477 | [AgentBeats:agentified agent assessment via A2A / MCP]() | A | 沿用既有阅读记录 | | 478 | [Agents-K1:agent-native scientific knowledge orchestration]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() · [来源 6]() | A | 沿用既有阅读记录 | | 479 | [EurekAgent:environment engineering for autonomous scientific discovery]() | A | 沿用既有阅读记录 | | 480 | [SkillSpector:agent skill supply-chain scanner]() | A | 沿用既有阅读记录 | | 481 | [GLM-5:from vibe coding to agentic engineering]() · [来源 2]() · [来源 3]() · [来源 4]() | A | 沿用既有阅读记录 | | 482 | [LongTraceRL:learning long-context reasoning from search-agent trajectories]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 483 | [Plan-RewardBench / VPR:trajectory-level reward modeling and verifiable process reward for agents]() · [来源 2]() · [来源 3]() · [来源 4]() · [来源 5]() | A | 沿用既有阅读记录 | | 484 | [EvoArena / EvoMem:tracking memory evolution for robust LLM agents]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 485 | [depthfirst 21 FFmpeg zero-days:autonomous security agent with reproducible PoCs]() | A | 沿用既有阅读记录 | | 486 | [OpenAI Codex enterprise workflow cases:Notion / Nextdoor / Wasmer outcome engineering]() | A | 沿用既有阅读记录 | | 487 | [Microsoft Discovery GA:governed agentic R&D workflows]() | A | 沿用既有阅读记录 | | 488 | [TensorZero archive signal:LLMOps open-source continuity risk]() · [来源 2]() | B | 沿用既有阅读记录 | | 489 | [OpenAI multistate investigation:AI product safety and personalization audit risk]() | B | 沿用既有阅读记录 | | 490 | [WeaveBench:hybrid-interface long-horizon computer-use agent benchmark]() · [来源 2]() | A | 沿用既有阅读记录 | | 491 | [TRACE:compiling user corrections into runtime enforcement for coding agents]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 492 | [HarnessBridge:learnable bidirectional controller for LLM agent harness]() · [来源 2]() | A | 沿用既有阅读记录 | | 493 | [EvoBrowseComp:benchmarking search agents on evolving knowledge]() · [来源 2]() | A | 沿用既有阅读记录 | | 494 | [Perplexity / HBS:How AI Agents Reshape Knowledge Work]() | A | 沿用既有阅读记录 | | 495 | [Google DeepMind From AGI to ASI:multi-agent collectives as one ASI pathway]() · [来源 2]() | B | 沿用既有阅读记录 | | 496 | [OpenAI Partner Network:enterprise AI delivery and specialization market]() | B | 沿用既有阅读记录 | | 497 | [Context window budget:Don't trust large context windows]() | B | 沿用既有阅读记录 | | 498 | [AI provenance failure signal:UK police fake-evidence allegation + KPMG hallucinated report]() | B | 沿用既有阅读记录 | | 499 | [Gabriel Weinberg:No, everyone is not using AI for everything]() | B | 沿用既有阅读记录 | | 500 | [HarnessX:composable, adaptive, evolvable agent harness foundry]() | A | 沿用既有阅读记录 | | 501 | [StreamMemBench:streaming evaluation of agent memory for future-oriented assistance]() · [来源 2]() | A | 沿用既有阅读记录 | | 502 | [Dialogue SWE-Bench:benchmarking dialogue-driven coding agents]() · [来源 2]() | A | 沿用既有阅读记录 | | 503 | [Parallel-Synthesis:direct latent-space synthesis for parallel branches in LLM-agent workflows]() | A | 沿用既有阅读记录 | | 504 | [SIMMER:latent failures in LLM executable planning with a world model]() · [来源 2]() | A | 沿用既有阅读记录 | | 505 | [Google Cloud OKF:Open Knowledge Format for agent-readable context bundles]() · [来源 2]() | A | 沿用既有阅读记录 | | 506 | [OpenRouter Fusion:model panels as an API-level reasoning primitive]() | A | 沿用既有阅读记录 | | 507 | [Apple Foundation Models / Xcode 27:native provider protocol and agentic coding workflow]() | A | 沿用既有阅读记录 | | 508 | [Enterprise agent identity and service-agent consolidation:NewCore + Salesforce / Fin]() | B | 沿用既有阅读记录 | | 509 | [Measuring Agents in Production:真实生产 agent 仍是高频人工介入系统]() | A | 沿用既有阅读记录 | | 510 | [Principles of Mixed-Initiative User Interfaces:混合主动权的经典设计原则]() | A | 沿用既有阅读记录 | | 511 | [Power to the People:Interactive ML 中人的角色]() | A | 沿用既有阅读记录 | | 512 | [Evaluation of Interactive Machine Learning Systems:algorithm-centered + human-centered 双验证]() | A | 沿用既有阅读记录 | | 513 | [A Benchmark for Scalable Oversight Mechanisms:监督机制也需要 benchmark]() | A | 沿用既有阅读记录 | | 514 | [Deep Reinforcement Learning from Human Preferences:少量偏好监督如何塑造复杂目标]() | A | 沿用既有阅读记录 | | 515 | [EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments]() | S | 沿用既有阅读记录 | | 516 | [Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents]() | S | 沿用既有阅读记录 | | 517 | [OpenRouter Fusion:多模型 deliberation 作为 API runtime primitive]() | A | 沿用既有阅读记录 | | 518 | [OpenAI Deployment Simulation:用真实分布预演模型上线行为]() | S | 沿用既有阅读记录 | | 519 | [AA-AgentPerf:agentic inference 的 SLO / agents-per-megawatt 口径]() | A | 沿用既有阅读记录 | | 520 | [GLM-5.2:open-weight long-horizon agent model]() · [来源 2]() | A | 沿用既有阅读记录 | | 521 | [AI Coding Agents Can Reproduce Social Science Findings / SocSci-Repro-Bench]() | A | 沿用既有阅读记录 | | 522 | [LifeSciBench + AI Chemist:science agent 的专家 rubric 与湿实验闭环]() | A | 沿用既有阅读记录 | | 523 | [Appia / Pramaana:AI trust 从 policy 走向 conformity + proof]() | A | 沿用既有阅读记录 | | 524 | [MCP Enterprise-Managed Authorization:Zero-touch OAuth for MCP]() | A | 沿用既有阅读记录 | | 525 | [Elastic agent memory:hybrid retrieval + DLS 的生产 memory 参考]() | A | 沿用既有阅读记录 | | 526 | [Decoupled Search Grounding:把 agent search 变成 MCP-compatible gateway]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 527 | [RODS:multi-turn tool-use agent 的 reward-driven online data synthesis]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 528 | [CEO-Bench:long-horizon business agent eval]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 529 | [SGCD:GUI agent 的 off-trajectory continuation distillation]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 530 | [Xcientist:AI scientist 的 research harness 与 claim drift 防线]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 531 | [Stanford PhD 回山东做传统企业 AI 落地:代码快,诊断慢]() | A | 沿用既有阅读记录 | | 532 | [Google Agentic Resource Discovery:agent capability discovery + trust manifest]() | A | 沿用既有阅读记录 | | 533 | [Claude Design + Claude Code sync:design-to-code workspace 进入真实组件回路]() | A | 沿用既有阅读记录 | | 534 | [SkillVetBench:LLM agent skills 的语义风险评估]() | A | 沿用既有阅读记录 | | 535 | [ADK Arena:用 LLM-as-a-Developer 测 agent framework 可用性]() | A | 沿用既有阅读记录 | | 536 | [Agent Planning Benchmark:把 planning failure 从执行失败里拆出来]() | A | 沿用既有阅读记录 | | 537 | [R3-Skill:skill routing 中的 rejection signal 不是垃圾数据]() | A | 沿用既有阅读记录 | | 538 | [Exploration Structure in LLM Agents:coding agent 的 repo traversal 结构会决定定位质量]() | A | 沿用既有阅读记录 | | 539 | [AgentFairBench:agent 公平性要测 action,不只测 answer]() | A | 沿用既有阅读记录 | | 540 | [Kimi Work / Kimi Code Goal Mode:本地长程 agent 的 goal state 与权限面]() | A | 沿用既有阅读记录 | | 541 | [OpenAI Patch the Planet:安全 agent 的 patch loop 与 maintainer agency]() | A | 沿用既有阅读记录 | | 542 | [Google Interactions API GA:managed agent API contract]() | A | 沿用既有阅读记录 | | 543 | [Multi-LCB + Contagion Networks:coding agent 评测的语言轴与 judge topology]() · [来源 2]() | A | 沿用既有阅读记录 | | 544 | [Liquid AI LFM2.5 Retrievers:本地多语言 memory/search 检索底座]() | A | 沿用既有阅读记录 | | 545 | [xAI Grok Build /goal:long-running coding agent 的目标状态与验证面]() | A | 沿用既有阅读记录 | | 546 | [Claude Tag:Slack 中的 scoped team agent 与组织级权限/成本控制]() | A | 沿用既有阅读记录 | | 547 | [The Coming Loop:harness-level loop 会放大工程质量债]() | A | 沿用既有阅读记录 | | 548 | [Mistral OCR 4 + Baidu Unlimited-OCR:文档 ingestion 从 OCR 走向结构化 context substrate]() · [来源 2]() | A | 沿用既有阅读记录 | | 549 | [Randomized YaRN:短上下文训练也能改善 16K-128K 长上下文推理泛化]() | A | 沿用既有阅读记录 | | 550 | [AIR:用 RL 学会何时在多模态推理中调用代码工具]() | A | 沿用既有阅读记录 | | 551 | [Can LLMs Reliably Self-Report Adversarial Prefills:模型自我报告不能当安全证据]() | A | 沿用既有阅读记录 | | 552 | [VibeThinker-3B:小模型 verifiable reasoning 的成本/能力边界]() | A | 沿用既有阅读记录 | | 553 | [Gemini 3.5 Flash computer use:UI agent 的 observe-act-screenshot contract]() | A | 沿用既有阅读记录 | | 554 | [Qwen-AgentWorld:language world model for agentic RL]() · [来源 2]() | A | 沿用既有阅读记录 | | 555 | [AOHP:Android Open Harness Project / OS-level agent harness]() · [来源 2]() | A | 沿用既有阅读记录 | | 556 | [OpenThoughts-Agent:agentic model data recipes]() · [来源 2]() | A | 沿用既有阅读记录 | | 557 | [Tmax:terminal-agent RL recipe and TMax-15K]() · [来源 2]() · [来源 3]() | A | 沿用既有阅读记录 | | 558 | [OpenAI/Broadcom Jalapeno:inference cost as agent runtime constraint]() | A | 沿用既有阅读记录 | | 559 | [Qualcomm 收购 Modular:portable AI serving/software stack signal]() | A | 沿用既有阅读记录 | | 560 | [Headroom:agent context compression / CCR gate]() | A | 沿用既有阅读记录 | | 561 | [Notion Mail agent takeover + General Intuition action-labeled world model data]() | A | 沿用既有阅读记录 | | 562 | [OpenAI GPT-5.6 Sol limited preview:frontier model release gate and ultra subagents]() | A | 沿用既有阅读记录 | | 563 | [OpenAI Codex economic research:agents transform work into delegated long-horizon tasks]() | A | 沿用既有阅读记录 | | 564 | [Cursor reward hacking in coding benchmarks:strict harness for aware coding agents]() | A | 沿用既有阅读记录 | | 565 | [Workweave Router:cache-aware model routing for Claude Code / Codex / Cursor]() | A | 沿用既有阅读记录 | | 566 | [NVIDIA NeMo AutoModel / Transformers v5 Expert Parallelism + DeepEP MoE fine-tuning path]() · [来源 2]() | A | 沿用既有阅读记录 | | 567 | [PEEU GUI agents:Autonomous Experience Exploration + Hindsight Experience Utilization for task planning]() | A | 沿用既有阅读记录 | | 568 | [When are likely answers right? Sequence Probability and Correctness in LLMs]() | A | 沿用既有阅读记录 | | 569 | [Un-0:open coupled-oscillator image generator as physical-compute substrate]() · [来源 2]() | A | 沿用既有阅读记录 | | 570 | [OpenAI + Broadcom Jalapeno inference chip:full-stack LLM inference platform]() | A | 沿用既有阅读记录 | | 571 | [DeepSeek DSpark / DeepSpec:confidence-scheduled speculative decoding for production serving]() · [来源 2]() | A | 沿用既有阅读记录 | | 572 | [AWS Lambda MicroVMs:full-lifecycle Firecracker sandboxes for AI agents]() | A | 沿用既有阅读记录 | | 573 | [DBOSify:Postgres-backed durable workflow as compact Temporal alternative]() | A | 沿用既有阅读记录 | | 574 | [Adrafinil:macOS activity assertion layer for long-running AI agents]() | A | 沿用既有阅读记录 | | 575 | [Cloud World Model:cloud-infra simulation product radar]() | A | 仅元信息;原记录仅有聚合入口;未补齐具体原文,不作为已读材料。 | | 576 | [Are We Ready For An Agent-Native Memory System? Trustworthy Memory Search]() · [来源 2]() | A | 沿用既有阅读记录 | | 577 | [Semgrep GLM-5.2 cyber benchmark:real security harness for coding agents]() | A | 沿用既有阅读记录 | | 578 | [GitHub Copilot App BYOK:provider/account boundary for agent sessions]() | A | 沿用既有阅读记录 | | 579 | [Wayfinder Router:deterministic local/cloud LLM routing]() | A | 沿用既有阅读记录 | | 580 | [OpenAI Codex issue:exclude sensitive files by default]() | A | 沿用既有阅读记录 | | 581 | [Tokenmaxxing / agent output budget policy]() | A | 沿用既有阅读记录 | | 582 | [百度千帆 Coding Plan -> Token Plan:coding agent usage ledger signal]() | B | 沿用既有阅读记录 | | 583 | [Dual Strix Halo vLLM cluster:local serving lab reference]() | B | 沿用既有阅读记录 | | 584 | [Claude Sonnet 5:agent default model cost-performance reset]() | A | 沿用既有阅读记录 | | 585 | [Claude Science AI workbench:auditable domain agent OS]() | A | 沿用既有阅读记录 | | 586 | [Gemini Spark updates:desktop automation + connected apps + custom MCP]() | A | 沿用既有阅读记录 | | 587 | [vLLM Micro-Agent:serving router as bounded agent collaboration]() | A | 沿用既有阅读记录 | | 588 | [SWE-Together:interactive user-session benchmark for coding agents]() | A | 沿用既有阅读记录 | | 589 | [OSWorld 2.0:long-horizon computer-use official runner boundary]() | A | 沿用既有阅读记录 | | 590 | [SWE-MeM:adaptive memory management for long-horizon coding agents]() | A | 沿用既有阅读记录 | | 591 | [Language Firewall:routing defense for multi-agent systems]() | A | 沿用既有阅读记录 | | 592 | [Couchbase AI Data Plane:enterprise memory/context substrate]() | A | 沿用既有阅读记录 | | 593 | [Lingtai:local-first lifelong Agent runtime]() | A | 沿用既有阅读记录 | | 594 | [Ephemeral Sandbox:COW workspace 与 OCC publication]() | A | 沿用既有阅读记录 | | 595 | [Using Claude Code: The Unreasonable Effectiveness of HTML]() | A | 沿用既有阅读记录 | | 596 | [梁文锋投资者交流会 · 网传录音文字稿(非官方)]() · [来源 2]() | A | 仅元信息;已取得 42 页 PDF、提取文本并检查首页,未逐段精读。网传转写稿,未经 DeepSeek 或本人公开确认;ASR 与 AI 整理可能引入错误,不作为官方口径。 | | 597 | [hai-stack / Geju:用 target model、falsifier 与收益账单抵抗局部补丁化]() | A | 沿用既有阅读记录 | | 598 | [Waza:把工程习惯产品化为可路由、可验证的 Agent skills]() | A | 沿用既有阅读记录 | | 599 | [BfdCampos Mermaid skill:把“语法正确”提升为“渲染后可读”]() | A | 沿用既有阅读记录 | | 600 | [Oracle:把第二模型意见封装成可追踪的 consult session]() | A | 沿用既有阅读记录 | | 601 | [Graph Engineering:给 Agent Harness 叠加显式控制流图,而不是换一个新名词]() | A | 沿用既有阅读记录 | | 602 | [赵克常《炒股挣钱》:风险认知课,不是可复制的投资研究方法]() | B | 沿用既有阅读记录 | | 603 | [未核实的 MLSys 截图线索:集中式推理、确定性信号与知识图谱控制面]() | B | 沿用既有阅读记录;来源真实性未确认;保留为核验案例,不作为论文结论。 | | 604 | [Evolvent AI GitHub 组织复核:与既有研究目录的重复记录]() | B | 沿用既有阅读记录;保留稳定 ID 与重复记录;不因重复而静默删除。 |