{"id":"age-0001","section":"Start Here","subcategory":"Architecture guide","rtype":"Blog","title":"A practical guide to building agents","url":"https://openai.com/business/guides-and-resources/a-practical-guide-to-building-ai-agents/","venue":"OpenAI Guides & Resources","year":2025,"authors":"OpenAI","description":"Explains when agents are appropriate, how to define their tools and instructions, and how to move from one agent to manager and handoff-based multi-agent patterns.","why":"A useful first check on whether distinct roles justify the coordination cost of a graph.","evidence":"Practitioner analysis","layer":"Roles"} {"id":"age-0002","section":"Start Here","subcategory":"Architecture guide","rtype":"Blog","title":"Building effective agents","url":"https://www.anthropic.com/engineering/building-effective-agents","venue":"Anthropic Engineering","year":2024,"authors":"Erik Schluntz; Barry Zhang","description":"Distinguishes workflows from autonomous agents and presents routing, parallelization, orchestrator-worker, and evaluator-optimizer patterns.","why":"Provides a compact topology vocabulary and a strong case for adding coordination only when the task demands it.","evidence":"Practitioner analysis","layer":"Topology"} {"id":"age-0003","section":"Research Foundations","subcategory":"Agent foundations","rtype":"Paper","title":"Agent-Oriented Programming","url":"https://doi.org/10.1016/0004-3702(93)90034-9","venue":"Artificial Intelligence","year":1993,"authors":"Yoav Shoham","description":"Defines a programming paradigm in which agents are first-class components described through mental state and governed by explicit interaction rules.","why":"Establishes the intellectual lineage for treating an agent role as a programmable organizational unit.","evidence":"Peer-reviewed research","layer":"Roles"} {"id":"age-0004","section":"Research Foundations","subcategory":"Agent foundations","rtype":"Paper","title":"Intelligent Agents: Theory and Practice","url":"https://doi.org/10.1017/S0269888900008122","venue":"The Knowledge Engineering Review","year":1995,"authors":"Michael Wooldridge; Nicholas R. Jennings","description":"Surveys the properties, architectures, and engineering approaches that distinguish autonomous agents from ordinary software modules.","why":"Grounds the agency-at-the-nodes boundary that separates an agent graph from a deterministic workflow.","evidence":"Peer-reviewed research","layer":"Roles"} {"id":"age-0005","section":"Research Foundations","subcategory":"Shared-state architectures","rtype":"Paper","title":"The Blackboard Model of Problem Solving and the Evolution of Blackboard Architectures","url":"https://doi.org/10.1609/aimag.v7i2.537","venue":"AI Magazine","year":1986,"authors":"H. Penny Nii","description":"Describes systems in which independent specialists coordinate opportunistically through a shared problem state and a control component.","why":"Supplies a durable model for shared state without requiring every node to exchange its full context directly.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0006","section":"Research Foundations","subcategory":"Learned communication","rtype":"Paper","title":"Learning to Communicate with Deep Multi-Agent Reinforcement Learning","url":"https://proceedings.neurips.cc/paper_files/paper/2016/hash/c7635bfd99248a2cdef8249ef7bfbef4-Abstract.html","venue":"NeurIPS","year":2016,"authors":"Jakob Foerster; Ioannis Alexandros Assael; Nando de Freitas; Shimon Whiteson","description":"Introduces reinforcement-learning methods that let agents learn communication protocols alongside their task policies, including discrete messages for execution.","why":"Shows that edge content and communication policy can be engineered or learned rather than treated as free-form chat.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0007","section":"Research Foundations","subcategory":"Learned communication","rtype":"Paper","title":"TarMAC: Targeted Multi-Agent Communication","url":"https://proceedings.mlr.press/v97/das19a.html","venue":"ICML","year":2019,"authors":"Abhishek Das et al.","description":"Uses attention to let agents address different messages to selected recipients instead of broadcasting the same information to the whole team.","why":"Motivates selective, recipient-aware handoffs when all-to-all communication is wasteful or distracting.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0008","section":"Research Foundations","subcategory":"LLM multi-agent systems","rtype":"Paper","title":"AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations","url":"https://openreview.net/forum?id=BAakY1hNKS","venue":"COLM","year":2024,"authors":"Qingyun Wu; Gagan Bansal; Jieyu Zhang; Yiran Wu; Beibin Li; Erkang Zhu; Li Jiang; Xiaoyun Zhang; Shaokun Zhang; Jiale Liu; Ahmed Hassan Awadallah; Ryen W. White; Doug Burger; Chi Wang","description":"Presents a framework for composing customizable conversational agents that can combine language models, tools, code execution, and human input.","why":"An early, influential demonstration that agent roles and conversation links can be expressed as an executable topology.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0009","section":"Research Foundations","subcategory":"Topology optimization","rtype":"Paper","title":"GPTSwarm: Language Agents as Optimizable Graphs","url":"https://proceedings.mlr.press/v235/zhuge24a.html","venue":"ICML","year":2024,"authors":"Mingchen Zhuge et al.","description":"Represents language-agent systems as computational graphs and optimizes graph components from task feedback.","why":"Makes the graph itself an optimization target rather than a fixed orchestration diagram.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0010","section":"Research Foundations","subcategory":"Dynamic topology","rtype":"Paper","title":"A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration","url":"https://openreview.net/forum?id=XII0Wp1XA9","venue":"COLM","year":2024,"authors":"Zijun Liu et al.","description":"Constructs task-specific collaboration networks that can vary which agents participate and how they communicate instead of relying on one fixed team.","why":"Provides evidence for adapting the work graph to the task while keeping the available agent roles reusable.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0011","section":"Research Foundations","subcategory":"Communication efficiency","rtype":"Paper","title":"Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems","url":"https://openreview.net/forum?id=LkzuPorQ5L","venue":"ICLR","year":2025,"authors":"Guibin Zhang et al.","description":"Studies an economical communication pipeline that reduces redundant information exchanged among language-model agents.","why":"Treats edge traffic as a measurable cost and tests whether less communication can preserve useful collaboration.","evidence":"Peer-reviewed research","layer":"Observability & cost"} {"id":"age-0012","section":"Research Foundations","subcategory":"Topology optimization","rtype":"Paper","title":"G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks","url":"https://proceedings.mlr.press/v267/zhang25cu.html","venue":"ICML","year":2025,"authors":"Guibin Zhang et al.","description":"Uses graph neural networks to design communication structures for multi-agent systems instead of assuming a complete or manually chosen graph.","why":"Connects task performance to explicit topology search and exposes communication structure as an engineering variable.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0013","section":"Research Foundations","subcategory":"System search","rtype":"Paper","title":"Automated Design of Agentic Systems","url":"https://openreview.net/forum?id=t9U3LW7JVX","venue":"ICLR","year":2025,"authors":"Shengran Hu; Cong Lu; Jeff Clune","description":"Uses a meta-agent to propose, evaluate, and iteratively improve code-defined agentic systems across tasks.","why":"Demonstrates automated search over coordination logic while retaining executable artifacts that engineers can inspect.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0014","section":"Research Foundations","subcategory":"Workflow search","rtype":"Paper","title":"AFlow: Automating Agentic Workflow Generation","url":"https://openreview.net/forum?id=z5uVAKwmjf","venue":"ICLR","year":2025,"authors":"Jiayi Zhang et al.","description":"Searches over reusable workflow operators to generate task-specific agentic workflows and improve them from evaluation results.","why":"Offers a concrete method for evolving work graphs against measurable objectives rather than intuition alone.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0015","section":"Research Foundations","subcategory":"Joint optimization","rtype":"Paper","title":"Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies","url":"https://openreview.net/forum?id=I05H9RUzHB","venue":"ICLR","year":2026,"authors":"Han Zhou et al.","description":"Studies joint optimization of agent prompts and communication topology instead of tuning either component in isolation.","why":"Shows that node behavior and graph structure interact and may need to evolve together.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0016","section":"Research Foundations","subcategory":"Debate and councils","rtype":"Paper","title":"Improving Factuality and Reasoning in Language Models through Multiagent Debate","url":"https://proceedings.mlr.press/v235/du24e.html","venue":"ICML","year":2024,"authors":"Yilun Du et al.","description":"Tests rounds of proposal and critique among multiple language-model instances as a way to improve factual and reasoning answers.","why":"Supplies an empirical basis for debate-style gates while leaving room to examine correlated errors and added cost.","evidence":"Peer-reviewed research","layer":"Gates"} {"id":"age-0017","section":"Research Foundations","subcategory":"Debate and councils","rtype":"Paper","title":"Improving Multi-Agent Debate with Sparse Communication Topology","url":"https://aclanthology.org/2024.findings-emnlp.427/","venue":"Findings of EMNLP","year":2024,"authors":"Yunxuan Li et al.","description":"Examines multi-agent debate under sparse communication structures rather than defaulting to full information exchange among every participant.","why":"Isolates topology as a factor in debate quality and communication efficiency.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0018","section":"Research Foundations","subcategory":"Feedback and memory","rtype":"Paper","title":"Reflexion: Language Agents with Verbal Reinforcement Learning","url":"https://openreview.net/forum?id=vAElhFcKW6","venue":"NeurIPS","year":2023,"authors":"Noah Shinn et al.","description":"Lets an agent convert feedback into textual reflections stored in episodic memory and reused on later attempts.","why":"Clarifies how a node loop can persist learning across retries without changing model weights.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0019","section":"Research Foundations","subcategory":"Tool-grounded verification","rtype":"Paper","title":"CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing","url":"https://openreview.net/forum?id=Sx038qxjek","venue":"ICLR","year":2024,"authors":"Zhibin Gou et al.","description":"Uses external tools to obtain feedback that a language model can apply when critiquing and revising its outputs.","why":"Supports verification gates grounded in observable evidence instead of another ungrounded model opinion.","evidence":"Peer-reviewed research","layer":"Gates"} {"id":"age-0020","section":"Production Case Studies","subcategory":"Generalist agent team","rtype":"Paper","title":"Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks","url":"https://www.microsoft.com/en-us/research/publication/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/","venue":"Microsoft Research MSR-TR-2024-47","year":2024,"authors":"Adam Fourney et al.","description":"Documents an orchestrator-led team of specialized agents for web, file, coding, and terminal tasks, with progress tracking and replanning.","why":"Provides a concrete generalist topology whose role boundaries and recovery behavior can be examined as a system.","evidence":"Research preprint","layer":"Roles"} {"id":"age-0021","section":"Production Case Studies","subcategory":"Parallel research","rtype":"Blog","title":"How we built our multi-agent research system","url":"https://www.anthropic.com/engineering/multi-agent-research-system","venue":"Anthropic Engineering","year":2025,"authors":"Jeremy Hadfield et al.","description":"Describes a lead research agent that creates parallel subagents, delegates searches, and synthesizes their findings, including operational lessons from production.","why":"A detailed case study of dynamic fan-out and fan-in, context separation, evaluation, and the token cost of an adaptive work graph.","evidence":"Practitioner analysis","layer":"Work graphs"} {"id":"age-0022","section":"Frameworks & SDKs","subcategory":"Role-based teams","rtype":"Docs","title":"CrewAI Crews","url":"https://docs.crewai.com/en/concepts/crews","venue":"CrewAI Documentation","year":2026,"authors":"CrewAI","description":"Documents role-based agent crews, assigned tasks, delegation, and sequential or hierarchical execution processes.","why":"Offers accessible primitives for testing explicit ownership and manager-worker coordination.","evidence":"Official documentation","layer":"Roles"} {"id":"age-0023","section":"Frameworks & SDKs","subcategory":"Agent orchestration","rtype":"Docs","title":"OpenAI Agents SDK: Agent orchestration","url":"https://openai.github.io/openai-agents-python/multi_agent/","venue":"OpenAI Agents SDK Documentation","year":2026,"authors":"OpenAI","description":"Explains manager-style orchestration with agents exposed as tools and decentralized orchestration through handoffs.","why":"Makes the centralized-versus-decentralized topology choice explicit in a production SDK.","evidence":"Official documentation","layer":"Topology"} {"id":"age-0024","section":"Frameworks & SDKs","subcategory":"Pattern catalog","rtype":"Docs","title":"Strands Agents: Multi-Agent Patterns","url":"https://strandsagents.com/docs/user-guide/concepts/multi-agent/multi-agent-patterns/","venue":"Strands Agents Documentation","year":2026,"authors":"Strands Agents; AWS","description":"Documents several coordination shapes for composing agents, including supervisor, swarm, workflow, and graph-oriented patterns.","why":"Lets builders compare topology choices within one SDK instead of treating one pattern as universal.","evidence":"Official documentation","layer":"Topology"} {"id":"age-0025","section":"Frameworks & SDKs","subcategory":"Typed agents","rtype":"Docs","title":"Pydantic AI: Multi-Agent Applications","url":"https://pydantic.dev/docs/ai/guides/multi-agent-applications/","venue":"Pydantic AI Documentation","year":2026,"authors":"Pydantic","description":"Shows delegation, programmatic control flow, and graph-based state machines for composing typed Python agents.","why":"Useful for expressing node inputs, outputs, dependencies, and orchestration boundaries in ordinary application code.","evidence":"Official documentation","layer":"Topology"} {"id":"age-0026","section":"Frameworks & SDKs","subcategory":"Multi-agent orchestration","rtype":"Docs","title":"LlamaIndex: Multi-Agent Patterns","url":"https://developers.llamaindex.ai/python/framework/understanding/agent/multi_agent/","venue":"LlamaIndex Documentation","year":2026,"authors":"LlamaIndex","description":"Covers agent workflow, orchestrator, and planner-oriented approaches for coordinating specialized agents.","why":"Provides implementation patterns for choosing who owns delegation and how results return to the coordinating node.","evidence":"Official documentation","layer":"Handoffs"} {"id":"age-0027","section":"Frameworks & SDKs","subcategory":"Graph workflows","rtype":"Docs","title":"Google ADK: Graph-based Agent Workflows","url":"https://adk.dev/graphs/","venue":"Google ADK Documentation","year":2026,"authors":"Google","description":"Documents declarative workflows whose nodes combine agents, tools, functions, and human input through explicit edges, typed data passing, routing, branching, state, fan-out and join, loops, escalation, and nesting.","why":"Makes the graph load-bearing and inspectable while separating deterministic process control from model reasoning.","evidence":"Official documentation","layer":"Work graphs"} {"id":"age-0028","section":"Frameworks & SDKs","subcategory":"Graph workflows","rtype":"Docs","title":"Microsoft Agent Framework: Workflows","url":"https://learn.microsoft.com/en-us/agent-framework/workflows/","venue":"Microsoft Learn","year":2026,"authors":"Microsoft","description":"Describes workflows built from executors and explicit edges, with support for branching, aggregation, state, and checkpointing.","why":"Exposes the work graph as an inspectable program rather than hiding coordination inside prompts.","evidence":"Official documentation","layer":"Work graphs"} {"id":"age-0029","section":"Frameworks & SDKs","subcategory":"Graph runtime","rtype":"Docs","title":"LangGraph overview","url":"https://docs.langchain.com/oss/python/langgraph/overview","venue":"LangGraph Documentation","year":2025,"authors":"LangChain","description":"Introduces a low-level runtime for stateful agent graphs with durable execution, streaming, memory, and human intervention.","why":"A widely used substrate for implementing explicit nodes, edges, state transitions, and resumable work graphs.","evidence":"Official documentation","layer":"Work graphs"} {"id":"age-0030","section":"Protocols & Handoffs","subcategory":"In-process transfer","rtype":"Docs","title":"OpenAI Agents SDK: Handoffs","url":"https://openai.github.io/openai-agents-python/handoffs/","venue":"OpenAI Agents SDK Documentation","year":2026,"authors":"OpenAI","description":"Documents transfers from one agent to another, including tool-shaped handoff schemas, input filters, and callbacks.","why":"Turns an edge into an explicit contract controlling when ownership moves and what context crosses with it.","evidence":"Official documentation","layer":"Handoffs"} {"id":"age-0031","section":"Protocols & Handoffs","subcategory":"Tool and context protocol","rtype":"Standard","title":"Model Context Protocol Specification 2025-11-25","url":"https://modelcontextprotocol.io/specification/2025-11-25","venue":"MCP Specification 2025-11-25","year":2025,"authors":"Model Context Protocol project; Agentic AI Foundation","description":"Specifies a client-server protocol through which AI applications discover and use tools, resources, prompts, and contextual data.","why":"Standardizes capability and context edges, while remaining distinct from a protocol for delegating work between autonomous agents.","evidence":"Industry standard","layer":"Handoffs"} {"id":"age-0032","section":"Protocols & Handoffs","subcategory":"Agent interoperability","rtype":"Standard","title":"Agent2Agent Protocol Specification v1.0.0","url":"https://a2a-protocol.org/v1.0.0/specification/","venue":"A2A Specification","year":2026,"authors":"A2A Protocol Working Group; Linux Foundation","description":"Defines interoperable agent discovery, task lifecycle, messages, artifacts, streaming, asynchronous updates, version negotiation, and multiple protocol bindings across service boundaries.","why":"Provides a version-pinned wire contract for cross-system handoffs where agents cannot share an in-process runtime.","evidence":"Industry standard","layer":"Handoffs"} {"id":"age-0033","section":"State, Memory & Artifacts","subcategory":"Checkpointing","rtype":"Docs","title":"LangGraph Persistence","url":"https://docs.langchain.com/oss/python/langgraph/persistence","venue":"LangGraph Documentation","year":2025,"authors":"LangChain","description":"Documents thread-scoped checkpoints, saved state, replay, state inspection, and memory storage for LangGraph runs.","why":"Shows how graph state can become the recoverable system of record instead of living only in model context.","evidence":"Official documentation","layer":"State"} {"id":"age-0034","section":"State, Memory & Artifacts","subcategory":"Conversation state","rtype":"Docs","title":"OpenAI Agents SDK: Sessions","url":"https://openai.github.io/openai-agents-python/sessions/","venue":"OpenAI Agents SDK Documentation","year":2026,"authors":"OpenAI","description":"Documents persistent conversation history that can be loaded and updated across repeated agent runs.","why":"Provides a bounded mechanism for carrying state across nodes and turns without manually rebuilding every prompt.","evidence":"Official documentation","layer":"State"} {"id":"age-0035","section":"Verification & Evals","subcategory":"Workflow evaluation","rtype":"Docs","title":"Evaluate agent workflows","url":"https://developers.openai.com/api/docs/guides/agent-evals","venue":"OpenAI API Documentation","year":2026,"authors":"OpenAI","description":"Explains how to build datasets, graders, trace-based evaluations, and reproducible checks for agent workflows.","why":"Makes gates testable at both the final outcome and the intermediate handoff level.","evidence":"Official documentation","layer":"Gates"} {"id":"age-0036","section":"Verification & Evals","subcategory":"Evaluation practice","rtype":"Blog","title":"Demystifying evals for AI agents","url":"https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents","venue":"Anthropic Engineering","year":2026,"authors":"Mikaela Grace et al.","description":"Presents practical guidance for defining tasks, outcomes, graders, and evaluation suites for agents with non-deterministic trajectories.","why":"Helps turn vague reviewer judgments into evidence gates that can guide graph changes.","evidence":"Practitioner analysis","layer":"Gates"} {"id":"age-0037","section":"Verification & Evals","subcategory":"Evaluation framework","rtype":"Tool","title":"Inspect AI","url":"https://inspect.aisi.org.uk/","venue":"Inspect AI Documentation","year":2024,"authors":"UK AI Security Institute","description":"Provides an open-source evaluation framework with tasks, solvers, scorers, sandboxed tools, and structured logs.","why":"Supports reproducible gate nodes and traceable evidence for both individual agents and composed systems.","evidence":"Maintained OSS project","layer":"Gates"} {"id":"age-0038","section":"Reliability & Durable Execution","subcategory":"Durable actors and workflows","rtype":"Docs","title":"Dapr Agents introduction","url":"https://docs.dapr.io/developing-ai/dapr-agents/dapr-agents-introduction/","venue":"Dapr Documentation","year":2026,"authors":"Dapr; CNCF","description":"Introduces an agent framework built on durable actors and workflows, with state, messaging, and recovery supplied by the Dapr runtime.","why":"Shows how graph nodes can inherit distributed-systems durability instead of implementing recovery only in prompts.","evidence":"Official documentation","layer":"Reliability"} {"id":"age-0039","section":"Reliability & Durable Execution","subcategory":"Durable execution integration","rtype":"Tool","title":"Temporal integration for OpenAI Agents SDK","url":"https://github.com/temporalio/sdk-python/tree/main/temporalio/contrib/openai_agents","venue":"Temporal Python SDK","year":2025,"authors":"Temporal","description":"Integrates OpenAI agent runs, model calls, and tools with Temporal workflows and activities for durable execution.","why":"Provides replay, retry, timeout, and recovery semantics beneath an agent graph without asking the model to manage them.","evidence":"Maintained OSS project","layer":"Reliability"} {"id":"age-0040","section":"Reliability & Durable Execution","subcategory":"Database-backed workflows","rtype":"Docs","title":"DBOS AI Quickstart","url":"https://docs.dbos.dev/ai/ai-quickstart","venue":"DBOS Documentation","year":2026,"authors":"DBOS","description":"Shows how to place AI application steps inside durable workflows whose progress is recorded and recoverable after interruption.","why":"Offers a compact path from an agent prototype to resumable execution with explicit step boundaries.","evidence":"Official documentation","layer":"Reliability"} {"id":"age-0041","section":"Reliability & Durable Execution","subcategory":"Durable agent patterns","rtype":"Docs","title":"Restate Durable Agents","url":"https://docs.restate.dev/ai/patterns/durable-agents","venue":"Restate Documentation","year":2026,"authors":"Restate","description":"Documents durable agent patterns using persisted execution, reliable calls, retries, timers, and stateful services.","why":"Maps common mid-graph failures to runtime guarantees rather than fragile application-level retry code.","evidence":"Official documentation","layer":"Reliability"} {"id":"age-0042","section":"Observability & Cost","subcategory":"Tracing","rtype":"Docs","title":"OpenAI Agents SDK: Tracing","url":"https://openai.github.io/openai-agents-python/tracing/","venue":"OpenAI Agents SDK Documentation","year":2026,"authors":"OpenAI","description":"Documents traces and spans for agent runs, model generations, tool calls, handoffs, and guardrails.","why":"Makes node and edge behavior inspectable so latency, failures, and expensive paths can be attributed correctly.","evidence":"Official documentation","layer":"Observability & cost"} {"id":"age-0043","section":"Observability & Cost","subcategory":"Telemetry standard","rtype":"Standard","title":"OpenTelemetry GenAI Semantic Conventions","url":"https://github.com/open-telemetry/semantic-conventions-genai","venue":"OpenTelemetry GenAI repository","year":2026,"authors":"OpenTelemetry GenAI SIG","description":"Develops shared telemetry names and attributes for generative-AI model, tool, and agent operations in traces, metrics, and events.","why":"Helps graph telemetry remain portable across runtimes and observability vendors.","evidence":"Maintained OSS project","layer":"Observability & cost"} {"id":"age-0044","section":"Observability & Cost","subcategory":"Cost attribution","rtype":"Docs","title":"Arize Phoenix: Cost Tracking","url":"https://arize.com/docs/phoenix/tracing/how-to-tracing/cost-tracking","venue":"Arize Phoenix Documentation","year":2026,"authors":"Arize AI","description":"Explains how Phoenix derives and displays token usage and model cost from traced generative-AI calls.","why":"Supports per-node and per-path cost accounting instead of treating a multi-agent run as one opaque bill.","evidence":"Official documentation","layer":"Observability & cost"} {"id":"age-0045","section":"Observability & Cost","subcategory":"Observability platform","rtype":"Docs","title":"LangSmith Observability","url":"https://docs.langchain.com/langsmith/observability","venue":"LangSmith Documentation","year":2026,"authors":"LangChain","description":"Documents tracing, dashboards, alerts, feedback, and evaluation views for language-model and agent applications.","why":"Provides operational views for following execution across nodes and locating failures on the critical path.","evidence":"Official documentation","layer":"Observability & cost"} {"id":"age-0046","section":"Benchmarks & Datasets","subcategory":"Collaboration benchmark","rtype":"Benchmark","title":"MultiAgentBench: Evaluating Collaboration and Competition of LLM Agents","url":"https://aclanthology.org/2025.acl-long.421/","venue":"ACL","year":2025,"authors":"Kunlun Zhu et al.","description":"Benchmarks language-model agents in collaborative and competitive settings while examining coordination processes as well as task outcomes.","why":"Measures properties of the team interaction that single-agent benchmarks cannot expose.","evidence":"Benchmark/dataset","layer":"Observability & cost"} {"id":"age-0047","section":"Benchmarks & Datasets","subcategory":"Failure diagnosis","rtype":"Benchmark","title":"Why Do Multi-Agent LLM Systems Fail?","url":"https://nips.cc/virtual/2025/poster/121528","venue":"NeurIPS Datasets & Benchmarks","year":2025,"authors":"Mert Cemri et al.","description":"Provides a structured taxonomy and evaluation approach for diagnosing coordination failures in multi-agent language-model systems.","why":"Turns reliability incidents into recurring, attributable failure classes that can guide graph redesign.","evidence":"Benchmark/dataset","layer":"Reliability"} {"id":"age-0048","section":"Critiques & Limits","subcategory":"Self-correction limits","rtype":"Paper","title":"Large Language Models Cannot Self-Correct Reasoning Yet","url":"https://openreview.net/forum?id=IkmD3fKBPQ","venue":"ICLR","year":2024,"authors":"Jie Huang et al.","description":"Finds that intrinsic self-correction without reliable external feedback often fails to improve reasoning and can reduce accuracy.","why":"Warns against using an ungrounded critic node as evidence simply because it is separate from the producing node.","evidence":"Peer-reviewed research","layer":"Gates"} {"id":"age-0049","section":"Critiques & Limits","subcategory":"Scaling evidence","rtype":"Paper","title":"Towards a Science of Scaling Agent Systems","url":"https://arxiv.org/abs/2512.08296","venue":"arXiv; Google Research","year":2025,"authors":"Yubin Kim et al.","description":"Studies how agent-system performance changes across tasks, models, coordination structures, and scaling choices under controlled experiments.","why":"Tests the assumption that adding agents reliably helps and frames scaling as an empirical topology decision.","evidence":"Research preprint","layer":"Topology"} {"id":"age-0050","section":"Critiques & Limits","subcategory":"Token-budget comparison","rtype":"Paper","title":"Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets","url":"https://arxiv.org/abs/2604.02460","venue":"arXiv","year":2026,"authors":"Dat Tran; Douwe Kiela","description":"Compares single-agent and multi-agent approaches to multi-hop reasoning while holding the total thinking-token budget constant.","why":"Provides the cost-controlled baseline needed before claiming that coordination, rather than extra inference, caused an improvement.","evidence":"Research preprint","layer":"Topology"} {"id":"age-0051","section":"Start Here","subcategory":"Contemporary framing","rtype":"Blog","title":"From Loop Engineering to Graph Engineering?","url":"https://x.com/IntuitMachine/status/2078419526354378975","venue":"X Articles","year":2026,"authors":"Carlos E. Perez","description":"Frames the shift as loop architecture: networks of improvement cycles that monitor, feed, constrain, and correct one another, with reliability located in their edges.","why":"Adds the grounding requirement missing from topology-only accounts: independent counter-metrics, frozen tests or rules, external anchors, and human ownership of root objectives.","evidence":"Practitioner analysis","layer":"Gates"} {"id":"age-0052","section":"Research Foundations","subcategory":"Blackboard coordination","rtype":"Paper","title":"A Multi-Level Organization for Problem Solving Using Many, Diverse, Cooperating Sources of Knowledge","url":"https://www.ijcai.org/Proceedings/75/Papers/072.pdf","venue":"IJCAI","year":1975,"authors":"Lee D. Erman; Victor R. Lesser","description":"Presents the Hearsay-II multi-level blackboard, where independent knowledge sources react to shared hypotheses, create explicit structural dependencies, and verify or revise one another's contributions.","why":"Provides an early architecture for loosely coupled specialist nodes coordinating through inspectable shared state instead of direct all-to-all calls.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0053","section":"Research Foundations","subcategory":"Negotiated delegation","rtype":"Paper","title":"The Contract Net Protocol: High-Level Communication and Control in a Distributed Problem Solver","url":"https://doi.org/10.1109/TC.1980.1675516","venue":"IEEE Transactions on Computers","year":1980,"authors":"Reid G. Smith","description":"Defines a negotiation protocol in which managers announce tasks, potential contractors bid, and awards establish temporary problem-solving relationships.","why":"Supplies the classic contract for capability-aware delegation and auditable assignment edges between autonomous nodes.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0054","section":"Start Here","subcategory":"Communication survey","rtype":"Paper","title":"The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why – A Survey from MARL to Emergent Language and LLMs","url":"https://openreview.net/pdf?id=LGsed0QQVq","venue":"Transactions on Machine Learning Research","year":2026,"authors":"Jingdi Chen; Hanqing Yang; Zongjun Liu; Carlee Joe-Wong","description":"Synthesizes multi-agent communication across reinforcement learning, emergent language, and LLM systems through sender, recipient, timing, content, and purpose decisions.","why":"Provides a design-oriented map for engineering edge selection, message timing, payloads, grounding, scalability, and interpretability.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0055","section":"Research Foundations","subcategory":"Collaboration scaling","rtype":"Paper","title":"Scaling Large Language Model-based Multi-Agent Collaboration","url":"https://openreview.net/forum?id=K3n5jPkrU6","venue":"ICLR","year":2025,"authors":"Chen Qian; Zihao Xie; YiFei Wang; Wei Liu; Kunlun Zhu; Hanchen Xia; Yufan Dang; Zhuoyun Du; Weize Chen; Cheng Yang; Zhiyuan Liu; Maosong Sun","description":"Introduces MacNet, a DAG-based collaboration architecture executed in topological order, and studies communication structure while scaling experiments beyond 1,000 agents.","why":"Makes topology a causal scaling variable rather than assuming that larger teams or denser communication automatically help.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0056","section":"Research Foundations","subcategory":"Holistic orchestration","rtype":"Paper","title":"MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks","url":"https://openreview.net/forum?id=3fGXBm4c3S","venue":"ICML","year":2026,"authors":"Zixuan Ke et al.","description":"Formulates orchestration as reinforcement-learned generation of a complete multi-agent program and evaluates it across controlled dimensions including depth, horizon, breadth, parallelism, and robustness.","why":"Tests when whole-system graph structure helps instead of attributing gains to coordination without controlled task evidence.","evidence":"Peer-reviewed research","layer":"Work graphs"} {"id":"age-0057","section":"Research Foundations","subcategory":"Conditional topology","rtype":"Paper","title":"CARD: Towards Conditional Design of Multi-agent Topological Structures","url":"https://openreview.net/forum?id=JgvJdICc6P","venue":"ICLR","year":2026,"authors":"Tongtong Wu et al.","description":"Generates communication graphs conditioned on agent roles, models, tools, and data sources, and adapts topology as task resources change.","why":"Treats the available capabilities and directed edges as a versionable organizational artifact rather than a fixed team template.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0058","section":"Reliability & Durable Execution","subcategory":"Resilient topology","rtype":"Paper","title":"ResMAS: Resilience Optimization in LLM-based Multi-agent Systems","url":"https://ojs.aaai.org/index.php/AAAI/article/view/40824","venue":"AAAI","year":2026,"authors":"Zhilun Zhou et al.","description":"Learns task-specific resilient communication topologies and topology-aware prompts after measuring how graph structure and node instructions affect performance under agent failures and other perturbations.","why":"Moves resilience from reactive recovery into the design of the graph itself and evaluates transfer to new tasks and models.","evidence":"Peer-reviewed research","layer":"Reliability"} {"id":"age-0059","section":"Research Foundations","subcategory":"System evolution","rtype":"Paper","title":"EvoMAS: Evolutionary Generation of Multi-Agent Systems","url":"https://openreview.net/forum?id=ic0AGRIkmY","venue":"ICML","year":2026,"authors":"Yuntong Hu et al.","description":"Evolves structured multi-agent configurations through trace-guided mutation, crossover, selection, and an experience memory across reasoning, coding, and tool-use tasks.","why":"Shows how an inspectable team specification can evolve from execution evidence while retaining executability and runtime robustness.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0060","section":"Critiques & Limits","subcategory":"Error propagation","rtype":"Paper","title":"Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems","url":"https://aclanthology.org/2025.emnlp-main.623/","venue":"EMNLP","year":2025,"authors":"Xu Shen et al.","description":"Causally studies correct and erroneous information propagation across communication densities and finds that moderately sparse structures can preserve useful diffusion while suppressing errors.","why":"Provides evidence against defaulting to dense graphs and links topology decisions to measured error amplification.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0061","section":"Frameworks & SDKs","subcategory":"Graph-centric orchestration","rtype":"Paper","title":"MASFactory: A Graph-centric Framework for Orchestrating LLM-Based Multi-Agent Systems with Vibe Graphing","url":"https://aclanthology.org/2026.acl-demo.35/","venue":"ACL System Demonstrations","year":2026,"authors":"Yang Liu et al.","description":"Compiles natural-language intent into an editable workflow specification and executable directed graph, with reusable components, topology preview, runtime tracing, multimodal messages, and human interaction.","why":"Provides a direct implementation path from an inspectable organizational graph to execution and evaluates it on seven public benchmarks.","evidence":"Peer-reviewed research","layer":"Work graphs"} {"id":"age-0062","section":"Frameworks & SDKs","subcategory":"Directed agent graphs","rtype":"Docs","title":"AutoGen GraphFlow (Workflows)","url":"https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/graph-flow.html","venue":"Microsoft AutoGen Documentation","year":2026,"authors":"Microsoft","description":"Implements directed multi-agent execution graphs with sequential, parallel, conditional, fan-in, and cyclic paths, edge conditions, activation groups, safe loop exits, and separately configurable message filtering. The feature is explicitly experimental.","why":"Distinguishes the execution graph from the message graph, exposing both who acts next and what context each agent receives.","evidence":"Official documentation","layer":"Work graphs"} {"id":"age-0063","section":"Protocols & Handoffs","subcategory":"Durable remote tasks","rtype":"Standard","title":"Model Context Protocol: Tasks","url":"https://modelcontextprotocol.io/specification/2025-11-25/basic/utilities/tasks","venue":"MCP Specification 2025-11-25","year":2025,"authors":"Model Context Protocol project; Agentic AI Foundation","description":"Specifies experimental durable asynchronous request state machines with capability negotiation, polling, deferred results, progress, input-required states, cancellation, TTLs, and task-message correlation.","why":"Turns a remote tool or context edge into a recoverable task contract that can outlive one synchronous request.","evidence":"Official documentation","layer":"Reliability"} {"id":"age-0064","section":"Protocols & Handoffs","subcategory":"Secure agent messaging","rtype":"Docs","title":"Secure Low-Latency Interactive Messaging (SLIM)","url":"https://datatracker.ietf.org/doc/draft-mpsb-agntcy-slim/","venue":"IETF Datatracker; individual Internet-Draft","year":2026,"authors":"Luca Muscariello; Michele Papalini; Mauro Sardara; Sam Betts","description":"Proposes a transport layer for A2A and MCP using gRPC over HTTP/2 and HTTP/3 with stream multiplexing, flow control, group communication, native RPC semantics, and MLS end-to-end encryption. It is an individual informational Internet-Draft with no formal IETF standing.","why":"Adds a concrete secure transport substrate for high-volume graph edges that must cross process and organizational boundaries.","evidence":"Official documentation","layer":"Handoffs"} {"id":"age-0065","section":"Protocols & Handoffs","subcategory":"Agent identity and authorization","rtype":"Docs","title":"Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization","url":"https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd","venue":"NIST NCCoE Initial Public Draft","year":2026,"authors":"Harold Booth; William Fisher; Ryan Galluzzo; Joshua Roberts","description":"Outlines considerations and open questions for standards-based identity, authorization, auditing, and non-repudiation when software and AI agents access enterprise systems and take actions. The concept paper remains an initial public draft under review.","why":"Grounds graph roles and permissions in identity practice so delegation does not silently transfer more authority than an edge contract allows.","evidence":"Official documentation","layer":"Roles"} {"id":"age-0066","section":"State, Memory & Artifacts","subcategory":"Versioned work products","rtype":"Docs","title":"Google ADK: Artifacts","url":"https://adk.dev/artifacts/","venue":"Google ADK Documentation","year":2026,"authors":"Google","description":"Defines named, automatically versioned binary work products that agents and tools can save, load, list, and exchange within session-scoped or persistent user-scoped namespaces.","why":"Provides explicit, inspectable edge artifacts instead of forcing large or structured outputs through conversational context.","evidence":"Official documentation","layer":"State"} {"id":"age-0067","section":"State, Memory & Artifacts","subcategory":"Procedural memory","rtype":"Paper","title":"LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation","url":"https://www.microsoft.com/en-us/research/publication/legomem-modular-procedural-memory-for-multi-agent-llm-systems-for-workflow-automation/","venue":"AAMAS","year":2026,"authors":"Dongge Han; Camille Couturier; Daniel Madrigal; Xuchao Zhang; Victor Ruehle; Saravan Rajmohan","description":"Decomposes execution trajectories into reusable procedural memories and allocates them to an orchestrator and specialist agents for workflow automation.","why":"Shows how durable experience can improve decomposition and delegation without collapsing every node's memory into one undifferentiated store.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0068","section":"Verification & Evals","subcategory":"Adversarial agent detection","rtype":"Paper","title":"When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems","url":"https://openreview.net/forum?id=BnduUW8izq","venue":"ICML","year":2026,"authors":"Haowen Xu et al.","description":"Evaluates activation-space detection and restorative steering of compromised agents across five attack scenarios, multiple models, and synchronous and asynchronous multi-agent interactions.","why":"Tests a topology-agnostic gate for locating and repairing a malicious node when its messages appear superficially benign.","evidence":"Peer-reviewed research","layer":"Reliability"} {"id":"age-0069","section":"Reliability & Durable Execution","subcategory":"Intervention-driven debugging","rtype":"Paper","title":"DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems","url":"https://iclr.cc/virtual/2026/poster/10007537","venue":"ICLR","year":2026,"authors":"Ming Ma et al.","description":"Tests failure hypotheses by editing messages or plans and measuring whether each intervention repairs the outcome or advances execution in branching multi-agent traces.","why":"Turns graph debugging into causal repair experiments instead of relying on plausible but unverified post-hoc explanations.","evidence":"Peer-reviewed research","layer":"Reliability"} {"id":"age-0070","section":"Reliability & Durable Execution","subcategory":"Durable graph workflows","rtype":"Docs","title":"Microsoft Agent Framework: Durable Extension","url":"https://learn.microsoft.com/en-us/agent-framework/integrations/durable-extension","venue":"Microsoft Learn","year":2026,"authors":"Microsoft","description":"Adds persisted sessions, checkpointed progress, failure recovery, external-event waits, and distributed hosting to agents, multi-agent orchestrations, and graph workflows. The documented packages remain prerelease.","why":"Shows how an agent graph can resume without losing context or repeating completed work after interruption.","evidence":"Official documentation","layer":"Reliability"} {"id":"age-0071","section":"Critiques & Limits","subcategory":"Control-flow security","rtype":"Paper","title":"Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems","url":"https://openreview.net/forum?id=PNU9Rj5RDQ","venue":"ICLR","year":2026,"authors":"Rishi Dev Jha; Harold Triedman; Justin Wagle; Vitaly Shmatikov","description":"Demonstrates attacks against alignment-check defenses and introduces ControlValve, which generates permitted control-flow graphs and enforces least privilege for each agent invocation.","why":"Makes invocation authority and contextual permissions explicit on every edge instead of trusting a separate checker that can also be hijacked.","evidence":"Peer-reviewed research","layer":"Reliability"} {"id":"age-0072","section":"Observability & Cost","subcategory":"Budget-aware topology","rtype":"Paper","title":"BAMAS: Structuring Budget-Aware Multi-Agent Systems","url":"https://ojs.aaai.org/index.php/AAAI/article/view/40226","venue":"AAAI","year":2026,"authors":"Liming Yang; Junyu Luo; Xuanzhe Liu; Yiling Lou; Zhenpeng Chen","description":"Selects an LLM team with integer programming and then learns its collaboration topology under an explicit budget, reporting comparable performance with cost reductions of up to 86% on three tasks.","why":"Jointly engineers node selection, edge structure, and spend instead of optimizing quality while treating inference cost as an afterthought.","evidence":"Peer-reviewed research","layer":"Observability & cost"} {"id":"age-0073","section":"Observability & Cost","subcategory":"Workflow reconstruction","rtype":"Paper","title":"AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction","url":"https://arxiv.org/abs/2602.05353","venue":"ICML","year":2026,"authors":"Ruijie Shi; Houbin Zhang; Yuecheng Han; Yuheng Wang; Jingru Fan; Runde Yang; Yufan Dang; Huatao Li; Dewen Liu; Yuan Cheng; Chen Qian","description":"Reconstructs an editable explicit stand-in workflow for a black-box agentic system from input-output behavior using iterative search and evaluation.","why":"Offers a path to inspect, compare, and modify systems whose load-bearing workflow is hidden behind an API.","evidence":"Peer-reviewed research","layer":"Observability & cost"} {"id":"age-0074","section":"Benchmarks & Datasets","subcategory":"Distributed coordination","rtype":"Benchmark","title":"SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems","url":"https://aclanthology.org/2026.acl-long.1354/","venue":"ACL","year":2026,"authors":"Yuzhe Zhang et al.","description":"Benchmarks free-form coordination under information silos across 30 exact-answer tasks, three communication protocols, six agent scales, and three models while recording success, tokens, and communication density.","why":"Tests whether more nodes and denser edges overcome distributed information constraints under explicit coordination and cost measures.","evidence":"Benchmark/dataset","layer":"Topology"} {"id":"age-0075","section":"Benchmarks & Datasets","subcategory":"Dynamic asynchronous evaluation","rtype":"Benchmark","title":"Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments","url":"https://openreview.net/forum?id=9gw03JpKK4","venue":"ICLR","year":2026,"authors":"Romain Froger et al.","description":"Evaluates agents in asynchronous simulated environments across execution, search, ambiguity, adaptation, temporal reasoning, noise, and agent-to-agent collaboration, with action-level verifiers and structured traces.","why":"Makes delays, environmental change, peer communication, and externally checked writes first-class work-graph conditions.","evidence":"Benchmark/dataset","layer":"Work graphs"} {"id":"age-0076","section":"Benchmarks & Datasets","subcategory":"Adversarial multi-agent safety","rtype":"Benchmark","title":"TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems","url":"https://aclanthology.org/2026.acl-long.1442/","venue":"ACL","year":2026,"authors":"Ishan Kavathekar; Hemang Jain; Ameya Rathod; Ponnurangam Kumaraguru; Tanuja Ganu","description":"Provides five scenarios with 300 adversarial instances, six attack types, 211 tools, 100 harmless tasks, and multiple AutoGen and CrewAI interaction configurations, plus an Effective Robustness Score.","why":"Evaluates whether safety controls preserve useful work while attacks exploit the distinctive trust and communication surfaces of a multi-agent graph.","evidence":"Benchmark/dataset","layer":"Reliability"} {"id":"age-0077","section":"Benchmarks & Datasets","subcategory":"Failure attribution","rtype":"Benchmark","title":"Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems","url":"https://aclanthology.org/2026.acl-long.912/","venue":"ACL","year":2026,"authors":"Mengzhuo Chen et al.","description":"Introduces TraceElephant with full execution traces and reproducible environments for attributing failures to responsible agents and decisive steps.","why":"Measures causal diagnosis over nodes, messages, plans, and dependencies instead of evaluating only the final team output.","evidence":"Benchmark/dataset","layer":"Observability & cost"} {"id":"age-0078","section":"Production Case Studies","subcategory":"Secure delegated access","rtype":"Blog","title":"Creating AI agent solutions for warehouse data access and security","url":"https://engineering.fb.com/2025/08/13/data-infrastructure/agentic-solution-for-warehouse-data-access/","venue":"Engineering at Meta","year":2025,"authors":"Can Lin; Uday Ramesh Savagaonkar; Iuliu Rus; Komal Mangtani","description":"Describes collaborating data-user and data-owner agents, specialized subagents, triage, permission negotiation, human oversight, access budgets, analytical risk rules, output guardrails, traces, and daily regression evaluation.","why":"Shows plural bounded agency governed by independent rule-based gates rather than relying on model judgment for security decisions.","evidence":"Practitioner analysis","layer":"Gates"} {"id":"age-0079","section":"Production Case Studies","subcategory":"Staged agent swarm","rtype":"Blog","title":"How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines","url":"https://engineering.fb.com/2026/04/06/developer-tools/how-meta-used-ai-to-map-tribal-knowledge-in-large-scale-data-pipelines/","venue":"Engineering at Meta","year":2026,"authors":"Krishna Ganeriwal; Plawan Rath; Ashwini Verma","description":"Reports a staged swarm of more than 50 explorer, analyst, writer, critic, fixer, upgrader, tester, and gap-filling tasks, with repeated review rounds and recurring automated refresh runs.","why":"Provides a concrete fan-out, critique, repair, integration, and recurrence graph with explicit artifacts and quality gates; efficiency figures are company-reported and preliminary.","evidence":"Practitioner analysis","layer":"Work graphs"} {"id":"age-0080","section":"Production Case Studies","subcategory":"Multi-hop identity and provenance","rtype":"Blog","title":"Solving the Identity Crisis for AI Agents","url":"https://www.uber.com/by/en/blog/solving-the-agent-identity-crisis/","venue":"Uber Engineering","year":2026,"authors":"Matt Mathew; Prasad Borole; Meng Huang; Sergey Burykin; Gaurav Goel; Bayard Walsh","description":"Describes an internal agent mesh with registered identities, SPIRE-backed workload attestation, short-lived audience-scoped tokens for every hop, actor-chain provenance, MCP gateway enforcement, and a standardized A2A client.","why":"Treats identity, delegated authority, and provenance as mandatory edge state across a multi-agent graph; adoption and latency metrics are company-reported.","evidence":"Practitioner analysis","layer":"Handoffs"} {"id":"age-0081","section":"Research Foundations","subcategory":"Task-adaptive topology synthesis","rtype":"Paper","title":"Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models","url":"https://aclanthology.org/2026.acl-long.1764/","venue":"ACL","year":2026,"authors":"Eric Hanchen Jiang et al.","description":"Introduces Guided Topology Diffusion, which iteratively synthesizes sparse, task-adaptive communication graphs using a proxy model for objectives such as accuracy, utility, and cost.","why":"Treats topology as a multi-objective, per-task design artifact rather than a static collaboration template.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0082","section":"Reliability & Durable Execution","subcategory":"Topology-conditioned memory leakage","rtype":"Paper","title":"Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs","url":"https://aclanthology.org/2026.findings-acl.1980/","venue":"Findings of ACL","year":2026,"authors":"Jinbo Liu et al.","description":"Measures private-information leakage across six communication topologies, agent counts, and attacker-target placements, finding higher leakage with denser connectivity, shorter graph distance, and greater target centrality.","why":"Makes privacy a measurable consequence of edge structure and node placement, motivating topology-aware access controls.","evidence":"Peer-reviewed research","layer":"Reliability"} {"id":"age-0083","section":"Critiques & Limits","subcategory":"Collaboration-induced diversity collapse","rtype":"Paper","title":"Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation","url":"https://aclanthology.org/2026.findings-acl.13/","venue":"Findings of ACL","year":2026,"authors":"Nuo Chen et al.","description":"Studies multi-agent ideation across model, cognition, and system levels, finding diminishing group-size returns, authority effects, and faster premature convergence under dense communication.","why":"Shows that interaction can contract a search space, so graph designs for exploration must preserve independence and meaningful disagreement.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0084","section":"Research Foundations","subcategory":"Dynamic node and edge elimination","rtype":"Paper","title":"AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration","url":"https://aclanthology.org/2025.acl-long.1170/","venue":"ACL","year":2025,"authors":"Zhexuan Wang; Yutong Wang; Xuebo Liu; Liang Ding; Miao Zhang; Jie Liu; Min Zhang","description":"Optimizes communication-graph adjacency matrices across collaboration rounds to remove redundant agents and messages, reducing both prompt and completion token consumption in its evaluations.","why":"Makes node participation and edge density explicit cost-quality decisions instead of assuming every role should run on every task.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0085","section":"Benchmarks & Datasets","subcategory":"Process-level collaboration","rtype":"Benchmark","title":"Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents","url":"https://aclanthology.org/2025.emnlp-main.249/","venue":"EMNLP","year":2025,"authors":"Haochen Sun; Shuwen Zhang; Lujie Niu; Lei Ren; Hao Xu; Hao Fu; Fangkun Zhao; Caixia Yuan; Xiaojie Wang","description":"Provides 30 open-ended cooperative tasks and process-oriented measures for goal interpretation, active collaboration, communication, and continuous adaptation across 13 language models.","why":"Distinguishes reaching an answer from coordinating well, exposing collaboration failures that final-outcome metrics can hide.","evidence":"Benchmark/dataset","layer":"Observability & cost"} {"id":"age-0086","section":"Critiques & Limits","subcategory":"Inter-agent communication attacks","rtype":"Paper","title":"Red-Teaming LLM Multi-Agent Systems via Communication Attacks","url":"https://aclanthology.org/2025.findings-acl.349/","venue":"Findings of ACL","year":2025,"authors":"Pengfei He; Yuping Lin; Shen Dong; Han Xu; Yue Xing; Hui Liu","description":"Introduces an Agent-in-the-Middle attack that intercepts and manipulates inter-agent messages, evaluated across multiple frameworks, communication structures, and applications.","why":"Shows that graph edges can compromise the whole system without altering its nodes, making message integrity and provenance part of the handoff contract.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0087","section":"Research Foundations","subcategory":"Query-dependent architecture search","rtype":"Paper","title":"Multi-agent Architecture Search via Agentic Supernet","url":"https://proceedings.mlr.press/v267/zhang25bi.html","venue":"ICML","year":2025,"authors":"Guibin Zhang; Luyang Niu; Junfeng Fang; Kun Wang; Lei Bai; Xiang Wang","description":"Represents possible agentic architectures as a probabilistic supernet and samples query-dependent systems with tailored language-model calls, tool calls, and token costs across six benchmarks.","why":"Replaces a one-size-fits-all graph with per-query architecture and resource allocation while keeping the design space explicit.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0088","section":"Critiques & Limits","subcategory":"Unequal agent contribution","rtype":"Paper","title":"Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation","url":"https://openreview.net/forum?id=5J6u03ObRZ","venue":"ICLR","year":2026,"authors":"Zhiwei Zhang et al.","description":"Identifies lazy-agent behavior in which one role dominates a nominally collaborative system, then measures causal influence and uses verifiable rewards to encourage deliberation and selective reasoning restarts.","why":"Requires builders to verify that each node contributes causal value rather than allowing a multi-agent graph to collapse into one effective agent.","evidence":"Peer-reviewed research","layer":"Roles"} {"id":"age-0089","section":"Benchmarks & Datasets","subcategory":"Multi-party negotiation","rtype":"Benchmark","title":"Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation","url":"https://openreview.net/forum?id=59E19c6yrN","venue":"NeurIPS","year":2024,"authors":"Sahar Abdelnabi; Amr Gomaa; Sarath Sivaprasad; Lea Schönherr; Mario Fritz","description":"Provides scorable, multi-agent, multi-issue negotiation games with metrics for task performance and role alignment, including cooperative, competitive, greedy, and adversarial participants.","why":"Tests communication and collective decisions when graph nodes hold conflicting objectives or attempt manipulation rather than cooperating by default.","evidence":"Benchmark/dataset","layer":"Handoffs"} {"id":"age-0090","section":"Reliability & Durable Execution","subcategory":"Agentic application security risks","rtype":"Standard","title":"OWASP Top 10 for Agentic Applications for 2026","url":"https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/","venue":"OWASP GenAI Security Project","year":2025,"authors":"OWASP GenAI Security Project; Agentic Security Initiative","description":"Provides a globally peer-reviewed operational framework for critical risks in autonomous and agentic applications, developed with contributions from more than 100 security experts, researchers, and practitioners.","why":"Supplies a practical threat checklist for node authority, tool permissions, memory provenance, inter-agent trust, and cascading failures across an agent graph.","evidence":"Industry standard","layer":"Reliability"} {"id":"age-0091","section":"Research Foundations","subcategory":"Task-adaptive graph routing","rtype":"Paper","title":"AMAS: Adaptively Determining Communication Topology for LLM-based Multi-agent System","url":"https://aclanthology.org/2025.emnlp-industry.144/","venue":"EMNLP Industry Track","year":2025,"authors":"Hui Yi Leong; Yuheng Li; Yuqing Wu; Wenwen Ouyang; Wei Zhu; Jiechao Gao; Wei Han","description":"Introduces a dynamic graph selector that uses lightweight LLM adaptation to choose task-specific communication structures and route each query through a matching agent pathway.","why":"Operationalizes topology as an input-conditioned decision rather than a fixed template; evidence is currently concentrated in question answering, mathematics, and code generation.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0092","section":"Research Foundations","subcategory":"Collaboration-mode, role, and model routing","rtype":"Paper","title":"MasRouter: Learning to Route LLMs for Multi-Agent Systems","url":"https://aclanthology.org/2025.acl-long.757/","venue":"ACL","year":2025,"authors":"Yanwei Yue; Guibin Zhang; Boyang Liu; Guancheng Wan; Kun Wang; Dawei Cheng; Yiyan Qi","description":"Defines multi-agent system routing as a joint decision over collaboration mode, role allocation, and language-model selection, implemented with a cascaded controller that progressively constructs a task-specific system.","why":"Makes node roles, model assignment, and whether collaboration is needed explicit routing decisions with measurable cost-quality trade-offs.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0093","section":"Research Foundations","subcategory":"Multi-hop evidence propagation","rtype":"Paper","title":"MOC: Multi-Order Communication in LLM-based Multi-Agent Systems","url":"https://openreview.net/forum?id=wyynWicO5s","venue":"ICML","year":2026,"authors":"Yao Guan; Lin Wang; Zhihui Lu; Ziyi Wang; Wenzhu Yan; Qiang Duan","description":"Constructs structured multi-order evidence streams so agents can receive relevant information from multiple upstream hops, then applies semantic-topological merging under token constraints.","why":"Shows that graph engineering includes what evidence survives across paths, not only which nodes and edges exist; it improves communication over a supplied topology rather than selecting that topology.","evidence":"Peer-reviewed research","layer":"Handoffs"} {"id":"age-0094","section":"Research Foundations","subcategory":"Hierarchical node and topology optimization","rtype":"Paper","title":"HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent Systems","url":"https://openreview.net/forum?id=p7p5foAXaB","venue":"ICML","year":2026,"authors":"Tianjun Yao; Zhaoyi Li; Zhiqiang Shen","description":"Models each functional role as a supernode containing heterogeneous language models in a propose-synthesis structure, then uses multi-level reward attribution and graph classification to select the inter-supernode topology.","why":"Jointly engineers node composition, role capability, credit assignment, and communication edges instead of optimizing each dimension independently.","evidence":"Peer-reviewed research","layer":"Evolution"} {"id":"age-0095","section":"State, Memory & Artifacts","subcategory":"Hierarchical graph memory","rtype":"Paper","title":"GAM: Hierarchical Graph-based Agentic Memory for LLM Agents","url":"https://aclanthology.org/2026.acl-long.1600/","venue":"ACL","year":2026,"authors":"Zhaofen Wu; Hanrong Zhang; Fulin Lin; Wujiang Xu; Xinran Xu; Yankai Chen; Henry Peng Zou; Shaowen Chen; Weizhi Zhang; Xue Liu; Philip S. Yu; Hongwei Wang","description":"Separates rapid memory encoding from stable consolidation through an event-progression graph, a topic-associative network, and graph-guided multi-factor retrieval.","why":"Provides a concrete graph contract for encoding, consolidating, connecting, and retrieving state; it does not address shared-memory permissions or concurrent writes among agents.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0096","section":"State, Memory & Artifacts","subcategory":"Selective shared memory","rtype":"Paper","title":"Learning to Share: Selective Memory for Efficient Parallel Agentic Systems","url":"https://openreview.net/forum?id=cCFyY2LmF5","venue":"ICML","year":2026,"authors":"Joseph Fioresi; Parth Parag Kulkarni; Ashmal Vayani; Song Wang; Mubarak Shah","description":"Adds a global memory bank to parallel agent teams and trains an admission controller with stepwise reinforcement learning and usage-aware credit assignment to retain reusable intermediate work.","why":"Treats cross-team memory writes as governed graph operations that can reduce duplicate work while limiting indiscriminate context growth.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0097","section":"Observability & Cost","subcategory":"Graph-aware cache reuse","rtype":"Paper","title":"Accelerating Language Model Workflows with Prompt Choreography","url":"https://aclanthology.org/2026.tacl-1.13/","venue":"TACL","year":2026,"authors":"TJ Bai; Jason Eisner","description":"Executes multi-agent workflows with a dynamic global key-value cache in which each call can attend to a reordered subset of previously encoded messages, including parallel branches.","why":"Makes workflow dependencies and reusable message state explicit execution concerns; its primary result concerns speed, and cache reuse can alter model behavior.","evidence":"Peer-reviewed research","layer":"Work graphs"} {"id":"age-0098","section":"Benchmarks & Datasets","subcategory":"Enterprise workflow benchmark","rtype":"Benchmark","title":"Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments","url":"https://aclanthology.org/2025.emnlp-main.466/","venue":"EMNLP","year":2025,"authors":"Harsh Vishwakarma; Ankush Agarwal; Ojas Patil; Chaitanya Devaguptapu; Mahesh Chandran","description":"Introduces EnterpriseBench with 500 tasks across software engineering, HR, finance, and administration, including fragmented data, access-control hierarchies, and cross-functional workflows.","why":"Tests work graphs that retrieve, modify, and transfer artifacts across services while respecting authority boundaries; the organization and services are simulated.","evidence":"Benchmark/dataset","layer":"Work graphs"} {"id":"age-0099","section":"Benchmarks & Datasets","subcategory":"Stateful tool-use evaluation","rtype":"Benchmark","title":"ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities","url":"https://aclanthology.org/2025.findings-naacl.65/","venue":"Findings of NAACL","year":2025,"authors":"Jiarui Lu; Thomas Holleis; Yizhe Zhang; Bernhard Aumayer; Feng Nan; Haoping Bai; Shuang Ma; Shen Ma; Mengyu Li; Guoli Yin; Zirui Wang; Ruoming Pang","description":"Evaluates stateful tool execution, implicit dependencies between tools, on-policy user interaction, and intermediate and final milestones over arbitrary trajectories.","why":"Transfers milestone and state-transition testing to graph prerequisites, although the evaluated system is an agent-user-tool loop rather than an agent-agent topology.","evidence":"Benchmark/dataset","layer":"State"} {"id":"age-0100","section":"Benchmarks & Datasets","subcategory":"Trajectory-level evaluation","rtype":"Benchmark","title":"AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents","url":"https://openreview.net/forum?id=4S8agvKjle","venue":"NeurIPS Datasets & Benchmarks","year":2024,"authors":"Chang Ma; Junlei Zhang; Zhihao Zhu; Cheng Yang; Yujiu Yang; Yaohui Jin; Zhenzhong Lan; Lingpeng Kong; Junxian He","description":"Unifies partially observable, multi-round agent environments and adds a fine-grained progress-rate metric plus interactive trajectory analysis beyond final success.","why":"Supplies milestone-level evaluation that can localize incomplete work, while its primarily single-agent tasks do not identify the responsible graph component by themselves.","evidence":"Benchmark/dataset","layer":"Gates"} {"id":"age-0101","section":"Benchmarks & Datasets","subcategory":"Agent security benchmark","rtype":"Benchmark","title":"Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents","url":"https://openreview.net/forum?id=V4y0CpX4hK","venue":"ICLR","year":2025,"authors":"Hanrong Zhang; Jingyuan Huang; Kai Mei; Yifei Yao; Zhenting Wang; Chenlu Zhan; Hongwei Wang; Yongfeng Zhang","description":"Benchmarks prompt injection, memory poisoning, backdoors, mixed attacks, and defenses across 10 scenarios, more than 400 tools, 13 model backbones, and seven metrics.","why":"Maps security failures to prompts, plans, tools, and memory surfaces that graph designs must isolate separately; inter-agent trust propagation is outside its main evaluation target.","evidence":"Benchmark/dataset","layer":"Reliability"} {"id":"age-0102","section":"Benchmarks & Datasets","subcategory":"Prompt-injection evaluation","rtype":"Benchmark","title":"AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents","url":"https://openreview.net/forum?id=m1YYAQjO3w","venue":"NeurIPS Datasets & Benchmarks","year":2024,"authors":"Edoardo Debenedetti; Jie Zhang; Mislav Balunovic; Luca Beurer-Kellner; Marc Fischer; Florian Tramèr","description":"Provides an extensible adversarial environment with 97 realistic tasks and 629 security test cases for agents executing tools over untrusted data.","why":"Tests the instruction-data boundary that every external-input and artifact edge must preserve; compromised peer-agent messages require separate evaluation.","evidence":"Benchmark/dataset","layer":"Reliability"} {"id":"age-0103","section":"Benchmarks & Datasets","subcategory":"Agent misuse benchmark","rtype":"Benchmark","title":"AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents","url":"https://openreview.net/forum?id=AC5n7xHuR1","venue":"ICLR","year":2025,"authors":"Maksym Andriushchenko; Alexandra Souly; Mateusz Dziemian; Derek Duenas; Maxwell Lin; Justin Wang; Dan Hendrycks; Andy Zou; J. Zico Kolter; Matt Fredrikson; Yarin Gal; Xander Davies","description":"Evaluates refusal and retained task capability on 110 explicitly malicious multi-stage agent tasks with 440 augmented variants across 11 harm categories.","why":"Tests whether safety gates reject prohibited work and stop a compromised node from completing a harmful path; the score is not a general measure of agent-system safety.","evidence":"Benchmark/dataset","layer":"Gates"} {"id":"age-0104","section":"Reliability & Durable Execution","subcategory":"Structural prompt-injection defense","rtype":"Paper","title":"IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents","url":"https://aclanthology.org/2025.emnlp-main.53/","venue":"EMNLP","year":2025,"authors":"Hengyu An; Jinghuai Zhang; Tianyu Du; Chunyi Zhou; Qingming Li; Tao Lin; Shouling Ji","description":"Models task execution as traversal over a planned Tool Dependency Graph and separates action planning from interaction with untrusted external data.","why":"Shows how structural constraints on allowed tool edges can block unauthorized actions instead of relying only on prompts or classifiers.","evidence":"Peer-reviewed research","layer":"Work graphs"} {"id":"age-0105","section":"Protocols & Handoffs","subcategory":"Agent identity and delegated authorization","rtype":"Docs","title":"Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world","url":"https://openid.net/wp-content/uploads/2025/10/Identity-Management-for-Agentic-AI.pdf","venue":"OpenID Foundation","year":2025,"authors":"OpenID Foundation Artificial Intelligence Identity Management Community Group","description":"Maps identity, authentication, delegated authority, consent, governance, and accountability gaps for autonomous agents against OAuth and OpenID Connect foundations.","why":"Provides a standards-grounded vocabulary for identity-bearing nodes and auditable delegation across agent and tool edges; it is a non-normative landscape whitepaper, not a protocol.","evidence":"Official documentation","layer":"Handoffs"} {"id":"age-0106","section":"Protocols & Handoffs","subcategory":"Agent capability and discovery schema","rtype":"Standard","title":"Open Agentic Schema Framework (OASF) v1.1.0","url":"https://github.com/agntcy/oasf/releases/tag/v1.1.0","venue":"AGNTCY Project, Linux Foundation","year":2026,"authors":"AGNTCY Project","description":"Defines versioned schemas and taxonomies for agent identity metadata, capabilities, skills, modules, interactions, and discovery records, including MCP and A2A alignment modules.","why":"Gives graph builders a portable node descriptor for capability-based discovery and routing validation; it describes capabilities but does not execute a graph or authorize an edge.","evidence":"Industry standard","layer":"Roles"} {"id":"age-0107","section":"State, Memory & Artifacts","subcategory":"Artifact provenance","rtype":"Standard","title":"PROV-DM: The PROV Data Model","url":"https://www.w3.org/TR/prov-dm/","venue":"W3C Recommendation","year":2013,"authors":"Luc Moreau; Paolo Missier","description":"Defines entities, activities, agents, derivation, attribution, responsibility, bundles, and collections for interoperable provenance records.","why":"Maps artifacts, execution steps, and responsible nodes into a portable lineage model; storage, access control, and cryptographic integrity remain implementation concerns.","evidence":"Industry standard","layer":"State"} {"id":"age-0108","section":"State, Memory & Artifacts","subcategory":"Cryptographic artifact provenance","rtype":"Paper","title":"in-toto: Providing farm-to-table guarantees for bits and bytes","url":"https://www.usenix.org/conference/usenixsecurity19/presentation/torres-arias","venue":"USENIX Security","year":2019,"authors":"Santiago Torres-Arias; Hammad Afzali; Trishank Karthik Kuppusamy; Reza Curtmola; Justin Cappos","description":"Introduces signed material and product links for supply-chain steps and verifies them against an expected layout, supported by analysis of 30 historical compromises and deployed integrations.","why":"Supplies a transferable receipt pattern that binds each artifact transformation to an authorized step and identity; it verifies process integrity, not semantic correctness.","evidence":"Peer-reviewed research","layer":"State"} {"id":"age-0109","section":"Reliability & Durable Execution","subcategory":"Portable workflow semantics","rtype":"Standard","title":"Serverless Workflow Specification v1.0.0","url":"https://serverlessworkflow.io/blog/releases/release-100/","venue":"Serverless Workflow, Cloud Native Computing Foundation","year":2025,"authors":"Serverless Workflow Project","description":"Defines a vendor-neutral workflow language with sequential and concurrent tasks, event correlation, service calls, error handling, retries, and timeouts.","why":"Provides an inspectable execution contract for production work graphs; conformance does not itself supply a durable runtime, an agent protocol, or model-level correctness.","evidence":"Industry standard","layer":"Reliability"} {"id":"age-0110","section":"Critiques & Limits","subcategory":"Compound-system call scaling","rtype":"Paper","title":"Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems","url":"https://proceedings.neurips.cc/paper_files/paper/2024/hash/51173cf34c5faac9796a47dc2fdd3a71-Abstract-Conference.html","venue":"NeurIPS","year":2024,"authors":"Lingjiao Chen; Jared Davis; Boris Hanin; Peter Bailis; Ion Stoica; Matei Zaharia; James Zou","description":"Analyzes Vote and Filter-Vote compound systems and finds that performance can rise and then fall as the number of language-model calls increases because query difficulty is heterogeneous.","why":"Bounds claims that adding calls, voters, or nodes automatically improves a system; the experiments cover simple aggregation systems rather than rich agent organizations.","evidence":"Peer-reviewed research","layer":"Observability & cost"} {"id":"age-0111","section":"Verification & Evals","subcategory":"Verification-aware work graphs","rtype":"Paper","title":"Verification-Aware Planning for Multi-Agent Systems","url":"https://aclanthology.org/2026.eacl-long.353/","venue":"EACL","year":2026,"authors":"Tianyang Xu; Dan Zhang; Kushan Mitra; Estevam Hruschka","description":"Introduces VeriMAP, which decomposes tasks into a dependency graph and attaches planner-defined Python and natural-language verification functions to subtasks before execution.","why":"Makes acceptance criteria part of the work graph so handoff failures can trigger local repair instead of surfacing only in a final answer.","evidence":"Peer-reviewed research","layer":"Gates"} {"id":"age-0112","section":"Research Foundations","subcategory":"Human-in-the-loop work-graph planning","rtype":"Paper","title":"AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems","url":"https://aclanthology.org/2025.emnlp-demos.7/","venue":"EMNLP System Demonstrations","year":2025,"authors":"Hannah Kim; Kushan Mitra; Chen Shen; Dan Zhang; Estevam Hruschka","description":"Presents conversational and graph-based interfaces for inspecting, editing, and collaboratively guiding plans in orchestrated multi-agent systems.","why":"Treats the work graph as a human-legible control surface; its non-autonomous agents execute assigned tasks, so it is a planning substrate rather than a complete agent organization.","evidence":"Peer-reviewed research","layer":"Work graphs"} {"id":"age-0113","section":"Frameworks & SDKs","subcategory":"Graph workflow runtime","rtype":"Blog","title":"Build reliable multi-agent applications with ADK Go 2.0","url":"https://developers.googleblog.com/announcing-adk-go-20/","venue":"Google Developers Blog","year":2026,"authors":"Toni Klopfenstein; Sampath Kumar Maddula","description":"Documents a graph-based multi-agent workflow engine with typed nodes, conditional edges, fan-out and fan-in, nested graphs, cycles, retries, durable human input, state, and unified telemetry.","why":"Provides a first-party implementation of explicit, resumable work graphs while showing that function and tool nodes remain distinct from agent nodes.","evidence":"Practitioner analysis","layer":"Work graphs"} {"id":"age-0114","section":"Research Foundations","subcategory":"Decentralized adaptive coordination","rtype":"Paper","title":"AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems","url":"https://proceedings.neurips.cc/paper_files/paper/2025/hash/9a379c1b05793d1c42dc832269834515-Abstract-Conference.html","venue":"NeurIPS","year":2025,"authors":"Yingxuan Yang; Huacan Chai; Shuai Shao; Yuanyi Song; Siyuan Qi; Renting Rui; Weinan Zhang","description":"Organizes specialized agents in a decentralized DAG, uses retrieval-backed local experience for capability refinement, and adapts routing without a single central orchestrator.","why":"Offers a contrasting topology for settings where central coordination is a bottleneck or trust boundary; its reported evidence is concentrated in reasoning and coding tasks.","evidence":"Peer-reviewed research","layer":"Topology"} {"id":"age-0115","section":"Reliability & Durable Execution","subcategory":"Effect-typed agent contracts","rtype":"Paper","title":"ETAS: An Effect-Typed Language for Agent Systems","url":"https://arxiv.org/abs/2607.17780","venue":"arXiv","year":2026,"authors":"Huiri Tan; Yikun Wang; Puyang Zhang; Shangyu Li; Jiasi Shen","description":"Defines a language in which agents, tools, typed memory, approvals, policies, effects, and execution traces are semantic program elements, with static obligations and runtime monitors.","why":"Provides a transferable contract model for authorizing and auditing node actions before and during execution; it is a new preprint and not yet deployment evidence.","evidence":"Research preprint","layer":"Reliability"} {"id":"age-0116","section":"Critiques & Limits","subcategory":"Consensus-induced search collapse","rtype":"Paper","title":"The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search","url":"https://arxiv.org/abs/2607.18045","venue":"arXiv","year":2026,"authors":"Yohei Nakajima","description":"Uses an exactly solvable multi-searcher benchmark to show that pooling information into one repeated recommendation can improve belief accuracy while sharply reducing collective discovery coverage.","why":"Separates information quality from allocation policy and warns that consensus edges can collapse parallel exploration unless a coordinator preserves a portfolio of actions.","evidence":"Research preprint","layer":"Topology"} {"id":"age-0117","section":"Critiques & Limits","subcategory":"Relay information bottlenecks","rtype":"Paper","title":"When Do Multi-Agent Systems Help? An Information Bottleneck Perspective","url":"https://arxiv.org/abs/2607.16133","venue":"arXiv","year":2026,"authors":"Wendi Yu; Lianhao Zhou; Xiangjue Dong; Sai Sudarshan Barath; Declan Staunton; Byung-Jun Yoon; Xiaoning Qian; James Caverlee; Shuiwang Ji","description":"Formalizes bounded inter-agent relays as an information bottleneck and reports 18 controlled experiments across five benchmarks and three model scales.","why":"Explains when isolated contexts and compressed handoffs save useful context and when they discard task-relevant information, providing a direct test for whether a graph earns its edge losses.","evidence":"Research preprint","layer":"Handoffs"} {"id":"age-0118","section":"Critiques & Limits","subcategory":"Critique uptake failure","rtype":"Paper","title":"Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning","url":"https://arxiv.org/abs/2607.15388","venue":"arXiv","year":2026,"authors":"Chih-Hsuan Yang; Jingyan Jiang; Vikram Vasudevan; Cheng-Hau Yang; Huihuo Zheng; Le Chen; Eliu A. Huerta; Venkatram Vishwanath; Ian T. Foster; Rajeev Thakur","description":"Evaluates 4,181 verifier-grounded math problems and finds that a precise reviewer can still yield weak repair when the protocol does not carry useful critique into the next candidate.","why":"Shows that a verifier node is insufficient by itself: the outgoing edge must couple accepted evidence to a concrete state change or retry action.","evidence":"Research preprint","layer":"Gates"}