id,section,subcategory,rtype,title,url,venue,year,authors,description,why,evidence,layer age-0001,Start Here,Architecture guide,Blog,A practical guide to building agents,https://openai.com/business/guides-and-resources/a-practical-guide-to-building-ai-agents/,OpenAI Guides & Resources,2025,OpenAI,"Explains when agents are appropriate, how to define their tools and instructions, and how to move from one agent to manager and handoff-based multi-agent patterns.",A useful first check on whether distinct roles justify the coordination cost of a graph.,Practitioner analysis,Roles age-0002,Start Here,Architecture guide,Blog,Building effective agents,https://www.anthropic.com/engineering/building-effective-agents,Anthropic Engineering,2024,Erik Schluntz; Barry Zhang,"Distinguishes workflows from autonomous agents and presents routing, parallelization, orchestrator-worker, and evaluator-optimizer patterns.",Provides a compact topology vocabulary and a strong case for adding coordination only when the task demands it.,Practitioner analysis,Topology age-0003,Research Foundations,Agent foundations,Paper,Agent-Oriented Programming,https://doi.org/10.1016/0004-3702(93)90034-9,Artificial Intelligence,1993,Yoav Shoham,Defines a programming paradigm in which agents are first-class components described through mental state and governed by explicit interaction rules.,Establishes the intellectual lineage for treating an agent role as a programmable organizational unit.,Peer-reviewed research,Roles age-0004,Research Foundations,Agent foundations,Paper,Intelligent Agents: Theory and Practice,https://doi.org/10.1017/S0269888900008122,The Knowledge Engineering Review,1995,Michael Wooldridge; Nicholas R. Jennings,"Surveys the properties, architectures, and engineering approaches that distinguish autonomous agents from ordinary software modules.",Grounds the agency-at-the-nodes boundary that separates an agent graph from a deterministic workflow.,Peer-reviewed research,Roles age-0005,Research Foundations,Shared-state architectures,Paper,The Blackboard Model of Problem Solving and the Evolution of Blackboard Architectures,https://doi.org/10.1609/aimag.v7i2.537,AI Magazine,1986,H. Penny Nii,Describes systems in which independent specialists coordinate opportunistically through a shared problem state and a control component.,Supplies a durable model for shared state without requiring every node to exchange its full context directly.,Peer-reviewed research,State age-0006,Research Foundations,Learned communication,Paper,Learning to Communicate with Deep Multi-Agent Reinforcement Learning,https://proceedings.neurips.cc/paper_files/paper/2016/hash/c7635bfd99248a2cdef8249ef7bfbef4-Abstract.html,NeurIPS,2016,Jakob Foerster; Ioannis Alexandros Assael; Nando de Freitas; Shimon Whiteson,"Introduces reinforcement-learning methods that let agents learn communication protocols alongside their task policies, including discrete messages for execution.",Shows that edge content and communication policy can be engineered or learned rather than treated as free-form chat.,Peer-reviewed research,Handoffs age-0007,Research Foundations,Learned communication,Paper,TarMAC: Targeted Multi-Agent Communication,https://proceedings.mlr.press/v97/das19a.html,ICML,2019,Abhishek Das et al.,Uses attention to let agents address different messages to selected recipients instead of broadcasting the same information to the whole team.,"Motivates selective, recipient-aware handoffs when all-to-all communication is wasteful or distracting.",Peer-reviewed research,Handoffs age-0008,Research Foundations,LLM multi-agent systems,Paper,AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversations,https://openreview.net/forum?id=BAakY1hNKS,COLM,2024,Qingyun Wu; Gagan Bansal; Jieyu Zhang; Yiran Wu; Beibin Li; Erkang Zhu; Li Jiang; Xiaoyun Zhang; Shaokun Zhang; Jiale Liu; Ahmed Hassan Awadallah; Ryen W. White; Doug Burger; Chi Wang,"Presents a framework for composing customizable conversational agents that can combine language models, tools, code execution, and human input.","An early, influential demonstration that agent roles and conversation links can be expressed as an executable topology.",Peer-reviewed research,Topology age-0009,Research Foundations,Topology optimization,Paper,GPTSwarm: Language Agents as Optimizable Graphs,https://proceedings.mlr.press/v235/zhuge24a.html,ICML,2024,Mingchen Zhuge et al.,Represents language-agent systems as computational graphs and optimizes graph components from task feedback.,Makes the graph itself an optimization target rather than a fixed orchestration diagram.,Peer-reviewed research,Evolution age-0010,Research Foundations,Dynamic topology,Paper,A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration,https://openreview.net/forum?id=XII0Wp1XA9,COLM,2024,Zijun Liu et al.,Constructs task-specific collaboration networks that can vary which agents participate and how they communicate instead of relying on one fixed team.,Provides evidence for adapting the work graph to the task while keeping the available agent roles reusable.,Peer-reviewed research,Evolution age-0011,Research Foundations,Communication efficiency,Paper,Cut the Crap: An Economical Communication Pipeline for LLM-based Multi-Agent Systems,https://openreview.net/forum?id=LkzuPorQ5L,ICLR,2025,Guibin Zhang et al.,Studies an economical communication pipeline that reduces redundant information exchanged among language-model agents.,Treats edge traffic as a measurable cost and tests whether less communication can preserve useful collaboration.,Peer-reviewed research,Observability & cost age-0012,Research Foundations,Topology optimization,Paper,G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks,https://proceedings.mlr.press/v267/zhang25cu.html,ICML,2025,Guibin Zhang et al.,Uses graph neural networks to design communication structures for multi-agent systems instead of assuming a complete or manually chosen graph.,Connects task performance to explicit topology search and exposes communication structure as an engineering variable.,Peer-reviewed research,Evolution age-0013,Research Foundations,System search,Paper,Automated Design of Agentic Systems,https://openreview.net/forum?id=t9U3LW7JVX,ICLR,2025,Shengran Hu; Cong Lu; Jeff Clune,"Uses a meta-agent to propose, evaluate, and iteratively improve code-defined agentic systems across tasks.",Demonstrates automated search over coordination logic while retaining executable artifacts that engineers can inspect.,Peer-reviewed research,Evolution age-0014,Research Foundations,Workflow search,Paper,AFlow: Automating Agentic Workflow Generation,https://openreview.net/forum?id=z5uVAKwmjf,ICLR,2025,Jiayi Zhang et al.,Searches over reusable workflow operators to generate task-specific agentic workflows and improve them from evaluation results.,Offers a concrete method for evolving work graphs against measurable objectives rather than intuition alone.,Peer-reviewed research,Evolution age-0015,Research Foundations,Joint optimization,Paper,Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies,https://openreview.net/forum?id=I05H9RUzHB,ICLR,2026,Han Zhou et al.,Studies joint optimization of agent prompts and communication topology instead of tuning either component in isolation.,Shows that node behavior and graph structure interact and may need to evolve together.,Peer-reviewed research,Evolution age-0016,Research Foundations,Debate and councils,Paper,Improving Factuality and Reasoning in Language Models through Multiagent Debate,https://proceedings.mlr.press/v235/du24e.html,ICML,2024,Yilun Du et al.,Tests rounds of proposal and critique among multiple language-model instances as a way to improve factual and reasoning answers.,Supplies an empirical basis for debate-style gates while leaving room to examine correlated errors and added cost.,Peer-reviewed research,Gates age-0017,Research Foundations,Debate and councils,Paper,Improving Multi-Agent Debate with Sparse Communication Topology,https://aclanthology.org/2024.findings-emnlp.427/,Findings of EMNLP,2024,Yunxuan Li et al.,Examines multi-agent debate under sparse communication structures rather than defaulting to full information exchange among every participant.,Isolates topology as a factor in debate quality and communication efficiency.,Peer-reviewed research,Topology age-0018,Research Foundations,Feedback and memory,Paper,Reflexion: Language Agents with Verbal Reinforcement Learning,https://openreview.net/forum?id=vAElhFcKW6,NeurIPS,2023,Noah Shinn et al.,Lets an agent convert feedback into textual reflections stored in episodic memory and reused on later attempts.,Clarifies how a node loop can persist learning across retries without changing model weights.,Peer-reviewed research,State age-0019,Research Foundations,Tool-grounded verification,Paper,CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing,https://openreview.net/forum?id=Sx038qxjek,ICLR,2024,Zhibin Gou et al.,Uses external tools to obtain feedback that a language model can apply when critiquing and revising its outputs.,Supports verification gates grounded in observable evidence instead of another ungrounded model opinion.,Peer-reviewed research,Gates age-0020,Production Case Studies,Generalist agent team,Paper,Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks,https://www.microsoft.com/en-us/research/publication/magentic-one-a-generalist-multi-agent-system-for-solving-complex-tasks/,Microsoft Research MSR-TR-2024-47,2024,Adam Fourney et al.,"Documents an orchestrator-led team of specialized agents for web, file, coding, and terminal tasks, with progress tracking and replanning.",Provides a concrete generalist topology whose role boundaries and recovery behavior can be examined as a system.,Research preprint,Roles age-0021,Production Case Studies,Parallel research,Blog,How we built our multi-agent research system,https://www.anthropic.com/engineering/multi-agent-research-system,Anthropic Engineering,2025,Jeremy Hadfield et al.,"Describes a lead research agent that creates parallel subagents, delegates searches, and synthesizes their findings, including operational lessons from production.","A detailed case study of dynamic fan-out and fan-in, context separation, evaluation, and the token cost of an adaptive work graph.",Practitioner analysis,Work graphs age-0022,Frameworks & SDKs,Role-based teams,Docs,CrewAI Crews,https://docs.crewai.com/en/concepts/crews,CrewAI Documentation,2026,CrewAI,"Documents role-based agent crews, assigned tasks, delegation, and sequential or hierarchical execution processes.",Offers accessible primitives for testing explicit ownership and manager-worker coordination.,Official documentation,Roles age-0023,Frameworks & SDKs,Agent orchestration,Docs,OpenAI Agents SDK: Agent orchestration,https://openai.github.io/openai-agents-python/multi_agent/,OpenAI Agents SDK Documentation,2026,OpenAI,Explains manager-style orchestration with agents exposed as tools and decentralized orchestration through handoffs.,Makes the centralized-versus-decentralized topology choice explicit in a production SDK.,Official documentation,Topology age-0024,Frameworks & SDKs,Pattern catalog,Docs,Strands Agents: Multi-Agent Patterns,https://strandsagents.com/docs/user-guide/concepts/multi-agent/multi-agent-patterns/,Strands Agents Documentation,2026,Strands Agents; AWS,"Documents several coordination shapes for composing agents, including supervisor, swarm, workflow, and graph-oriented patterns.",Lets builders compare topology choices within one SDK instead of treating one pattern as universal.,Official documentation,Topology age-0025,Frameworks & SDKs,Typed agents,Docs,Pydantic AI: Multi-Agent Applications,https://pydantic.dev/docs/ai/guides/multi-agent-applications/,Pydantic AI Documentation,2026,Pydantic,"Shows delegation, programmatic control flow, and graph-based state machines for composing typed Python agents.","Useful for expressing node inputs, outputs, dependencies, and orchestration boundaries in ordinary application code.",Official documentation,Topology age-0026,Frameworks & SDKs,Multi-agent orchestration,Docs,LlamaIndex: Multi-Agent Patterns,https://developers.llamaindex.ai/python/framework/understanding/agent/multi_agent/,LlamaIndex Documentation,2026,LlamaIndex,"Covers agent workflow, orchestrator, and planner-oriented approaches for coordinating specialized agents.",Provides implementation patterns for choosing who owns delegation and how results return to the coordinating node.,Official documentation,Handoffs age-0027,Frameworks & SDKs,Graph workflows,Docs,Google ADK: Graph-based Agent Workflows,https://adk.dev/graphs/,Google ADK Documentation,2026,Google,"Documents declarative workflows whose nodes combine agents, tools, functions, and human input through explicit edges, typed data passing, routing, branching, state, fan-out and join, loops, escalation, and nesting.",Makes the graph load-bearing and inspectable while separating deterministic process control from model reasoning.,Official documentation,Work graphs age-0028,Frameworks & SDKs,Graph workflows,Docs,Microsoft Agent Framework: Workflows,https://learn.microsoft.com/en-us/agent-framework/workflows/,Microsoft Learn,2026,Microsoft,"Describes workflows built from executors and explicit edges, with support for branching, aggregation, state, and checkpointing.",Exposes the work graph as an inspectable program rather than hiding coordination inside prompts.,Official documentation,Work graphs age-0029,Frameworks & SDKs,Graph runtime,Docs,LangGraph overview,https://docs.langchain.com/oss/python/langgraph/overview,LangGraph Documentation,2025,LangChain,"Introduces a low-level runtime for stateful agent graphs with durable execution, streaming, memory, and human intervention.","A widely used substrate for implementing explicit nodes, edges, state transitions, and resumable work graphs.",Official documentation,Work graphs age-0030,Protocols & Handoffs,In-process transfer,Docs,OpenAI Agents SDK: Handoffs,https://openai.github.io/openai-agents-python/handoffs/,OpenAI Agents SDK Documentation,2026,OpenAI,"Documents transfers from one agent to another, including tool-shaped handoff schemas, input filters, and callbacks.",Turns an edge into an explicit contract controlling when ownership moves and what context crosses with it.,Official documentation,Handoffs age-0031,Protocols & Handoffs,Tool and context protocol,Standard,Model Context Protocol Specification 2025-11-25,https://modelcontextprotocol.io/specification/2025-11-25,MCP Specification 2025-11-25,2025,Model Context Protocol project; Agentic AI Foundation,"Specifies a client-server protocol through which AI applications discover and use tools, resources, prompts, and contextual data.","Standardizes capability and context edges, while remaining distinct from a protocol for delegating work between autonomous agents.",Industry standard,Handoffs age-0032,Protocols & Handoffs,Agent interoperability,Standard,Agent2Agent Protocol Specification v1.0.0,https://a2a-protocol.org/v1.0.0/specification/,A2A Specification,2026,A2A Protocol Working Group; Linux Foundation,"Defines interoperable agent discovery, task lifecycle, messages, artifacts, streaming, asynchronous updates, version negotiation, and multiple protocol bindings across service boundaries.",Provides a version-pinned wire contract for cross-system handoffs where agents cannot share an in-process runtime.,Industry standard,Handoffs age-0033,"State, Memory & Artifacts",Checkpointing,Docs,LangGraph Persistence,https://docs.langchain.com/oss/python/langgraph/persistence,LangGraph Documentation,2025,LangChain,"Documents thread-scoped checkpoints, saved state, replay, state inspection, and memory storage for LangGraph runs.",Shows how graph state can become the recoverable system of record instead of living only in model context.,Official documentation,State age-0034,"State, Memory & Artifacts",Conversation state,Docs,OpenAI Agents SDK: Sessions,https://openai.github.io/openai-agents-python/sessions/,OpenAI Agents SDK Documentation,2026,OpenAI,Documents persistent conversation history that can be loaded and updated across repeated agent runs.,Provides a bounded mechanism for carrying state across nodes and turns without manually rebuilding every prompt.,Official documentation,State age-0035,Verification & Evals,Workflow evaluation,Docs,Evaluate agent workflows,https://developers.openai.com/api/docs/guides/agent-evals,OpenAI API Documentation,2026,OpenAI,"Explains how to build datasets, graders, trace-based evaluations, and reproducible checks for agent workflows.",Makes gates testable at both the final outcome and the intermediate handoff level.,Official documentation,Gates age-0036,Verification & Evals,Evaluation practice,Blog,Demystifying evals for AI agents,https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents,Anthropic Engineering,2026,Mikaela Grace et al.,"Presents practical guidance for defining tasks, outcomes, graders, and evaluation suites for agents with non-deterministic trajectories.",Helps turn vague reviewer judgments into evidence gates that can guide graph changes.,Practitioner analysis,Gates age-0037,Verification & Evals,Evaluation framework,Tool,Inspect AI,https://inspect.aisi.org.uk/,Inspect AI Documentation,2024,UK AI Security Institute,"Provides an open-source evaluation framework with tasks, solvers, scorers, sandboxed tools, and structured logs.",Supports reproducible gate nodes and traceable evidence for both individual agents and composed systems.,Maintained OSS project,Gates age-0038,Reliability & Durable Execution,Durable actors and workflows,Docs,Dapr Agents introduction,https://docs.dapr.io/developing-ai/dapr-agents/dapr-agents-introduction/,Dapr Documentation,2026,Dapr; CNCF,"Introduces an agent framework built on durable actors and workflows, with state, messaging, and recovery supplied by the Dapr runtime.",Shows how graph nodes can inherit distributed-systems durability instead of implementing recovery only in prompts.,Official documentation,Reliability age-0039,Reliability & Durable Execution,Durable execution integration,Tool,Temporal integration for OpenAI Agents SDK,https://github.com/temporalio/sdk-python/tree/main/temporalio/contrib/openai_agents,Temporal Python SDK,2025,Temporal,"Integrates OpenAI agent runs, model calls, and tools with Temporal workflows and activities for durable execution.","Provides replay, retry, timeout, and recovery semantics beneath an agent graph without asking the model to manage them.",Maintained OSS project,Reliability age-0040,Reliability & Durable Execution,Database-backed workflows,Docs,DBOS AI Quickstart,https://docs.dbos.dev/ai/ai-quickstart,DBOS Documentation,2026,DBOS,Shows how to place AI application steps inside durable workflows whose progress is recorded and recoverable after interruption.,Offers a compact path from an agent prototype to resumable execution with explicit step boundaries.,Official documentation,Reliability age-0041,Reliability & Durable Execution,Durable agent patterns,Docs,Restate Durable Agents,https://docs.restate.dev/ai/patterns/durable-agents,Restate Documentation,2026,Restate,"Documents durable agent patterns using persisted execution, reliable calls, retries, timers, and stateful services.",Maps common mid-graph failures to runtime guarantees rather than fragile application-level retry code.,Official documentation,Reliability age-0042,Observability & Cost,Tracing,Docs,OpenAI Agents SDK: Tracing,https://openai.github.io/openai-agents-python/tracing/,OpenAI Agents SDK Documentation,2026,OpenAI,"Documents traces and spans for agent runs, model generations, tool calls, handoffs, and guardrails.","Makes node and edge behavior inspectable so latency, failures, and expensive paths can be attributed correctly.",Official documentation,Observability & cost age-0043,Observability & Cost,Telemetry standard,Standard,OpenTelemetry GenAI Semantic Conventions,https://github.com/open-telemetry/semantic-conventions-genai,OpenTelemetry GenAI repository,2026,OpenTelemetry GenAI SIG,"Develops shared telemetry names and attributes for generative-AI model, tool, and agent operations in traces, metrics, and events.",Helps graph telemetry remain portable across runtimes and observability vendors.,Maintained OSS project,Observability & cost age-0044,Observability & Cost,Cost attribution,Docs,Arize Phoenix: Cost Tracking,https://arize.com/docs/phoenix/tracing/how-to-tracing/cost-tracking,Arize Phoenix Documentation,2026,Arize AI,Explains how Phoenix derives and displays token usage and model cost from traced generative-AI calls.,Supports per-node and per-path cost accounting instead of treating a multi-agent run as one opaque bill.,Official documentation,Observability & cost age-0045,Observability & Cost,Observability platform,Docs,LangSmith Observability,https://docs.langchain.com/langsmith/observability,LangSmith Documentation,2026,LangChain,"Documents tracing, dashboards, alerts, feedback, and evaluation views for language-model and agent applications.",Provides operational views for following execution across nodes and locating failures on the critical path.,Official documentation,Observability & cost age-0046,Benchmarks & Datasets,Collaboration benchmark,Benchmark,MultiAgentBench: Evaluating Collaboration and Competition of LLM Agents,https://aclanthology.org/2025.acl-long.421/,ACL,2025,Kunlun Zhu et al.,Benchmarks language-model agents in collaborative and competitive settings while examining coordination processes as well as task outcomes.,Measures properties of the team interaction that single-agent benchmarks cannot expose.,Benchmark/dataset,Observability & cost age-0047,Benchmarks & Datasets,Failure diagnosis,Benchmark,Why Do Multi-Agent LLM Systems Fail?,https://nips.cc/virtual/2025/poster/121528,NeurIPS Datasets & Benchmarks,2025,Mert Cemri et al.,Provides a structured taxonomy and evaluation approach for diagnosing coordination failures in multi-agent language-model systems.,"Turns reliability incidents into recurring, attributable failure classes that can guide graph redesign.",Benchmark/dataset,Reliability age-0048,Critiques & Limits,Self-correction limits,Paper,Large Language Models Cannot Self-Correct Reasoning Yet,https://openreview.net/forum?id=IkmD3fKBPQ,ICLR,2024,Jie Huang et al.,Finds that intrinsic self-correction without reliable external feedback often fails to improve reasoning and can reduce accuracy.,Warns against using an ungrounded critic node as evidence simply because it is separate from the producing node.,Peer-reviewed research,Gates age-0049,Critiques & Limits,Scaling evidence,Paper,Towards a Science of Scaling Agent Systems,https://arxiv.org/abs/2512.08296,arXiv; Google Research,2025,Yubin Kim et al.,"Studies how agent-system performance changes across tasks, models, coordination structures, and scaling choices under controlled experiments.",Tests the assumption that adding agents reliably helps and frames scaling as an empirical topology decision.,Research preprint,Topology age-0050,Critiques & Limits,Token-budget comparison,Paper,Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets,https://arxiv.org/abs/2604.02460,arXiv,2026,Dat Tran; Douwe Kiela,Compares single-agent and multi-agent approaches to multi-hop reasoning while holding the total thinking-token budget constant.,"Provides the cost-controlled baseline needed before claiming that coordination, rather than extra inference, caused an improvement.",Research preprint,Topology age-0051,Start Here,Contemporary framing,Blog,From Loop Engineering to Graph Engineering?,https://x.com/IntuitMachine/status/2078419526354378975,X Articles,2026,Carlos E. Perez,"Frames the shift as loop architecture: networks of improvement cycles that monitor, feed, constrain, and correct one another, with reliability located in their edges.","Adds the grounding requirement missing from topology-only accounts: independent counter-metrics, frozen tests or rules, external anchors, and human ownership of root objectives.",Practitioner analysis,Gates age-0052,Research Foundations,Blackboard coordination,Paper,"A Multi-Level Organization for Problem Solving Using Many, Diverse, Cooperating Sources of Knowledge",https://www.ijcai.org/Proceedings/75/Papers/072.pdf,IJCAI,1975,Lee D. Erman; Victor R. Lesser,"Presents the Hearsay-II multi-level blackboard, where independent knowledge sources react to shared hypotheses, create explicit structural dependencies, and verify or revise one another's contributions.",Provides an early architecture for loosely coupled specialist nodes coordinating through inspectable shared state instead of direct all-to-all calls.,Peer-reviewed research,State age-0053,Research Foundations,Negotiated delegation,Paper,The Contract Net Protocol: High-Level Communication and Control in a Distributed Problem Solver,https://doi.org/10.1109/TC.1980.1675516,IEEE Transactions on Computers,1980,Reid G. Smith,"Defines a negotiation protocol in which managers announce tasks, potential contractors bid, and awards establish temporary problem-solving relationships.",Supplies the classic contract for capability-aware delegation and auditable assignment edges between autonomous nodes.,Peer-reviewed research,Handoffs age-0054,Start Here,Communication survey,Paper,"The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why – A Survey from MARL to Emergent Language and LLMs",https://openreview.net/pdf?id=LGsed0QQVq,Transactions on Machine Learning Research,2026,Jingdi Chen; Hanqing Yang; Zongjun Liu; Carlee Joe-Wong,"Synthesizes multi-agent communication across reinforcement learning, emergent language, and LLM systems through sender, recipient, timing, content, and purpose decisions.","Provides a design-oriented map for engineering edge selection, message timing, payloads, grounding, scalability, and interpretability.",Peer-reviewed research,Handoffs age-0055,Research Foundations,Collaboration scaling,Paper,Scaling Large Language Model-based Multi-Agent Collaboration,https://openreview.net/forum?id=K3n5jPkrU6,ICLR,2025,Chen Qian; Zihao Xie; YiFei Wang; Wei Liu; Kunlun Zhu; Hanchen Xia; Yufan Dang; Zhuoyun Du; Weize Chen; Cheng Yang; Zhiyuan Liu; Maosong Sun,"Introduces MacNet, a DAG-based collaboration architecture executed in topological order, and studies communication structure while scaling experiments beyond 1,000 agents.",Makes topology a causal scaling variable rather than assuming that larger teams or denser communication automatically help.,Peer-reviewed research,Topology age-0056,Research Foundations,Holistic orchestration,Paper,MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks,https://openreview.net/forum?id=3fGXBm4c3S,ICML,2026,Zixuan Ke et al.,"Formulates orchestration as reinforcement-learned generation of a complete multi-agent program and evaluates it across controlled dimensions including depth, horizon, breadth, parallelism, and robustness.",Tests when whole-system graph structure helps instead of attributing gains to coordination without controlled task evidence.,Peer-reviewed research,Work graphs age-0057,Research Foundations,Conditional topology,Paper,CARD: Towards Conditional Design of Multi-agent Topological Structures,https://openreview.net/forum?id=JgvJdICc6P,ICLR,2026,Tongtong Wu et al.,"Generates communication graphs conditioned on agent roles, models, tools, and data sources, and adapts topology as task resources change.",Treats the available capabilities and directed edges as a versionable organizational artifact rather than a fixed team template.,Peer-reviewed research,Evolution age-0058,Reliability & Durable Execution,Resilient topology,Paper,ResMAS: Resilience Optimization in LLM-based Multi-agent Systems,https://ojs.aaai.org/index.php/AAAI/article/view/40824,AAAI,2026,Zhilun Zhou et al.,Learns task-specific resilient communication topologies and topology-aware prompts after measuring how graph structure and node instructions affect performance under agent failures and other perturbations.,Moves resilience from reactive recovery into the design of the graph itself and evaluates transfer to new tasks and models.,Peer-reviewed research,Reliability age-0059,Research Foundations,System evolution,Paper,EvoMAS: Evolutionary Generation of Multi-Agent Systems,https://openreview.net/forum?id=ic0AGRIkmY,ICML,2026,Yuntong Hu et al.,"Evolves structured multi-agent configurations through trace-guided mutation, crossover, selection, and an experience memory across reasoning, coding, and tool-use tasks.",Shows how an inspectable team specification can evolve from execution evidence while retaining executability and runtime robustness.,Peer-reviewed research,Evolution age-0060,Critiques & Limits,Error propagation,Paper,Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent Systems,https://aclanthology.org/2025.emnlp-main.623/,EMNLP,2025,Xu Shen et al.,Causally studies correct and erroneous information propagation across communication densities and finds that moderately sparse structures can preserve useful diffusion while suppressing errors.,Provides evidence against defaulting to dense graphs and links topology decisions to measured error amplification.,Peer-reviewed research,Topology age-0061,Frameworks & SDKs,Graph-centric orchestration,Paper,MASFactory: A Graph-centric Framework for Orchestrating LLM-Based Multi-Agent Systems with Vibe Graphing,https://aclanthology.org/2026.acl-demo.35/,ACL System Demonstrations,2026,Yang Liu et al.,"Compiles natural-language intent into an editable workflow specification and executable directed graph, with reusable components, topology preview, runtime tracing, multimodal messages, and human interaction.",Provides a direct implementation path from an inspectable organizational graph to execution and evaluates it on seven public benchmarks.,Peer-reviewed research,Work graphs age-0062,Frameworks & SDKs,Directed agent graphs,Docs,AutoGen GraphFlow (Workflows),https://microsoft.github.io/autogen/dev/user-guide/agentchat-user-guide/graph-flow.html,Microsoft AutoGen Documentation,2026,Microsoft,"Implements directed multi-agent execution graphs with sequential, parallel, conditional, fan-in, and cyclic paths, edge conditions, activation groups, safe loop exits, and separately configurable message filtering. The feature is explicitly experimental.","Distinguishes the execution graph from the message graph, exposing both who acts next and what context each agent receives.",Official documentation,Work graphs age-0063,Protocols & Handoffs,Durable remote tasks,Standard,Model Context Protocol: Tasks,https://modelcontextprotocol.io/specification/2025-11-25/basic/utilities/tasks,MCP Specification 2025-11-25,2025,Model Context Protocol project; Agentic AI Foundation,"Specifies experimental durable asynchronous request state machines with capability negotiation, polling, deferred results, progress, input-required states, cancellation, TTLs, and task-message correlation.",Turns a remote tool or context edge into a recoverable task contract that can outlive one synchronous request.,Official documentation,Reliability age-0064,Protocols & Handoffs,Secure agent messaging,Docs,Secure Low-Latency Interactive Messaging (SLIM),https://datatracker.ietf.org/doc/draft-mpsb-agntcy-slim/,IETF Datatracker; individual Internet-Draft,2026,Luca Muscariello; Michele Papalini; Mauro Sardara; Sam Betts,"Proposes a transport layer for A2A and MCP using gRPC over HTTP/2 and HTTP/3 with stream multiplexing, flow control, group communication, native RPC semantics, and MLS end-to-end encryption. It is an individual informational Internet-Draft with no formal IETF standing.",Adds a concrete secure transport substrate for high-volume graph edges that must cross process and organizational boundaries.,Official documentation,Handoffs age-0065,Protocols & Handoffs,Agent identity and authorization,Docs,Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization,https://csrc.nist.gov/pubs/other/2026/02/05/accelerating-the-adoption-of-software-and-ai-agent/ipd,NIST NCCoE Initial Public Draft,2026,Harold Booth; William Fisher; Ryan Galluzzo; Joshua Roberts,"Outlines considerations and open questions for standards-based identity, authorization, auditing, and non-repudiation when software and AI agents access enterprise systems and take actions. The concept paper remains an initial public draft under review.",Grounds graph roles and permissions in identity practice so delegation does not silently transfer more authority than an edge contract allows.,Official documentation,Roles age-0066,"State, Memory & Artifacts",Versioned work products,Docs,Google ADK: Artifacts,https://adk.dev/artifacts/,Google ADK Documentation,2026,Google,"Defines named, automatically versioned binary work products that agents and tools can save, load, list, and exchange within session-scoped or persistent user-scoped namespaces.","Provides explicit, inspectable edge artifacts instead of forcing large or structured outputs through conversational context.",Official documentation,State age-0067,"State, Memory & Artifacts",Procedural memory,Paper,LEGOMem: Modular Procedural Memory for Multi-agent LLM Systems for Workflow Automation,https://www.microsoft.com/en-us/research/publication/legomem-modular-procedural-memory-for-multi-agent-llm-systems-for-workflow-automation/,AAMAS,2026,Dongge Han; Camille Couturier; Daniel Madrigal; Xuchao Zhang; Victor Ruehle; Saravan Rajmohan,Decomposes execution trajectories into reusable procedural memories and allocates them to an orchestrator and specialist agents for workflow automation.,Shows how durable experience can improve decomposition and delegation without collapsing every node's memory into one undifferentiated store.,Peer-reviewed research,State age-0068,Verification & Evals,Adversarial agent detection,Paper,When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems,https://openreview.net/forum?id=BnduUW8izq,ICML,2026,Haowen Xu et al.,"Evaluates activation-space detection and restorative steering of compromised agents across five attack scenarios, multiple models, and synchronous and asynchronous multi-agent interactions.",Tests a topology-agnostic gate for locating and repairing a malicious node when its messages appear superficially benign.,Peer-reviewed research,Reliability age-0069,Reliability & Durable Execution,Intervention-driven debugging,Paper,DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems,https://iclr.cc/virtual/2026/poster/10007537,ICLR,2026,Ming Ma et al.,Tests failure hypotheses by editing messages or plans and measuring whether each intervention repairs the outcome or advances execution in branching multi-agent traces.,Turns graph debugging into causal repair experiments instead of relying on plausible but unverified post-hoc explanations.,Peer-reviewed research,Reliability age-0070,Reliability & Durable Execution,Durable graph workflows,Docs,Microsoft Agent Framework: Durable Extension,https://learn.microsoft.com/en-us/agent-framework/integrations/durable-extension,Microsoft Learn,2026,Microsoft,"Adds persisted sessions, checkpointed progress, failure recovery, external-event waits, and distributed hosting to agents, multi-agent orchestrations, and graph workflows. The documented packages remain prerelease.",Shows how an agent graph can resume without losing context or repeating completed work after interruption.,Official documentation,Reliability age-0071,Critiques & Limits,Control-flow security,Paper,Breaking and Fixing Defenses Against Control Flow Hijacking in Multi-Agent Systems,https://openreview.net/forum?id=PNU9Rj5RDQ,ICLR,2026,Rishi Dev Jha; Harold Triedman; Justin Wagle; Vitaly Shmatikov,"Demonstrates attacks against alignment-check defenses and introduces ControlValve, which generates permitted control-flow graphs and enforces least privilege for each agent invocation.",Makes invocation authority and contextual permissions explicit on every edge instead of trusting a separate checker that can also be hijacked.,Peer-reviewed research,Reliability age-0072,Observability & Cost,Budget-aware topology,Paper,BAMAS: Structuring Budget-Aware Multi-Agent Systems,https://ojs.aaai.org/index.php/AAAI/article/view/40226,AAAI,2026,Liming Yang; Junyu Luo; Xuanzhe Liu; Yiling Lou; Zhenpeng Chen,"Selects an LLM team with integer programming and then learns its collaboration topology under an explicit budget, reporting comparable performance with cost reductions of up to 86% on three tasks.","Jointly engineers node selection, edge structure, and spend instead of optimizing quality while treating inference cost as an afterthought.",Peer-reviewed research,Observability & cost age-0073,Observability & Cost,Workflow reconstruction,Paper,AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction,https://arxiv.org/abs/2602.05353,ICML,2026,Ruijie Shi; Houbin Zhang; Yuecheng Han; Yuheng Wang; Jingru Fan; Runde Yang; Yufan Dang; Huatao Li; Dewen Liu; Yuan Cheng; Chen Qian,Reconstructs an editable explicit stand-in workflow for a black-box agentic system from input-output behavior using iterative search and evaluation.,"Offers a path to inspect, compare, and modify systems whose load-bearing workflow is hidden behind an API.",Peer-reviewed research,Observability & cost age-0074,Benchmarks & Datasets,Distributed coordination,Benchmark,SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems,https://aclanthology.org/2026.acl-long.1354/,ACL,2026,Yuzhe Zhang et al.,"Benchmarks free-form coordination under information silos across 30 exact-answer tasks, three communication protocols, six agent scales, and three models while recording success, tokens, and communication density.",Tests whether more nodes and denser edges overcome distributed information constraints under explicit coordination and cost measures.,Benchmark/dataset,Topology age-0075,Benchmarks & Datasets,Dynamic asynchronous evaluation,Benchmark,Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments,https://openreview.net/forum?id=9gw03JpKK4,ICLR,2026,Romain Froger et al.,"Evaluates agents in asynchronous simulated environments across execution, search, ambiguity, adaptation, temporal reasoning, noise, and agent-to-agent collaboration, with action-level verifiers and structured traces.","Makes delays, environmental change, peer communication, and externally checked writes first-class work-graph conditions.",Benchmark/dataset,Work graphs age-0076,Benchmarks & Datasets,Adversarial multi-agent safety,Benchmark,TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems,https://aclanthology.org/2026.acl-long.1442/,ACL,2026,Ishan Kavathekar; Hemang Jain; Ameya Rathod; Ponnurangam Kumaraguru; Tanuja Ganu,"Provides five scenarios with 300 adversarial instances, six attack types, 211 tools, 100 harmless tasks, and multiple AutoGen and CrewAI interaction configurations, plus an Effective Robustness Score.",Evaluates whether safety controls preserve useful work while attacks exploit the distinctive trust and communication surfaces of a multi-agent graph.,Benchmark/dataset,Reliability age-0077,Benchmarks & Datasets,Failure attribution,Benchmark,Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems,https://aclanthology.org/2026.acl-long.912/,ACL,2026,Mengzhuo Chen et al.,Introduces TraceElephant with full execution traces and reproducible environments for attributing failures to responsible agents and decisive steps.,"Measures causal diagnosis over nodes, messages, plans, and dependencies instead of evaluating only the final team output.",Benchmark/dataset,Observability & cost age-0078,Production Case Studies,Secure delegated access,Blog,Creating AI agent solutions for warehouse data access and security,https://engineering.fb.com/2025/08/13/data-infrastructure/agentic-solution-for-warehouse-data-access/,Engineering at Meta,2025,Can Lin; Uday Ramesh Savagaonkar; Iuliu Rus; Komal Mangtani,"Describes collaborating data-user and data-owner agents, specialized subagents, triage, permission negotiation, human oversight, access budgets, analytical risk rules, output guardrails, traces, and daily regression evaluation.",Shows plural bounded agency governed by independent rule-based gates rather than relying on model judgment for security decisions.,Practitioner analysis,Gates age-0079,Production Case Studies,Staged agent swarm,Blog,How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines,https://engineering.fb.com/2026/04/06/developer-tools/how-meta-used-ai-to-map-tribal-knowledge-in-large-scale-data-pipelines/,Engineering at Meta,2026,Krishna Ganeriwal; Plawan Rath; Ashwini Verma,"Reports a staged swarm of more than 50 explorer, analyst, writer, critic, fixer, upgrader, tester, and gap-filling tasks, with repeated review rounds and recurring automated refresh runs.","Provides a concrete fan-out, critique, repair, integration, and recurrence graph with explicit artifacts and quality gates; efficiency figures are company-reported and preliminary.",Practitioner analysis,Work graphs age-0080,Production Case Studies,Multi-hop identity and provenance,Blog,Solving the Identity Crisis for AI Agents,https://www.uber.com/by/en/blog/solving-the-agent-identity-crisis/,Uber Engineering,2026,Matt Mathew; Prasad Borole; Meng Huang; Sergey Burykin; Gaurav Goel; Bayard Walsh,"Describes an internal agent mesh with registered identities, SPIRE-backed workload attestation, short-lived audience-scoped tokens for every hop, actor-chain provenance, MCP gateway enforcement, and a standardized A2A client.","Treats identity, delegated authority, and provenance as mandatory edge state across a multi-agent graph; adoption and latency metrics are company-reported.",Practitioner analysis,Handoffs age-0081,Research Foundations,Task-adaptive topology synthesis,Paper,Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models,https://aclanthology.org/2026.acl-long.1764/,ACL,2026,Eric Hanchen Jiang et al.,"Introduces Guided Topology Diffusion, which iteratively synthesizes sparse, task-adaptive communication graphs using a proxy model for objectives such as accuracy, utility, and cost.","Treats topology as a multi-objective, per-task design artifact rather than a static collaboration template.",Peer-reviewed research,Evolution age-0082,Reliability & Durable Execution,Topology-conditioned memory leakage,Paper,Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs,https://aclanthology.org/2026.findings-acl.1980/,Findings of ACL,2026,Jinbo Liu et al.,"Measures private-information leakage across six communication topologies, agent counts, and attacker-target placements, finding higher leakage with denser connectivity, shorter graph distance, and greater target centrality.","Makes privacy a measurable consequence of edge structure and node placement, motivating topology-aware access controls.",Peer-reviewed research,Reliability age-0083,Critiques & Limits,Collaboration-induced diversity collapse,Paper,Diversity Collapse in Multi-Agent LLM Systems: Structural Coupling and Collective Failure in Open-Ended Idea Generation,https://aclanthology.org/2026.findings-acl.13/,Findings of ACL,2026,Nuo Chen et al.,"Studies multi-agent ideation across model, cognition, and system levels, finding diminishing group-size returns, authority effects, and faster premature convergence under dense communication.","Shows that interaction can contract a search space, so graph designs for exploration must preserve independence and meaningful disagreement.",Peer-reviewed research,Topology age-0084,Research Foundations,Dynamic node and edge elimination,Paper,AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration,https://aclanthology.org/2025.acl-long.1170/,ACL,2025,Zhexuan Wang; Yutong Wang; Xuebo Liu; Liang Ding; Miao Zhang; Jie Liu; Min Zhang,"Optimizes communication-graph adjacency matrices across collaboration rounds to remove redundant agents and messages, reducing both prompt and completion token consumption in its evaluations.",Makes node participation and edge density explicit cost-quality decisions instead of assuming every role should run on every task.,Peer-reviewed research,Evolution age-0085,Benchmarks & Datasets,Process-level collaboration,Benchmark,Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents,https://aclanthology.org/2025.emnlp-main.249/,EMNLP,2025,Haochen Sun; Shuwen Zhang; Lujie Niu; Lei Ren; Hao Xu; Hao Fu; Fangkun Zhao; Caixia Yuan; Xiaojie Wang,"Provides 30 open-ended cooperative tasks and process-oriented measures for goal interpretation, active collaboration, communication, and continuous adaptation across 13 language models.","Distinguishes reaching an answer from coordinating well, exposing collaboration failures that final-outcome metrics can hide.",Benchmark/dataset,Observability & cost age-0086,Critiques & Limits,Inter-agent communication attacks,Paper,Red-Teaming LLM Multi-Agent Systems via Communication Attacks,https://aclanthology.org/2025.findings-acl.349/,Findings of ACL,2025,Pengfei He; Yuping Lin; Shen Dong; Han Xu; Yue Xing; Hui Liu,"Introduces an Agent-in-the-Middle attack that intercepts and manipulates inter-agent messages, evaluated across multiple frameworks, communication structures, and applications.","Shows that graph edges can compromise the whole system without altering its nodes, making message integrity and provenance part of the handoff contract.",Peer-reviewed research,Handoffs age-0087,Research Foundations,Query-dependent architecture search,Paper,Multi-agent Architecture Search via Agentic Supernet,https://proceedings.mlr.press/v267/zhang25bi.html,ICML,2025,Guibin Zhang; Luyang Niu; Junfeng Fang; Kun Wang; Lei Bai; Xiang Wang,"Represents possible agentic architectures as a probabilistic supernet and samples query-dependent systems with tailored language-model calls, tool calls, and token costs across six benchmarks.",Replaces a one-size-fits-all graph with per-query architecture and resource allocation while keeping the design space explicit.,Peer-reviewed research,Evolution age-0088,Critiques & Limits,Unequal agent contribution,Paper,Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation,https://openreview.net/forum?id=5J6u03ObRZ,ICLR,2026,Zhiwei Zhang et al.,"Identifies lazy-agent behavior in which one role dominates a nominally collaborative system, then measures causal influence and uses verifiable rewards to encourage deliberation and selective reasoning restarts.",Requires builders to verify that each node contributes causal value rather than allowing a multi-agent graph to collapse into one effective agent.,Peer-reviewed research,Roles age-0089,Benchmarks & Datasets,Multi-party negotiation,Benchmark,"Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation",https://openreview.net/forum?id=59E19c6yrN,NeurIPS,2024,Sahar Abdelnabi; Amr Gomaa; Sarath Sivaprasad; Lea Schönherr; Mario Fritz,"Provides scorable, multi-agent, multi-issue negotiation games with metrics for task performance and role alignment, including cooperative, competitive, greedy, and adversarial participants.",Tests communication and collective decisions when graph nodes hold conflicting objectives or attempt manipulation rather than cooperating by default.,Benchmark/dataset,Handoffs age-0090,Reliability & Durable Execution,Agentic application security risks,Standard,OWASP Top 10 for Agentic Applications for 2026,https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/,OWASP GenAI Security Project,2025,OWASP GenAI Security Project; Agentic Security Initiative,"Provides a globally peer-reviewed operational framework for critical risks in autonomous and agentic applications, developed with contributions from more than 100 security experts, researchers, and practitioners.","Supplies a practical threat checklist for node authority, tool permissions, memory provenance, inter-agent trust, and cascading failures across an agent graph.",Industry standard,Reliability age-0091,Research Foundations,Task-adaptive graph routing,Paper,AMAS: Adaptively Determining Communication Topology for LLM-based Multi-agent System,https://aclanthology.org/2025.emnlp-industry.144/,EMNLP Industry Track,2025,Hui Yi Leong; Yuheng Li; Yuqing Wu; Wenwen Ouyang; Wei Zhu; Jiechao Gao; Wei Han,Introduces a dynamic graph selector that uses lightweight LLM adaptation to choose task-specific communication structures and route each query through a matching agent pathway.,"Operationalizes topology as an input-conditioned decision rather than a fixed template; evidence is currently concentrated in question answering, mathematics, and code generation.",Peer-reviewed research,Evolution age-0092,Research Foundations,"Collaboration-mode, role, and model routing",Paper,MasRouter: Learning to Route LLMs for Multi-Agent Systems,https://aclanthology.org/2025.acl-long.757/,ACL,2025,Yanwei Yue; Guibin Zhang; Boyang Liu; Guancheng Wan; Kun Wang; Dawei Cheng; Yiyan Qi,"Defines multi-agent system routing as a joint decision over collaboration mode, role allocation, and language-model selection, implemented with a cascaded controller that progressively constructs a task-specific system.","Makes node roles, model assignment, and whether collaboration is needed explicit routing decisions with measurable cost-quality trade-offs.",Peer-reviewed research,Evolution age-0093,Research Foundations,Multi-hop evidence propagation,Paper,MOC: Multi-Order Communication in LLM-based Multi-Agent Systems,https://openreview.net/forum?id=wyynWicO5s,ICML,2026,Yao Guan; Lin Wang; Zhihui Lu; Ziyi Wang; Wenzhu Yan; Qiang Duan,"Constructs structured multi-order evidence streams so agents can receive relevant information from multiple upstream hops, then applies semantic-topological merging under token constraints.","Shows that graph engineering includes what evidence survives across paths, not only which nodes and edges exist; it improves communication over a supplied topology rather than selecting that topology.",Peer-reviewed research,Handoffs age-0094,Research Foundations,Hierarchical node and topology optimization,Paper,HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent Systems,https://openreview.net/forum?id=p7p5foAXaB,ICML,2026,Tianjun Yao; Zhaoyi Li; Zhiqiang Shen,"Models each functional role as a supernode containing heterogeneous language models in a propose-synthesis structure, then uses multi-level reward attribution and graph classification to select the inter-supernode topology.","Jointly engineers node composition, role capability, credit assignment, and communication edges instead of optimizing each dimension independently.",Peer-reviewed research,Evolution age-0095,"State, Memory & Artifacts",Hierarchical graph memory,Paper,GAM: Hierarchical Graph-based Agentic Memory for LLM Agents,https://aclanthology.org/2026.acl-long.1600/,ACL,2026,Zhaofen Wu; Hanrong Zhang; Fulin Lin; Wujiang Xu; Xinran Xu; Yankai Chen; Henry Peng Zou; Shaowen Chen; Weizhi Zhang; Xue Liu; Philip S. Yu; Hongwei Wang,"Separates rapid memory encoding from stable consolidation through an event-progression graph, a topic-associative network, and graph-guided multi-factor retrieval.","Provides a concrete graph contract for encoding, consolidating, connecting, and retrieving state; it does not address shared-memory permissions or concurrent writes among agents.",Peer-reviewed research,State age-0096,"State, Memory & Artifacts",Selective shared memory,Paper,Learning to Share: Selective Memory for Efficient Parallel Agentic Systems,https://openreview.net/forum?id=cCFyY2LmF5,ICML,2026,Joseph Fioresi; Parth Parag Kulkarni; Ashmal Vayani; Song Wang; Mubarak Shah,Adds a global memory bank to parallel agent teams and trains an admission controller with stepwise reinforcement learning and usage-aware credit assignment to retain reusable intermediate work.,Treats cross-team memory writes as governed graph operations that can reduce duplicate work while limiting indiscriminate context growth.,Peer-reviewed research,State age-0097,Observability & Cost,Graph-aware cache reuse,Paper,Accelerating Language Model Workflows with Prompt Choreography,https://aclanthology.org/2026.tacl-1.13/,TACL,2026,TJ Bai; Jason Eisner,"Executes multi-agent workflows with a dynamic global key-value cache in which each call can attend to a reordered subset of previously encoded messages, including parallel branches.","Makes workflow dependencies and reusable message state explicit execution concerns; its primary result concerns speed, and cache reuse can alter model behavior.",Peer-reviewed research,Work graphs age-0098,Benchmarks & Datasets,Enterprise workflow benchmark,Benchmark,Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments,https://aclanthology.org/2025.emnlp-main.466/,EMNLP,2025,Harsh Vishwakarma; Ankush Agarwal; Ojas Patil; Chaitanya Devaguptapu; Mahesh Chandran,"Introduces EnterpriseBench with 500 tasks across software engineering, HR, finance, and administration, including fragmented data, access-control hierarchies, and cross-functional workflows.","Tests work graphs that retrieve, modify, and transfer artifacts across services while respecting authority boundaries; the organization and services are simulated.",Benchmark/dataset,Work graphs age-0099,Benchmarks & Datasets,Stateful tool-use evaluation,Benchmark,"ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities",https://aclanthology.org/2025.findings-naacl.65/,Findings of NAACL,2025,Jiarui Lu; Thomas Holleis; Yizhe Zhang; Bernhard Aumayer; Feng Nan; Haoping Bai; Shuang Ma; Shen Ma; Mengyu Li; Guoli Yin; Zirui Wang; Ruoming Pang,"Evaluates stateful tool execution, implicit dependencies between tools, on-policy user interaction, and intermediate and final milestones over arbitrary trajectories.","Transfers milestone and state-transition testing to graph prerequisites, although the evaluated system is an agent-user-tool loop rather than an agent-agent topology.",Benchmark/dataset,State age-0100,Benchmarks & Datasets,Trajectory-level evaluation,Benchmark,AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents,https://openreview.net/forum?id=4S8agvKjle,NeurIPS Datasets & Benchmarks,2024,Chang Ma; Junlei Zhang; Zhihao Zhu; Cheng Yang; Yujiu Yang; Yaohui Jin; Zhenzhong Lan; Lingpeng Kong; Junxian He,"Unifies partially observable, multi-round agent environments and adds a fine-grained progress-rate metric plus interactive trajectory analysis beyond final success.","Supplies milestone-level evaluation that can localize incomplete work, while its primarily single-agent tasks do not identify the responsible graph component by themselves.",Benchmark/dataset,Gates age-0101,Benchmarks & Datasets,Agent security benchmark,Benchmark,Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents,https://openreview.net/forum?id=V4y0CpX4hK,ICLR,2025,Hanrong Zhang; Jingyuan Huang; Kai Mei; Yifei Yao; Zhenting Wang; Chenlu Zhan; Hongwei Wang; Yongfeng Zhang,"Benchmarks prompt injection, memory poisoning, backdoors, mixed attacks, and defenses across 10 scenarios, more than 400 tools, 13 model backbones, and seven metrics.","Maps security failures to prompts, plans, tools, and memory surfaces that graph designs must isolate separately; inter-agent trust propagation is outside its main evaluation target.",Benchmark/dataset,Reliability age-0102,Benchmarks & Datasets,Prompt-injection evaluation,Benchmark,AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,https://openreview.net/forum?id=m1YYAQjO3w,NeurIPS Datasets & Benchmarks,2024,Edoardo Debenedetti; Jie Zhang; Mislav Balunovic; Luca Beurer-Kellner; Marc Fischer; Florian Tramèr,Provides an extensible adversarial environment with 97 realistic tasks and 629 security test cases for agents executing tools over untrusted data.,Tests the instruction-data boundary that every external-input and artifact edge must preserve; compromised peer-agent messages require separate evaluation.,Benchmark/dataset,Reliability age-0103,Benchmarks & Datasets,Agent misuse benchmark,Benchmark,AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents,https://openreview.net/forum?id=AC5n7xHuR1,ICLR,2025,Maksym Andriushchenko; Alexandra Souly; Mateusz Dziemian; Derek Duenas; Maxwell Lin; Justin Wang; Dan Hendrycks; Andy Zou; J. Zico Kolter; Matt Fredrikson; Yarin Gal; Xander Davies,Evaluates refusal and retained task capability on 110 explicitly malicious multi-stage agent tasks with 440 augmented variants across 11 harm categories.,Tests whether safety gates reject prohibited work and stop a compromised node from completing a harmful path; the score is not a general measure of agent-system safety.,Benchmark/dataset,Gates age-0104,Reliability & Durable Execution,Structural prompt-injection defense,Paper,IPIGuard: A Novel Tool Dependency Graph-Based Defense Against Indirect Prompt Injection in LLM Agents,https://aclanthology.org/2025.emnlp-main.53/,EMNLP,2025,Hengyu An; Jinghuai Zhang; Tianyu Du; Chunyi Zhou; Qingming Li; Tao Lin; Shouling Ji,Models task execution as traversal over a planned Tool Dependency Graph and separates action planning from interaction with untrusted external data.,Shows how structural constraints on allowed tool edges can block unauthorized actions instead of relying only on prompts or classifiers.,Peer-reviewed research,Work graphs age-0105,Protocols & Handoffs,Agent identity and delegated authorization,Docs,"Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world",https://openid.net/wp-content/uploads/2025/10/Identity-Management-for-Agentic-AI.pdf,OpenID Foundation,2025,OpenID Foundation Artificial Intelligence Identity Management Community Group,"Maps identity, authentication, delegated authority, consent, governance, and accountability gaps for autonomous agents against OAuth and OpenID Connect foundations.","Provides a standards-grounded vocabulary for identity-bearing nodes and auditable delegation across agent and tool edges; it is a non-normative landscape whitepaper, not a protocol.",Official documentation,Handoffs age-0106,Protocols & Handoffs,Agent capability and discovery schema,Standard,Open Agentic Schema Framework (OASF) v1.1.0,https://github.com/agntcy/oasf/releases/tag/v1.1.0,"AGNTCY Project, Linux Foundation",2026,AGNTCY Project,"Defines versioned schemas and taxonomies for agent identity metadata, capabilities, skills, modules, interactions, and discovery records, including MCP and A2A alignment modules.",Gives graph builders a portable node descriptor for capability-based discovery and routing validation; it describes capabilities but does not execute a graph or authorize an edge.,Industry standard,Roles age-0107,"State, Memory & Artifacts",Artifact provenance,Standard,PROV-DM: The PROV Data Model,https://www.w3.org/TR/prov-dm/,W3C Recommendation,2013,Luc Moreau; Paolo Missier,"Defines entities, activities, agents, derivation, attribution, responsibility, bundles, and collections for interoperable provenance records.","Maps artifacts, execution steps, and responsible nodes into a portable lineage model; storage, access control, and cryptographic integrity remain implementation concerns.",Industry standard,State age-0108,"State, Memory & Artifacts",Cryptographic artifact provenance,Paper,in-toto: Providing farm-to-table guarantees for bits and bytes,https://www.usenix.org/conference/usenixsecurity19/presentation/torres-arias,USENIX Security,2019,Santiago Torres-Arias; Hammad Afzali; Trishank Karthik Kuppusamy; Reza Curtmola; Justin Cappos,"Introduces signed material and product links for supply-chain steps and verifies them against an expected layout, supported by analysis of 30 historical compromises and deployed integrations.","Supplies a transferable receipt pattern that binds each artifact transformation to an authorized step and identity; it verifies process integrity, not semantic correctness.",Peer-reviewed research,State age-0109,Reliability & Durable Execution,Portable workflow semantics,Standard,Serverless Workflow Specification v1.0.0,https://serverlessworkflow.io/blog/releases/release-100/,"Serverless Workflow, Cloud Native Computing Foundation",2025,Serverless Workflow Project,"Defines a vendor-neutral workflow language with sequential and concurrent tasks, event correlation, service calls, error handling, retries, and timeouts.","Provides an inspectable execution contract for production work graphs; conformance does not itself supply a durable runtime, an agent protocol, or model-level correctness.",Industry standard,Reliability age-0110,Critiques & Limits,Compound-system call scaling,Paper,Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI Systems,https://proceedings.neurips.cc/paper_files/paper/2024/hash/51173cf34c5faac9796a47dc2fdd3a71-Abstract-Conference.html,NeurIPS,2024,Lingjiao Chen; Jared Davis; Boris Hanin; Peter Bailis; Ion Stoica; Matei Zaharia; James Zou,Analyzes Vote and Filter-Vote compound systems and finds that performance can rise and then fall as the number of language-model calls increases because query difficulty is heterogeneous.,"Bounds claims that adding calls, voters, or nodes automatically improves a system; the experiments cover simple aggregation systems rather than rich agent organizations.",Peer-reviewed research,Observability & cost age-0111,Verification & Evals,Verification-aware work graphs,Paper,Verification-Aware Planning for Multi-Agent Systems,https://aclanthology.org/2026.eacl-long.353/,EACL,2026,Tianyang Xu; Dan Zhang; Kushan Mitra; Estevam Hruschka,"Introduces VeriMAP, which decomposes tasks into a dependency graph and attaches planner-defined Python and natural-language verification functions to subtasks before execution.",Makes acceptance criteria part of the work graph so handoff failures can trigger local repair instead of surfacing only in a final answer.,Peer-reviewed research,Gates age-0112,Research Foundations,Human-in-the-loop work-graph planning,Paper,AIPOM: Agent-aware Interactive Planning for Multi-Agent Systems,https://aclanthology.org/2025.emnlp-demos.7/,EMNLP System Demonstrations,2025,Hannah Kim; Kushan Mitra; Chen Shen; Dan Zhang; Estevam Hruschka,"Presents conversational and graph-based interfaces for inspecting, editing, and collaboratively guiding plans in orchestrated multi-agent systems.","Treats the work graph as a human-legible control surface; its non-autonomous agents execute assigned tasks, so it is a planning substrate rather than a complete agent organization.",Peer-reviewed research,Work graphs age-0113,Frameworks & SDKs,Graph workflow runtime,Blog,Build reliable multi-agent applications with ADK Go 2.0,https://developers.googleblog.com/announcing-adk-go-20/,Google Developers Blog,2026,Toni Klopfenstein; Sampath Kumar Maddula,"Documents a graph-based multi-agent workflow engine with typed nodes, conditional edges, fan-out and fan-in, nested graphs, cycles, retries, durable human input, state, and unified telemetry.","Provides a first-party implementation of explicit, resumable work graphs while showing that function and tool nodes remain distinct from agent nodes.",Practitioner analysis,Work graphs age-0114,Research Foundations,Decentralized adaptive coordination,Paper,AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems,https://proceedings.neurips.cc/paper_files/paper/2025/hash/9a379c1b05793d1c42dc832269834515-Abstract-Conference.html,NeurIPS,2025,Yingxuan Yang; Huacan Chai; Shuai Shao; Yuanyi Song; Siyuan Qi; Renting Rui; Weinan Zhang,"Organizes specialized agents in a decentralized DAG, uses retrieval-backed local experience for capability refinement, and adapts routing without a single central orchestrator.",Offers a contrasting topology for settings where central coordination is a bottleneck or trust boundary; its reported evidence is concentrated in reasoning and coding tasks.,Peer-reviewed research,Topology age-0115,Reliability & Durable Execution,Effect-typed agent contracts,Paper,ETAS: An Effect-Typed Language for Agent Systems,https://arxiv.org/abs/2607.17780,arXiv,2026,Huiri Tan; Yikun Wang; Puyang Zhang; Shangyu Li; Jiasi Shen,"Defines a language in which agents, tools, typed memory, approvals, policies, effects, and execution traces are semantic program elements, with static obligations and runtime monitors.",Provides a transferable contract model for authorizing and auditing node actions before and during execution; it is a new preprint and not yet deployment evidence.,Research preprint,Reliability age-0116,Critiques & Limits,Consensus-induced search collapse,Paper,The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search,https://arxiv.org/abs/2607.18045,arXiv,2026,Yohei Nakajima,Uses an exactly solvable multi-searcher benchmark to show that pooling information into one repeated recommendation can improve belief accuracy while sharply reducing collective discovery coverage.,Separates information quality from allocation policy and warns that consensus edges can collapse parallel exploration unless a coordinator preserves a portfolio of actions.,Research preprint,Topology age-0117,Critiques & Limits,Relay information bottlenecks,Paper,When Do Multi-Agent Systems Help? An Information Bottleneck Perspective,https://arxiv.org/abs/2607.16133,arXiv,2026,Wendi Yu; Lianhao Zhou; Xiangjue Dong; Sai Sudarshan Barath; Declan Staunton; Byung-Jun Yoon; Xiaoning Qian; James Caverlee; Shuiwang Ji,Formalizes bounded inter-agent relays as an information bottleneck and reports 18 controlled experiments across five benchmarks and three model scales.,"Explains when isolated contexts and compressed handoffs save useful context and when they discard task-relevant information, providing a direct test for whether a graph earns its edge losses.",Research preprint,Handoffs age-0118,Critiques & Limits,Critique uptake failure,Paper,Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning,https://arxiv.org/abs/2607.15388,arXiv,2026,Chih-Hsuan Yang; Jingyan Jiang; Vikram Vasudevan; Cheng-Hau Yang; Huihuo Zheng; Le Chen; Eliu A. Huerta; Venkatram Vishwanath; Ian T. Foster; Rajeev Thakur,"Evaluates 4,181 verifier-grounded math problems and finds that a precise reviewer can still yield weak repair when the protocol does not carry useful critique into the next candidate.",Shows that a verifier node is insufficient by itself: the outgoing edge must couple accepted evidence to a concrete state change or retry action.,Research preprint,Gates