Single-agent systems are the default. Multi-agent systems split work across specialized agents. Sometimes that's worth the coordination overhead; sometimes it's premature.
When single-agent suffices
If one LLM with a focused tool set can handle the task:
- Customer support agent with 10 specific tools.
- Coding assistant with read/write file tools.
- Research assistant that searches, summarizes, and writes.
Single-agent is simpler, cheaper, easier to debug. Default to it.
When multi-agent makes sense
Case 1 — Specialized roles
Different parts of the task require fundamentally different prompts/tools:
- Researcher specializes in finding info.
- Writer specializes in synthesizing into prose.
- Reviewer specializes in critique.
Each agent has tightly scoped prompts and tools. Coordinator orchestrates.
Case 2 — Privilege separation
- Public agent answers user questions (limited tools).
- Admin agent has access to dangerous tools (delete, modify).
- Routing between them controls access.
Case 3 — Parallel work
Independent sub-tasks run in parallel:
- Compose-email task fans out to "research recipient", "draft body", "find references".
- Coordinator merges results.
Case 4 — Model heterogeneity
Different sub-tasks use different models:
- Frontier model for hard reasoning.
- Cheap model for triage.
- Fine-tuned model for specialized task.
Each agent uses the right model for its job.
When NOT to use multi-agent
- The "I want it to feel sophisticated" justification.
- A single agent with good prompting could do it.
- The tools are not actually role-specific.
- You haven't shipped a single-agent version first.
Multi-agent adds coordination cost. If single-agent is "good enough", that's the answer.
Pattern 1 — Supervisor (or Coordinator)
One agent dispatches to specialized sub-agents:
[Supervisor]
/ | \
[Research] [Write] [Review]
The supervisor sees the user's request, decides which sub-agent handles it (or which order to use them in), aggregates results.
LangGraph implementation:
from langgraph.graph import StateGraph
from langgraph.prebuilt import create_react_agent
# Specialized agents
research_agent = create_react_agent(llm, tools=[search_web, summarize])
writer_agent = create_react_agent(llm, tools=[draft_article])
reviewer_agent = create_react_agent(llm, tools=[critique])
# Supervisor decides routing
def supervisor(state):
response = supervisor_llm.invoke([
SystemMessage("You coordinate research, writing, and review. Output next agent: research, write, review, or FINISH."),
*state["messages"]
])
return {"next_agent": response.content.strip()}
def route_to_next(state):
if state["next_agent"] == "FINISH":
return END
return state["next_agent"] # research, write, or review
builder = StateGraph(MultiAgentState)
builder.add_node("supervisor", supervisor)
builder.add_node("research", research_agent)
builder.add_node("write", writer_agent)
builder.add_node("review", reviewer_agent)
builder.set_entry_point("supervisor")
builder.add_conditional_edges("supervisor", route_to_next)
builder.add_edge("research", "supervisor")
builder.add_edge("write", "supervisor")
builder.add_edge("review", "supervisor")
Each sub-agent returns control to the supervisor. The supervisor decides what to do next.
Pattern 2 — Hierarchical
Supervisors themselves can be supervised:
[Director]
/ \
[Research Team] [Writing Team]
/ | \ / \
[Web] [DB] [Files] [Draft] [Edit]
Each layer has its own coordinator. Useful for very complex agents but adds layers of latency.
Pattern 3 — Network / Mesh
Agents communicate freely without a central supervisor:
[A] ↔ [B]
↓ ↓
[C] ↔ [D]
Less common. Harder to control. Used in some research/simulation scenarios.
Pattern 4 — Plan and execute
One agent plans; others execute:
[Planner] → outputs a step-by-step plan
↓
[Executor 1] → step 1
↓
[Executor 2] → step 2
↓
...
Variant: re-plan after each step if conditions change.
Used for complex multi-step tasks (e.g., "research X and write a report").
How to design multi-agent state
State must be shared across agents. Two options:
Shared messages
All agents append to the same message list. Each agent sees the full conversation.
Pros: simple, all context available. Cons: messages grow long, expensive.
Scoped messages
Each agent has its own message scope. Supervisor passes specific info between them.
Pros: cheaper, focused contexts. Cons: requires explicit data passing.
For most systems, shared messages with smart trimming (summarize older messages).
Common multi-agent mistakes
- Premature multi-agent. Single agent would work; coordination overhead unjustified.
- No clear roles. Multiple agents that overlap → confusion.
- Tight coupling. Agents that constantly need to negotiate → death spiral.
- No supervisor. Anarchy. Always have an orchestrator (or a clear hand-off protocol).
- Same model for all. Defeats the purpose of role specialization.
Cost reality
Multi-agent systems typically cost 2-5x more than equivalent single-agent (more LLM calls). Justify with concrete benefits:
- "Better separation of concerns" — usually not enough.
- "Different sub-tasks need different models for cost/quality" — usually enough.
- "Privilege separation is required" — definitely enough.
Production checklist
- Justified multi-agent over single-agent with concrete reason.
- Clear role per agent (write down "what does each agent do?").
- Defined hand-off protocol.
- Supervisor / orchestrator in place.
- Max iterations cap (very important — agents calling each other can spiral).
- Observability per agent (LangSmith).
- Cost monitoring per agent.
Takeaway
Multi-agent for specialized roles, privilege separation, parallel work, or model heterogeneity. Single-agent otherwise. Supervisor pattern is the default; hierarchical for very complex; network/mesh rare. Cost is 2-5x single-agent; justify with concrete benefit. Cap iterations always.
Production Deep Dive: Multi-Agent Supervisors vs Autonomous Swarms
Two patterns dominate multi-agent design in 2026:
- Hierarchical Supervisor (
langgraph-supervisor): A centralized coordinator acts as the router, inspecting user intent, delegating to specialized workers, and synthesizing the final response. Recommended for structured enterprise business processes. - Decentralized Swarms / Handoffs: Workers communicate peer-to-peer, calling
transfer_to_another_agent()directly without returning to a central manager. Recommended for dynamic, exploratory conversations. Always configureoutput_mode="last_message"on subgraphs to avoid flooding the parent supervisor context with hundreds of intermediate scratchpad tokens.