Skip to content
Janeiro.ai
Pattern

Multi-Agent Supervisor Without Cosplay

Add a supervisor only when roles are real: typed handoffs, summarized worker output, cost caps, and one finish path—not a committee of chatbots in prod.

5 min read

A supervisor is a router with a budget, not a committee. You add agents when specialization or parallelism is real: a researcher that only searches, a coder that only patches, a critic that only scores. You do not add agents because a slide said “multi-agent.” Every extra role is extra latency, extra tokens, and another place for prompt injection to hide.

Most teams should stay on a single ReAct loop or a plan-and-execute graph until a named bottleneck appears. This pattern is how you add a supervisor without cosplay.

Context

2025–2026 frameworks made it easy to spawn roles. Parallel researchers can outperform a single agent on breadth, and they multiply spend. That cost lands harder when inference is a Brazil region or a self-hosted GPU. A supervisor plus two workers is three round-trips before the user sees a sentence.

AutoGen and the AutoGen paper showed conversational multi-agent setups. CrewAI packaged role-play for builders. LangGraph multi-agent made supervisors a graph node. Anthropic’s multi-agent research write-up is unusually honest about cost: parallel researchers can outperform a single agent on breadth, and they multiply spend.

That cost lands harder when inference is a Brazil region or a self-hosted GPU—see cost-aware model routing under FX pressure. If each worker also runs a tool loop, you have nested budgets you did not draw on the slide.

The production contract is typed handoffs, summarized worker output, a single finish path, and a rule for when to refuse the architecture. Market demos of “agent swarms” are usually not that contract—see the latest Radar.

The pattern

One supervisor owns the user thread. It emits a structured route: worker name or FINISH. Workers see a narrow brief and a narrow tool allow-list. They return a summary plus artifacts, not their full transcript. The supervisor reconciles and either routes again or finishes. Caps apply to supervisor turns and to each worker.

User → supervisor
  ├ research → researcher + search tools → summarize artifacts
  ├ code     → coder + repo tools        → summarize artifacts
  └ FINISH or budget → stop + reason
Workers do not talk to each other by default.

Typed handoffs

The route object is structured output: next: research | code | FINISH, a one-paragraph brief, and success criteria. Workers do not parse a novel. If the supervisor cannot name the worker and the criterion, you do not have a role boundary. You have two chats.

Workers do not gossip

Peer-to-peer agent chat is how context explodes and how injected text hops roles. Prefer star topology: supervisor in the middle. If two workers must share a file, share an artifact store with ACL, not a running dialogue. Each worker’s tool registry is a subset. The coder does not get refund tools.

When to use

Use a supervisor when (1) tools and prompts genuinely differ by role, or (2) you can run workers in parallel on independent evidence. Skip it when one allow-list and one loop would do. Do not use it to simulate a company org chart. Do not use it as a substitute for evals.

  • Ship now if you already have two specialists and they share one bloated prompt.
  • Prefer one loop if the “workers” would call the same tools.
  • Prefer plan-and-execute if the phases are a list, not competing specialties.

Implementation notes

Implement the supervisor as a graph node with a checkpointer. Persist the brief and the artifact ids, not the worker’s chain-of-thought. Budget in the same units as cost-aware model routing: max supervisor turns, max workers per turn, max spend.

Route object

class Route(BaseModel):
    next: Literal["research", "code", "FINISH"]
    brief: str
    success: str

def run(user_task, budget) -> Result:
    state = State(task=user_task, artifacts=[])
    for _ in range(budget.max_supervisor_turns):
        route = supervisor.route(state)  # structured
        if route.next == "FINISH" or budget.spent():
            return finalize(state)
        out = workers[route.next].run(route.brief, tools=ACL[route.next])
        state.artifacts.append(summarize(out, max_tokens=400))
    return stop("supervisor_budget", state)

Summarize aggressively

Worker transcripts will contain retrieved personal data and injected instructions. Summaries should carry citations and artifact pointers, not raw chunks. If the researcher used RAG, apply the same minimization you would in a single-agent path. The supervisor must not become a second index of everything every worker saw. Pair with LGPD-aware data paths.

Failure modes

  • Role-play without tool split. Three personas, one registry. You paid for theater.
  • Unbounded debate. Supervisor and critic loop. Cap turns; require FINISH.
  • Full transcript handoff. Context window and leak surface grow without bound.
  • Nested ReAct without budgets. Each worker is an unbounded loop. Nested caps.
  • Supervisor as bottleneck. Parallel workers wait on a single slow model. Measure queue depth.
  • Framework default gossip. AutoGen-style group chat enabled because it was the sample.

Trade-offs

Parallel specialists can beat a single agent on broad research. They cost more and fail in more places. A cheaper pattern is often a planner plus one executor, or one ReAct loop with a tight registry. Add agents when you can name the metric that improved: recall on a research pack, or time-to-patch on a repo—not “autonomy.”

Supervisor quality dominates. A weak router sends the coder a research brief and burns the budget. Eval the route object as its own fixture, independent of worker quality.

  • Star supervisor — Clear ownership and thinner leaks. Supervisor latency.
  • Parallel workers — Breadth on research. N-times spend.
  • Single ReAct loop — Simple SLO and traces. Weaker specialization.
  • Group-chat agents — Easy demo. Gossip, cost, injection surface.

A supervisor is a router with a budget. Extra agents pay latency and tokens; they only pay off when specialization or parallelism is real.

What to do next

  • Write the route schema and the per-worker tool ACL before adding a second model role.
  • Cap supervisor turns and worker spend separately.
  • Eval routes as fixtures: wrong worker and missing FINISH should fail CI via eval gates for shipping.
agentstoolsevals

Published by . Original editorial for operators. How this was made