Plan-and-Execute Instead of Wandering
Plan once, execute cheaply: JSON steps with success checks, a smaller executor, and replan only on failed gates—not an unbounded ReAct wander in prod.
When the phases are knowable, do not rediscover them every turn. Plan-and-execute writes an explicit plan—usually JSON steps with success checks—then runs a cheaper executor per step. Replan only when a check fails. That is the correct default for runbooks, research packs, and multi-file refactors. ReAct remains the default when the next hop is unknown.
Yao et al.’s ReAct paper optimized for interleaving thought and action in uncertain environments. Plan-and-execute optimizes for the opposite: you already know the outline, and you do not want to pay a frontier model to remember it on step seven.
Context
LangChain popularized the split: a planner produces steps; an executor (often a smaller model plus tools) works them. The split matters on BRL-priced regional APIs and on self-hosted weights. The failure mode is a confident bad plan that an executor diligently makes worse.
LangChain documented the split in plan-and-execute agents and later in LangGraph’s tutorial: a planner produces steps; an executor works them. LangGraph’s agentic concepts treat this as a graph you can checkpoint, not a hidden while-loop.
The split matters on BRL-priced regional APIs and on self-hosted weights—see cost-aware model routing under FX pressure. A planner call plus four cheap executor calls is often cheaper and more inspectable than eight ReAct turns on the large model. It is also easier to eval: each step has a success predicate you can fixture, which is how eval packs before the POC become a gate instead of a story.
The failure mode is a confident bad plan. The planner omits a rollback step; the executor diligently makes production worse. That is why plans are structured objects with checks, why high-risk steps still interrupt, and why replanning is a counted event—not an infinite “try something else.”
The pattern
The planner returns a versioned list of steps: id, action, success check, risk class. The executor runs one step at a time with a narrow tool allow-list. On success, continue. On failure, replan from remaining work with a bounded retry. On high risk, HITL before the executor. Finish when checks pass or the replan budget is gone.
Goal → planner JSON
→ executor step
├ check pass → next step or finish
├ check fail → replan remaining (bounded)
└ high risk → HITL
The plan is data
Use structured outputs for the plan object. A markdown checklist the executor must interpret is a prompt, not a plan. Each step names the tool class it may use. The executor does not get the whole registry “in case.”
Versus a supervisor
A supervisor routes among specialists when roles compete. Plan-and-execute sequences phases when the order is mostly known. If you need both—plan the research pack, then a specialist coder—compose them. Do not spawn a supervisor to walk a linear runbook.
When to use
Use plan-and-execute for work you could write as a runbook: gather sources, draft, cite, publish; or drain a queue, patch, verify, open the PR. Skip it for “see what the ticket is about” where the next tool depends on the last observation. Skip it when the planner would be guessing at a domain you have not encoded.
- Ship now if ReAct is rediscovering the same four steps on every run and the bill shows it.
- Stay on ReAct if success checks would be fake (“looks good”).
- Add HITL on any step that writes production or money, regardless of the plan’s confidence.
Implementation notes
Checkpoint the plan and the step index. Resume must not re-run a successful write. Route the planner to a stronger model and the executor to a cheaper one; that is the whole point of model routing. Tool arguments still go through tool contracts and least-privilege tool use.
Plan and check
class Step(BaseModel):
id: str
action: str
success: str
risk: Literal["low", "high"]
class Plan(BaseModel):
version: int
steps: list[Step]
def run(goal, budget) -> Result:
plan = planner(goal) # structured
replans = 0
for step in plan.steps:
if step.risk == "high":
interrupt(step)
out = executor(step, tools=tools_for(step))
if not check(step.success, out):
replans += 1
if replans > budget.max_replan:
return stop("replan_budget")
plan = planner.replan(goal, done=completed, failed=step)
return finish(plan)
Success checks must be executable
“User is happy” is not a check. “Tests pass,” “citation count ≥ 3,” “ticket status is pending-review” are. If you cannot write the check without a model, you do not have plan-and-execute. You have ReAct with extra ceremony. Model-as-judge checks belong in evals, not in the hot path, unless you accept the circularity.
Failure modes
- Narrative plans. The executor interprets prose. You are back to prompting.
- No replan cap. A bad planner burns the month.
- Executor inherits all tools. The plan’s narrowness was cosmetic.
- Replay without idempotency. Resume double-applies a step. Unbounded consumption plus a broken ledger.
- Planner overconfidence. Missing rollback. Human checkpoint on high-risk prefixes.
- Using it for exploration. The plan is fiction by step two. Switch to ReAct.
Trade-offs
You gain inspectability and cheaper steps. You lose adaptability mid-flight unless you replan, and replans are where cost returns. A rigid plan on an exploratory task is worse than ReAct. A ReAct loop on a stable runbook is worse than this pattern.
Planner quality dominates. Invest evals in plan completeness (missing rollback, missing verify) more than in executor prose. A perfect executor on a bad plan is a fast incident. Put those fixtures in eval gates for shipping.
- JSON plan + cheap executor — Lower token spend and checkable steps. Bad plans execute quickly.
- Bounded replan — Recovery without a wander. Some tasks abort honestly.
- ReAct — Adapts when the outline is unknown. Rediscovery cost.
- Supervisor of specialists — Parallel roles. Overkill for a linear runbook.
When the phases are knowable, a planner plus cheap executors beats a thought-act loop that rediscovers the outline every turn.
What to do next
- Write the plan schema with executable success checks before adding a planner model.
- Give the executor a per-step tool subset, not the full registry.
- Cap replans and checkpoint step ids so resume cannot double-write. Escalate high-risk steps through human-in-the-loop.
Published by Janeiro.ai. Original editorial for operators. How this was made