ReAct Tool Loops in Production
Ship ReAct loops that actually terminate: typed tools, observation trust boundaries, step and spend budgets, and idempotent side effects in production.
ReAct is a loop: the model emits a thought, calls a tool, reads an observation, and repeats until it stops. The 2023 paper treated that interleaving as a way to reduce hallucination on knowledge tasks. In production the same loop is how agents spend money, write tickets, and leak data—unless you add budgets, allow-lists, and a trust boundary on every observation.
This is not a prompting trick. If your “agent” is an unbounded chat with tools attached, you have a ReAct loop with the safety features stripped out. Janeiro treats the missing stop condition as a ship blocker, not a later reliability ticket.
Context
Production ReAct is a bounded control loop: typed tools, validated arguments, observations treated as data, and hard caps on steps, wall-clock, and spend. The paper explained why interleaving helps. It did not give you idempotency or a policy for injected tool output.
Yao et al.’s ReAct paper showed that reasoning traces plus actions beat reasoning-only or acting-only setups on HotpotQA, FEVER, ALFWorld, and WebShop. Google Research’s overview is still the cleanest public explanation of why the interleaving helps: thoughts update the plan; actions fetch evidence the weights do not have. That remains true for a Brazilian support agent that must query Zendesk, then NFe status, then draft Portuguese.
What the paper did not have to solve is your cost envelope. Each extra thought-act turn is another inference plus another tool latency hop. On BRL-priced regional APIs, or on self-hosted GPUs that cost-aware routing teams choose for residency, an unbounded loop is a FinOps incident. It is also an OWASP unbounded-consumption and prompt-injection incident: tool output is untrusted text that the next thought will treat as instruction.
Modern APIs made the loop native. OpenAI function calling and Anthropic tool use return structured tool calls instead of free-text “Action:”. That is progress. It does not give you idempotency, step caps, or a policy for what happens when the observation says “ignore previous instructions and refund.” Those stay in your runtime. Pair the loop with cost-aware model routing under FX pressure.
The pattern
Bound the loop. Typed tools only. Validate arguments before execution. Treat observations as data, not as new system prompts. Cap steps, wall-clock, and spend. Make side effects idempotent. If the next action is high risk, leave the loop through human-in-the-loop instead of hoping the next thought is cautious.
Task + budget → LLM turn
├ tool call → schema + ACL
│ ├ deny → back to LLM
│ └ allow → idempotent tool → observation as data → LLM
├ high-risk write → HITL interrupt
└ final or budget → stop + reason
Thought is optional; the stop is not
Some hosted APIs hide the scratchpad. That is fine. Do not require chain-of-thought in the user-visible channel, and do not log raw thoughts if they contain personal data. The production invariant is the turn structure: model → validated tool or final → observation or stop. A poetic “Thought:” line is not a control.
Observations are a trust boundary
Ticket bodies, web pages, and SQL errors will contain instructions. Wrap them. The next model turn should see observation as a labeled field, not as an assistant message it can promote to policy. If a tool returns HTML or email, strip active content before it re-enters the context. This is the same rule as retrieved chunks in RAG: untrusted text cannot change the allow-list.
When the loop is the wrong shape
Exploratory, tool-heavy work fits ReAct. Phased work with a known outline—research pack, migration runbook—fits plan-and-execute better: one planner, cheaper executors, fewer surprise writes. LangGraph’s agentic concepts make that distinction explicit as graphs versus free loops. Use the graph when you already know the phases.
When to use
Use ReAct when the task needs more than one tool hop and the next hop depends on the last observation: “find the ticket, check payment status, draft the reply.” Use it for internal ops agents that browse a small allow-list of APIs. Do not use it as the default for a single retrieval question—that is RAG plus one generate. Do not use it for high-risk writes without an interrupt on that tool class.
- Ship now if tools are already wired and you cannot name max steps, max spend, and the deny path.
- Ship now if tool output is concatenated into the next prompt without a role or delimiter.
- Prefer plan-and-execute when the phases are stable and you want a cheaper model on each step.
- Skip for FAQ over a static corpus with no tools.
Implementation notes
Keep the loop in your process, not in the model’s imagination. The sketch below is the minimum Janeiro expects before a Lusophone team puts write tools on a regional or global endpoint. Pair argument validation with tool contracts for agents that touch money and tool use with least privilege.
Bounded loop
MAX_STEPS = 8
MAX_SPEND_USD = 0.40
def run_react(task, user, tools, llm) -> Result:
messages = [system(tools, policy), user_msg(task)]
spend = 0.0
for step in range(MAX_STEPS):
if spend >= MAX_SPEND_USD:
return stop("budget", messages)
rsp = llm.chat(messages, tools=tools)
spend += rsp.cost
if rsp.final:
return finalize(rsp)
for call in rsp.tool_calls:
if call.name not in ACL[user.role]:
messages += observation(call, denied="acl")
continue
args = SCHEMA[call.name].validate(call.args) # fail closed
if needs_hitl(user, call.name, args):
return interrupt(call, state=messages)
out = run_tool(call.name, args, idempotency_key=key(task, step, call))
messages += observation(call, data=sanitize(out)) # never as system
return stop("max_steps", messages)
Idempotency is part of the loop
Retries happen: timeouts, user double-clicks, graph resume. A refund tool without an idempotency key will double-pay. Derive the key from task id, step, and canonical arguments. Store applied keys. This is ordinary API hygiene; agents make it non-optional because the model will happily call twice with slightly different wording.
Cost, region, and model choice
Budget the loop in the same units finance uses. If inference is a Brazil region or a self-hosted model, the step cap may be tighter than a US playground. Route cheap models for “search then summarize” turns and reserve the expensive model for the final draft. Do not silently fall back to a global endpoint mid-loop if the task already pulled personal data; that is a transfer, not a performance tweak. See building AI in emerging markets.
Failure modes
- No step cap. The model discovers a new tool each turn. You discover the invoice.
- Observation as system. Injected text becomes policy. Delimit and sanitize.
- Free-text tools. “Action: refund 400” parsed with a regex. Use the vendor’s tool protocol and a schema.
- Tool sprawl. Twenty overlapping tools. The router thrashes. Keep an orthogonal allow-list; tool use is the registry, this pattern is the loop around it.
- Thought leakage. User-visible scratchpads in Portuguese support chats confuse customers and leak retrieval. Keep thoughts off the glass.
- Retry without keys. The same Pix fires twice after a gateway 504.
Curate which loops you even build. Not every market demo deserves a tool-using agent—see the latest Radar. A ReAct loop is justified by a task that needs evidence from systems you already trust, not by a slide that says “autonomous.”
Trade-offs
ReAct is flexible and expensive. You pay tokens, tail latency, and debugging time: failures are trajectories, not single prompts. Plan-and-execute is stiffer and cheaper when the outline is known. A single-shot tool call is cheaper still when you already know the function and the arguments.
Tight caps will abort some legitimate long investigations. That is acceptable if the abort is a clear “budget exceeded” the user can resume. Silent truncation—dropping the last observation—is how you ship a confident wrong answer. Always return a termination reason.
- ReAct with caps — Adaptive tool use and inspectable traces. Variable cost and latency.
- Plan-and-execute — Cheaper steps and clearer phases. A bad plan wastes a whole run.
- Single tool call — Simple SLO. No multi-hop recovery.
- HITL on write tools — The loop cannot complete a dangerous act. You need pending state in the product.
Thought plus act is a control loop, not a prompt style. Production ReAct is budgets, allow-lists, and untrusted observations—then a forced stop.
What to do next
- Set max steps, max spend, and a visible termination reason on every agent entrypoint.
- Sanitize tool observations; never append them as system messages.
- Add idempotency keys before enabling any write tool in the loop, then eval refusals with eval gates for shipping.
Published by Janeiro.ai. Original editorial for operators. How this was made