Skip to content
Janeiro.ai
Pattern

Human-in-the-Loop for High-Risk Agent Actions

Pause agents before money, HR, or production changes: risk gates, durable interrupts, approval UX that shows blast radius, and fail-closed timeouts in ops.

8 min read

An agent that can draft a refund is a writing tool. An agent that can execute the refund is a control system. Human-in-the-loop (HITL) is the interrupt between those two states: durable pause, a named approver, a visible blast radius, and a timeout that fails closed. It is not a system prompt that says “be careful.”

Brazilian buyers already ask for this in procurement. LGPD art. 20 gives titulares a path to review decisions taken solely on automated processing. Corridor products that also sell into the EU inherit human-oversight language from the EU AI Act. The engineering job is to make the interrupt real in the runtime, not only in the annex.

Context

HITL is the product control for excessive agency: persist the proposed action, notify a named role, and resume only with approve, edit, or reject. Timeouts fail closed. Model-scored “confidence” is an input to a table, not a substitute for it.

Most agent demos hide the side effect. The model “updates the ticket,” “issues the credit,” or “restarts the pod” inside the same turn that produced the sentence. That is convenient until the observation was injected, the amount was hallucinated, or the wrong employee record was in context. The OWASP Top 10 for LLM applications treats excessive agency and insecure output handling as first-class failures. HITL is the product control for those classes when the blast radius is a person or a ledger.

Two regional facts change the design. First, Lei nº 13.709/2018 art. 20 is not satisfied by a chat hedge. If retrieval plus a model profiles credit, employment, or access, you need a review path a human can actually take. ANPD’s consultation on AI and automated decisions keeps that article in active regulatory scope. Second, Brazilian enterprise approvals already have owners: jurídico, risco, and the desk that will be paged at 2 a.m. An interrupt that lands in a generic Slack channel with no SLA is not oversight. It is a queue that trains people to click Approve.

The NIST AI Risk Management Framework is useful as a vocabulary—map the action, measure how often humans disagree, manage the residual risk—not as a substitute for LGPD artifacts. Pair this pattern with the Portugal–Brazil AI corridor and the procurement reality in when buyers ask for eval packs before the POC.

The pattern

Score the proposed action before it executes. If the score or the tool class is over the threshold, persist the proposal, notify a named role, and wait. Resume only with an explicit decision: approve, edit, or reject. Timeouts fail closed. The model does not get a second unsupervised try at the same side effect.

Agent proposes action
  ├ low-risk class  → execute idempotent command
  └ high-risk class → persist interrupt → named approver
                      ├ approve → execute
                      ├ edit    → re-propose
                      └ reject or timeout → safe stop + audit

Risk is a table, not a vibe

Put tool classes in Git: read, draft, write-internal, write-external, money, identity, production. Map roles to what they may auto-execute. A support agent can draft; a payments tool always interrupts; a production restart interrupts unless the caller is already on-call with a change ticket. Do not let the model classify its own risk. Models are optimistic about their own blast radius.

The interrupt must survive the process

A await input() inside a request handler dies when the function returns. LangGraph’s interrupt model is the right shape even if you do not use LangGraph: checkpoint the state, surface a payload, resume later with a command. Approvals that take four hours are normal in Brazilian banks and retailers. Design for a different machine, a different day, and a Portuguese UI. Tool contracts for agents that touch money is the schema half of the same gate.

Approval UX is the safety control

Show the diff, the amount, the target account or employee id, the retrieved evidence, and who will be affected. Hide the chain-of-thought. If the screen is a wall of tokens, people rubber-stamp. If the screen is a two-line “the agent is confident,” people rubber-stamp faster. The interrupt exists to make the blast radius boring and visible.

When to use

Use HITL when a wrong action is expensive to reverse: Pix or card credits, payroll, production config, CRM writes that notify a customer, HR status changes, or any decision that would be solely automated under LGPD art. 20. Use it also when the tool is new and the eval set is thin. Confidence scores from the model are not a substitute; they are one input to the table.

Skip HITL on read-only retrieval and on drafts that a human already sits in front of—email compose in an agent’s sidebar does not need a second queue. Do not put HITL on every tool call “to be safe.” That creates an approval factory and trains the organization to ignore it. If everything is high risk, nothing is.

  • Ship now if the agent can call a write tool in production and you cannot name the approver and the timeout.
  • Ship now if security or jurídico asked how a person reviews an automated decision.
  • Defer only for tools that cannot leave the tenant and cannot change external state—and write that exception in the same table.

Implementation notes

Implement three objects: a policy table, an interrupt record, and an idempotent command. The agent proposes a command; the runtime either executes it or stores an interrupt. Humans never execute the raw tool call from the model. They authorize the command id.

Policy table and command envelope

TOOL_CLASS = {
    "search_tickets": "read",
    "draft_reply": "draft",
    "issue_refund": "money",
    "restart_service": "production",
}

def needs_hitl(user, tool, args) -> bool:
    cls = TOOL_CLASS[tool]
    if cls in ("read", "draft"):
        return False
    if cls == "money" or cls == "production":
        return True
    return risk_score(tool, args, user) > user.auto_limit

def propose(user, tool, args) -> Command:
    cmd = Command(id=new_id(), tool=tool, args=validate(tool, args), actor=user)
    if not needs_hitl(user, tool, args):
        return execute(cmd)  # idempotency_key=cmd.id
    interrupt = store.put(cmd, assignee=route_approver(cmd), ttl_hours=4)
    notify(interrupt)
    return interrupt

def resume(interrupt_id, decision, editor) -> Result:
    cmd = store.load(interrupt_id)
    if decision.kind == "timeout" or decision.kind == "reject":
        return store.fail_closed(cmd, decision)
    if decision.kind == "edit":
        cmd.args = validate(cmd.tool, decision.args)
    return execute(cmd, approved_by=editor)

Routing, language, and SLA

Route by tool class and tenant, not by “whoever is online.” Money goes to finance on-call; production goes to the service owner; HR writes go to HRBP. The payload and the notification should be in Portuguese for Brazilian desks—English-only approval screens fail the same way English-only agents fail. Timeouts of four hours are a starting point for office-hours tools; overnight money movement should fail closed, not auto-approve at 04:00.

Log the proposal, the decision, and the command id in the audit trail you will show in a QBR. Redact raw personal fields the same way you would in LGPD-aware RAG paths. An approval log is another personal-data store.

Failure modes

  • In-process wait. The HTTP handler blocks on a human. The interrupt vanishes on deploy or timeout. Persist first.
  • Model-scored risk only. The agent labels a refund “low risk.” The table must win.
  • Rubber-stamp UX. No diff, no amount, no target. Measure reject and edit rates. If they are near zero, the queue is decoration.
  • Timeout auto-approve. Convenient, and how you issue the wrong Pix at dawn. Fail closed unless legal signed a narrow exception.
  • Second unsupervised retry. After reject, a ReAct loop tries a slightly different tool. Consume the decision as a hard constraint.
  • Constitution without a person. Critique-revise can catch tone. It cannot sign a transfer.

Eval this like any other gate. Put known-bad proposals in eval gates for shipping: wrong amount, injected “ignore policy,” cross-tenant target. A HITL path that never fires in fixtures will not fire in production either.

Trade-offs

HITL costs latency and people. A four-hour interrupt is a four-hour customer wait unless you designed a customer-visible pending state. That is still cheaper than an irreversible payment. Teams that optimize only for “autonomous” demos will fight this pattern until the first incident or the first RFP that asks for art. 20 review.

Over-gating destroys the product. If drafts and searches also interrupt, the queue becomes the product and the agent becomes a ticket router. Keep the table small. Expand it when evals or incidents say so, not when a slide says “human in the loop everywhere.”

  • Class-based interrupts — Predictable, reviewable policy. Coarse: some medium-risk writes still auto-run.
  • Durable checkpoint — Approvals survive deploys and nights. You pay a state store, resume semantics, and idempotency.
  • Fail-closed timeout — No dawn Pix. You accept abandoned carts and stale tickets.
  • Diff-first approval UX — Lower rubber-stamp rate. You pay design time, Portuguese copy, and training.

An interrupt is a product surface: durable state, a named approver, a visible blast radius, and a timeout that fails closed.

What to do next

  • Write the tool-class table in the repo and name the approver role per class.
  • Persist interrupts; stop blocking the request thread on a human.
  • Measure reject and edit rates monthly—near-zero means the queue is theater. Then run the 30-day stack check in evaluate an AI stack for LatAm.
agentssafetylgpdevals

Published by . Original editorial for operators. How this was made