Constitutional Critique and Revise
A second pass that checks a written constitution—tone, safety, Portuguese UX—then revises. It catches copy. It does not authorize money transfers.
A constitution is a written policy a second model can fail. The writer drafts; the critic checks the document; the reviser changes the draft or stops. That is useful for tone, safety, and Portuguese UX. It is not a signature. Money, identity, and LGPD art. 20 decisions still go through human-in-the-loop.
Teams that skip the document and say “be nice” have a vibe, not a constitution. Teams that use the critic as the only control have theater. Janeiro treats critique-revise as a cheap second pass on copy, layered on schemas and tool ACL—not as a replacement for them.
Context
Constitutional critique is a bounded draft → critique → revise loop against a versioned principle file. The research goal was training. The production pattern borrowed the loop, with a cap, and without write tools on the critic.
Bai et al.’s Constitutional AI showed that a model can critique and revise against an explicit principle list, reducing reliance on human labels for harmlessness. Anthropic’s research note is the primary public account. The research goal was training. The production pattern borrowed the loop: draft → critique against a file → revise, with a cap.
Customer-facing agents in Brazil fail on register as often as they fail on facts: você versus o senhor, anglicisms in a bank chat, a joke that does not survive jurídico. Dual-locale Portuguese is a product requirement. A constitution can encode it. A constitution cannot invent a legal basis for storing the transcript—see the Portugal–Brazil AI corridor and Lei nº 13.709/2018.
Corridor products that also sell into the EU pick up transparency and oversight language from the EU AI Act. A critique pass can check “did we disclose automation?” It cannot satisfy human oversight for high-risk actions. The NIST AI RMF is the right habit for measuring how often the critic disagrees with the writer—and how often a human later disagrees with both.
The pattern
Keep the constitution in Git: numbered principles, each testable. The writer produces a draft (ideally with structured fields plus a body). The critic returns a list of failed principle ids and spans. If the list is empty, ship the draft. If not, revise once or twice. If still failing, abstain or escalate. The critic does not get tools that write to the world.
Writer draft → critic vs constitution
├ pass → final
├ fail → reviser → critic again
└ still fail → abstain or HITL
Principles are tests
“Be respectful” is not a principle you can eval. “Do not invent a Pix amount,” “do not switch to English mid-reply unless the user did,” and “do not include CPF in the customer-visible body” are. Put examples of fail and pass next to each id. Run them in eval gates for shipping and golden sets in the languages you ship. If a principle never fails in fixtures, it is decoration.
Not a guardrail replacement
Input filters, schema validation, and tool ACL still run first. Critique-revise is for residual policy in language. OWASP still applies: a critic can be prompt-injected by the draft if the draft contains retrieved text. Treat the draft as untrusted.
When to use
Use critique-revise on customer-visible drafts and brand-sensitive copilots. Use it when a second cheap model pass is cheaper than a human queue. Do not use it as the only control on a write tool. Do not run it on every internal search snippet—latency will teach people to turn it off.
- Ship now if support replies go to customers and you have no written tone rules.
- Skip on read-only retrieval with no user-visible prose.
- Escalate when a failed principle is about a decision that affects credit, employment, or access.
Implementation notes
Two models can be the same weight with different prompts; a smaller critic is usually enough. Keep the constitution out of the writer’s system prompt if you want an independent check—otherwise the writer optimizes to the critic’s wording and you lose the second look.
The loop
CONSTITUTION = load("constitution/pt-BR.md") # versioned
class Critique(BaseModel):
failed: list[str] # principle ids
notes: str
def polish(draft, max_rounds=2) -> str:
text = draft
for _ in range(max_rounds):
c = critic(text, CONSTITUTION) # structured
if not c.failed:
return text
text = reviser(text, c)
return escalate_or_abstain(text, c)
What goes in the file
Separate legal facts from brand taste. Legal: no health data in the visible body, no automated credit decision without a review path, no transfer-looking promise the router cannot keep. Brand: Portuguese register, no hype comparatives, no English jargon the desk does not use. Retrieved chunks in LGPD-aware RAG should already be minimized; the critic is a backstop, not the classifier.
Failure modes
- Vibes constitution. Un-evaluable principles. Nobody can say it failed.
- Critic with write tools. The second pass becomes a second agent. No.
- Unbounded revise. Oscillation. Cap at two.
- Writer sees the constitution only. No independent critic. You have one prompt.
- Critique as HITL. A model approved the refund. A person did not.
- English constitution on a Portuguese desk. The critic misses the actual UX failures.
Trade-offs
Two-pass cost is real: you pay extra tokens and latency on every customer-visible turn. Use a small critic and skip the pass on internal drafts. Over-strict constitutions produce bland, refused replies; measure false-positive principle hits the way you measure filter false positives.
Independence costs ops: two prompts to version, two eval sets. It is still cheaper than training a custom harmlessness model for a mid-size Lusophone team. Training-time Constitutional AI is a different project.
- Git-versioned constitution — Reviewable policy. Change control.
- Independent critic — Second look at copy. Tokens and latency.
- HITL for side effects — Critique cannot sign. A real queue.
- Prompt “be careful” — Zero extra calls. No test, no audit.
Put policy in a document the critic can fail. Revise once or twice. Escalate to a human when the constitution is about money, identity, or art. 20 decisions.
What to do next
- Write ten numbered, testable principles in the language of the desk.
- Cap revise rounds at two; escalate instead of looping.
- Keep write tools off the critic. Point side effects at human-in-the-loop.
Published by Janeiro.ai. Original editorial for operators. How this was made