LGPD-Aware Data Paths in RAG Pipelines
Keep personal data off the wrong RAG hop in Brazil: classify at ingest, split indexes, control prompt and log copies, and make deletion a full graph walk.
A RAG pipeline is a copier. Every useful answer is assembled from chunks, embeddings, a packed prompt, a model call, traces, and—if you are honest—eval fixtures you copied from production. Under Lei nº 13.709/2018 (LGPD), each of those copies is a treatment whenever a person in Brazil can still be identified. This pattern is the engineering of that copier: what may be vectorized, where inference may run, how long each hop lives, and how a titular request walks the graph.
It is not legal advice. It is the contract Janeiro expects before a support, HR, or collections copilot indexes tickets in São Paulo or a Portugal–Brazil corridor tenant. Use it with the Portugal–Brazil AI corridor and the stack reset in how to evaluate an AI stack in LatAm in 30 days.
Context
Production RAG in Brazil fails compliance in the hops teams do not draw: the vector replica, the vendor prompt cache, the trace that still holds a CPF, and the “temporary” golden set. Embeddings are not anonymization. A latency fallback to a global model is a transfer.
The useful corpus is almost never a public wiki. It is Zendesk, NFs, CRM notes, Slack exports, and HR PDFs—public policy mixed with dados pessoais and, often, dados sensíveis (LGPD art. 5, II). If you embed first and classify later, the ANN already holds the span jurídico asked you not to store. So does the vendor backup you cannot unsay.
Two facts change the design. First, personal data in a prompt sent to a US or EU endpoint is an international transfer under arts. 33–36, now operationalized by Resolução CD/ANPD nº 19/2024. The DPA must match the router. A “temporary” fallback when southamerica-east1 is slow is still a transfer. Second, ANPD’s work on AI and automated decisions keeps art. 20 in scope: retrieval plus a model that profiles credit, employment, or access needs a human review path, not a hedging sentence.
Lewis et al.’s RAG paper solved grounding. It did not solve a regulated evidence store. Retrieval with regional context and jurisdiction filters before you rank are the quality half. This page is treatment, residency, retention, and erasure. Buyers already ask for the diagram—see when they ask for eval packs before the POC. If you cannot name every hop that can hold a CPF, you do not have a shippable path.
The Pattern
Treat every hop as a typed copy: owner, region, legal basis, TTL, subprocessor, delete key. Classify at ingest, before embed. Keep restricted personal data out of the ANN. Treat prompts, caches, traces, and goldens as extra copies on the same graph. Erase by walking subject keys, not by deleting one PDF.
Source → extract → classify (fail closed)
├ public / minimized → chunk + meta → regional ANN
├ personal restricted → ACL keyword/exact store (no vectors)
└ uncertain sensitive → quarantine (no index, no prompt)
Query: tenant + ACL + residency + jurisdiction THEN rank
Prompt pack → approved-region model only (no silent global fallback)
Prompt / cache / traces / goldens → same deletion graph
Titular request → walk hashed subject keys → EraseReport
Do this
- Label each span
public | internal | personal | sensitivebefore any embed job runs. - Minimize by rewrite: hash or drop CPF, RG, CNS, and contact fields the retrieval task does not need. Keep a source id and a breadcrumb so you can cite without pasting identifiers into the prompt.
- Namespace the index by tenant and residency. Apply jurisdiction filters inside the search request, not after fetch.
- Put hop-level subprocessors in Git: embed vendor, ANN, inference, safety classifier, log sink, eval host. Regional resellers count.
- Stamp a hashed subject key on every write so a DSR is a query, not an archaeology project.
Avoid that
- One “company index.” That is the usual cross-desk leak.
- “The model will be careful.” Uncertain sensitive text goes to quarantine.
- Prompt-only redaction. Instructions do not remove CPFs the retriever already selected.
- Silent global fallback when the Brazil region is slow.
- Golden sets copied from production with no TTL. They become the longest-lived store you have.
Sensitive data (art. 5, II) either stays in the system of record or requires an explicit basis and, usually, a Relatório de Impacto (RIPD) before you copy it again. This pattern is evidence for that report. It is not the report.
Prompts and traces are second copies. Redact before persist. Disable provider training on customer content. Treat prompt cache as a hop with region and TTL. Trace tenant, route, and policy decisions—not raw chunk text. The OWASP Top 10 for LLM applications names sensitive-information disclosure; in this deploy it is also an LGPD incident.
Golden sets in the languages you ship need the same keys. Provenance citations should be source ids plus short quotations from minimized text—not the top twenty raw chunks. Agent memory is a fourth index: same keys, TTLs, consent.
When to Use
Use this pattern when the corpus can contain personal data of people in Brazil, or when any hop—embed, rerank, generate, log, cache—may leave Brazil. That is the default for support, HR, collections, and any CRM-backed copilot.
Skip a full classify-and-split path only when the index is strictly public, non-personal material (published docs, open regulations). Even then, ship the metadata contract and in-query ACL so the path exists when the first ticket corpus arrives.
- Ship now if jurídico asked for hops, subprocessors, log retention, or how a DSR walks the index.
- Ship now if the router can still send CPF-bearing context to a global model, including as a latency fallback.
- Ship now if retrieval plus generation can decide credit, employment, or access. Pair with human-in-the-loop. Art. 20 is not a chat hedge.
- Defer the restricted store only when legal and a labeled ingest sample say the corpus has no personal data. Keep the metadata contract anyway.
- Do not substitute this pattern for a RIPD on high-risk treatment. Use the path as the engineering annex.
Do not perform classify-and-split theater on a ten-document wiki. Do not skip it on a “demo index” that already contains real tickets.
Implementation
Four jobs share one versioned schema: ingest, query, observe, erase. If a field is missing on a write, fail the write.
Metadata every hop must carry
tenant_id, source_id, legal_basis, categories, residency, ttl_days, subject_keys (hashed), subprocessors (that hop, not a PDF). Retrieval filters on tenant, ACL, residency, and category in the search call.
def ingest(doc, tenant, policy) -> IngestResult:
label = classify(extract(doc), policy) # fail closed
if label.blocks_index:
restricted.put(..., meta_from(doc, tenant, label))
return IngestResult(indexed=False, reason=label.reason)
text = redact(extract(doc), keep=policy.indexable_fields)
meta = meta_from(doc, tenant, label)
index.upsert(embed(chunk(text)), meta, namespace=f"{tenant}:{meta.residency}")
return IngestResult(indexed=True)
def retrieve(question, user):
return rerank(question, index.search(
embed(question), filters=acl_and_residency(user), top_k=24,
))[:8]
def erase_subject(tenant, subject_key) -> EraseReport:
return EraseReport(
index=index.delete(index.ids_by_subject(tenant, subject_key)),
restricted=restricted.delete(tenant, subject_key),
traces=traces.tombstone(tenant, subject_key),
cache=cache.purge(tenant, subject_key),
evals=evals.quarantine(tenant, subject_key),
)
Tune classify on labeled Brazilian tickets. An English PII regex that misses CPF, RG, CNS, and razão social is not a control. False negatives are incidents. False positives are desk friction—give a scoped, authorized lookup tool; do not disable the classifier.
Region, transfers, audit
Pick inference region the same way you pick index region. If the approved path is Brazil or an adequacy destination under Resolução 19, do not ship an unnamed global fallback. Cost-aware model routing is where you budget BRL and latency. This pattern is where you refuse to pack personal chunks onto an unapproved hop.
Keep the subprocessor list in Git and review it when a vendor changes region. Buyers will check the ANPD institutional hub and ANPD’s publications library, not a 2021 slide.
At query time: short citations from minimized text; abstain when retrieval is weak (a hallucinated personal fact is worse than “I don’t know”); escalate profiling answers through HITL. Treat retrieved text as untrusted—it must not change allow-lists, residency flags, or the approved model route.
Log what you will show in a QBR: ingest label, hop, region, deny reason, filters, model route, erase report id. Do not log the raw span. The NIST AI Risk Management Framework is a map/measure/manage vocabulary. It does not replace LGPD artifacts.
Fixture these in eval gates for shipping before the first production tenant: classify-after-embed; ACL after ANN; prompt-only redaction; router drift; goldens without TTL; erase-PDF-keep-embedding; English-only classifier; memory that persists health-adjacent text.
Trade-offs
You will lose recall. Questions a naive index “answers” by pasting a sensitive paragraph will abstain or take a slower restricted path. That is the point. Teams that optimize only for retrieval@k will fight this until the first incident or the first blocked RFP.
You will pay BRL and ops: two retrieval implementations, a classifier, a hop inventory, a quarterly deletion drill. The alternative is a global index that fails procurement. Budget the dual path as sprint-zero infrastructure, not a localization ticket—see building AI in emerging markets.
- Classify-then-embed — Index never holds fail-closed spans. You pay quarantine backlog and classifier error.
- Restricted store (no vectors) — Personal facts stay out of ANN. Semantic recall drops; you run two query paths.
- In-query ACL + residency — No cross-tenant neighbors in memory. The vendor must support pre-filters.
- Regional inference only — Transfer story matches the router. Latency and model quality may be worse.
- Erasure graph + TTL — DSRs are testable. Every write needs metadata discipline.
- Hop-level subprocessors — Questionnaires copy from Git. Vendor region changes become change-controlled deploys.
Do not “fix” false-positive redaction by turning classification off. Do not treat this page as a RIPD. Do not claim anonymization for dense embeddings of ticket text—pseudonymized IDs are still treatment; they only make the walk possible.
If you cannot draw the data path, you do not have a RAG system. You have a leak with embeddings.
Related Resources
Primary sources:
- Lei nº 13.709/2018 (LGPD) — official text
- ANPD — official guidance hub
- Resolução CD/ANPD nº 19/2024 — international transfers
- ANPD consultation on AI and automated decisions
- ANPD documents and publications
- Retrieval-Augmented Generation for Knowledge-Intensive NLP (Lewis et al., 2020)
- OWASP Top 10 for Large Language Model Applications
- NIST AI Risk Management Framework
On Janeiro:
Published by Janeiro.ai. Original editorial for operators. How this was made