Agent Memory with Consent and Eviction
Treat agent memory as a data plane: working, summary, and vector layers with tenant ACL, TTL, consent, and a deletion walk—not an unbounded transcript.
Agent memory is not “keep the chat.” It is three stores with different failure modes: working memory (the recent turns), summary memory (a compressed story that can go stale), and vector memory (retrieved snippets that can leak across tenants). If you cannot evict, export, or delete a person from all three, you do not have memory. You have an unbounded personal-data lake with a friendly name.
That matters in Brazil the moment a copilot remembers a CPF, a health-adjacent leave note, or a customer’s complaint across sessions. Memory is the highest-risk data plane in the agent stack. Treat it with the same path discipline as LGPD-aware RAG.
Context
Layered memory exists because long threads exceed context windows. Production systems also have to optimize for consent, TTL, and a data-subject request. A summary is a hypothesis. A vector fact is another index. Both are extra copies under LGPD.
Long threads exceed context windows. Teams then invent memory: sliding windows, periodic summaries, embeddings of “important facts.” The research is real. MemGPT treated the context window as a hierarchical memory hierarchy—working set versus external store—with explicit movement of information. Generative Agents showed that retrieval over a growing memory stream can produce coherent long-horizon behavior. Those papers optimize for believable continuity. Your production system also has to optimize for consent, TTL, and a data-subject request.
Under LGPD, a memory record that identifies a person is a treatment. Purpose limitation and retention apply. “The agent is more helpful if it remembers” is not a legal basis. Conversation logs, summaries, and vector snippets are three copies. The Field Notes checklist in the Portugal–Brazil AI corridor already tells you to set TTLs per store. This pattern is the engineering of those stores.
Two product facts make memory worse than a raw transcript. Summaries drop citations and invent a cleaner story—then that story is retrieved as truth next week. Vector memory retrieves by similarity, so a fact about one customer can surface in another thread if you forgot the tenant filter. The OWASP LLM Top 10 lists sensitive-information disclosure; memory is how disclosure becomes durable.
The pattern
Separate the layers. Working memory is the last k turns, bounded. Summary memory is regenerated on a trigger, versioned, and never the only copy of a fact you still have in source systems. Vector memory is optional, namespaced by tenant, and filtered in-query. Every write carries subject keys, purpose, and TTL. Eviction is a job, not a wiki promise.
Turn → working window
→ summary (if trigger)
→ vector facts (if consented)
All three → prompt pack
All three → evict / TTL → deletion graph
Working memory is a window
Keep the last 8–20 turns, not the year. Prefer the vendor thread or a checkpointer—see LangGraph persistence—over a custom blob you will forget to encrypt. Working memory should be easy to wipe at session end for sensitive desks (HR, collections). Do not embed every turn “just in case.”
Summaries are hypotheses
A summary is a model’s opinion about what mattered. It will drop the CPF you needed and keep a wrong account status. Version summaries. Rebuild them from working memory plus source-of-truth tools when a major state changes (refund issued, contract signed). Never let a stale summary override a tool result in the same turn. If Portuguese UX is the product—see dual-locale Portuguese without dual hallucinations—summarize in the language the next turn will speak, or you will mix registers and confuse the desk.
Vector memory is RAG with a longer half-life
If you store “facts” as embeddings, you have another index. Apply the same ACL, residency, and classify-before-embed rules as retrieval with regional context and the LGPD data-path article. Prefer writing only facts the user or policy allowed: “preferred language: pt-BR,” not “mentioned divorce in passing.” Subject keys on every vector are how a deletion walk finds them.
When to use
Use layered memory when the task is a long thread or a returning user and the alternative is stuffing 200 turns into the prompt. Use it for copilots that must recall a project’s constraints across days. Do not use long-horizon vector memory for a one-shot FAQ. Do not persist memory for a regulated desk until consent, purpose, and TTL are in the schema.
- Ship working memory on any multi-turn agent. Cap it.
- Ship summaries when threads regularly exceed the window and you can rebuild them.
- Ship vector memory only with tenant filters, consent, and a deletion test.
- Skip long-horizon memory for anonymous, single-task tools.
Implementation notes
One assembler, three readers, one evictor. The assembler never concatenates stores blindly: working first, then a fresh summary, then a handful of vector hits that passed ACL. Trace what you packed without logging raw memory text.
Prompt assembly
@dataclass
class MemoryMeta:
tenant_id: str
thread_id: str
purpose: str
subject_keys: list[str]
ttl_days: int
consented: bool
def assemble(thread_id, user, query) -> Prompt:
working = store.last_k(thread_id, k=12)
summary = store.summary(thread_id) # may be None
facts = []
if policy.allows_vector(user):
facts = vector.search(
query,
filters=acl_and_tenant(user),
k=6,
)
return render(working, summary, facts)
def write_fact(user, text, meta: MemoryMeta):
if not meta.consented:
return
label = classify(text)
if label.blocks_index:
return
vector.upsert(embed(redact(text)), meta)
def erase_subject(tenant, subject_key):
store.purge_working(tenant, subject_key)
store.purge_summaries(tenant, subject_key)
vector.delete_by_subject(tenant, subject_key)
Triggers, TTL, and checkpoints
Summarize on token budget, on topic shift, or on explicit “remember this”—not on every turn. TTL working memory in days; summaries in weeks unless the purpose is a long contract; vector facts no longer than the purpose requires. Checkpoints that enable replay are another copy. Encrypt them. They often contain tool payloads.
A ReAct loop will try to stuff observations into memory. Do not. Persist observations in the trace with a short TTL; promote to vector memory only through write_fact and the classifier. Otherwise every injected ticket becomes a long-term instruction.
Failure modes
- Unbounded transcript. You call it memory. Finance calls it storage. Legal calls it retention failure.
- Stale summary as source of truth. The contract changed; the summary did not. Rebuild on state change.
- Vector without tenant filter. Cross-customer “helpfulness.”
- Remember without consent. A support agent that persists health-adjacent text across months.
- Delete the PDF, keep the embedding. Same graph-walk failure as RAG.
- English summaries for a Portuguese desk. The next turn sounds foreign and drops local entities (Pix, CPF, razão social).
The NIST AI RMF “map / measure / manage” sequence is the right ops habit: map the stores, measure stale-summary incidents, manage TTL. It does not replace the LGPD inventory.
Trade-offs
Memory improves continuity and increases leak surface. A tighter window is safer and feels forgetful. A richer vector layer feels personal and becomes a second CRM you did not intend to build. Most teams should ship working memory plus conservative summaries, and wait on vector facts until deletion tests pass.
Summarization cost is real: you pay a model to write a paragraph you may discard next week. Batch it. Do not summarize after every user token on a high-QPS desk. Do not use the largest model to summarize “user prefers morning calls.”
- Working window — Local coherence and easy wipe. Forgets older constraints.
- Versioned summaries — Long threads stay in budget. You pay staleness and extra inference.
- Vector facts — Cross-session recall. You pay ACL, consent, and a deletion graph.
- No long-horizon memory — Smallest compliance surface. Users repeat themselves.
Working memory, summaries, and vector memory are three stores with different leak and staleness profiles. Eviction and consent are part of the design.
What to do next
- Name the three stores in the repo and give each a TTL and an owner.
- Cap working memory; stop embedding every turn.
- Run a deletion walk on staging that proves a subject key leaves all three stores, using the same hop list as LGPD-aware data paths.
Published by Janeiro.ai. Original editorial for operators. How this was made