primer.agents.memory

Memory and state

Run: python -m primer.agents.memory

This lesson builds on the assembled prompt from primer.agents.context, on similarity search from primer.ml.embeddings.similarity, and on the agent loop from primer.agents.agent_loop.

Level 1: The practitioner's guide

In one sentence. A model remembers nothing between calls, so "memory" is the system you build around it: what you keep from a conversation, what you store for months, what you put back into the prompt, who may see it, and where the state of a long task lives so that a crash doesn't lose it.

When you need it. Any product where the second conversation should know about the first, or where one conversation runs long enough to outgrow the window: assistants that remember preferences, agents that resume a multi-step job, support bots with long threads, anything multi-tenant. You don't need long-term memory for a one-shot task, and you don't need a database for a job that finishes in one call. The tells: users repeating themselves every session; a chat that gets slower, dearer and vaguer as it goes (this lesson's unmanaged history climbs without end, while the managed one levels off just under its 300-token budget from about turn 17); an agent that "forgot" a fact the user corrected an hour ago; or a job that has to start over because the process died halfway.

Your options. From nothing to a full system, each layer added on top of the previous one:

Option What it does What it guarantees What it costs Where it lives
The whole history in the prompt Resends every turn on every call Nothing is ever forgotten inside one session Tokens that grow every turn; quality that fades with length Your prompt
Short-term memory with a budget Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave The history never exceeds the budget, however long the session Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call Your code, per conversation
Notes the agent reads on demand The model writes and reads files or notes outside the window through a tool Facts and progress survive a reset of the window Tool calls per read and write, and storage you control A tool and a directory or table
Long-term memory, typed and recalled by similarity Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased An embedding per record, a write policy, a partitioned store, export and deletion paths A store your code owns
Task state in a database Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions A run survives a crash and resumes at the first unfinished step; no step can be skipped A schema and a validator per task type A database

How to choose. Decide separately for the conversation, for facts that outlive it, and for the state of a job, because they fail differently.

  • A chat that runs long: short-term memory with a budget, always. Keep the recent turns verbatim (the order number the user just typed) and summarize the rest.
  • Anything the user would be annoyed to repeat (preferences, corrections, how they like a task done): long-term memory, written selectively. A good assistant's notebook holds "prefers morning meetings", not "said thanks at 3:02pm", and never a password.
  • A job with steps that must happen in order and may outlive a process: task state in a database, with the model proposing and code holding the pen. In this lesson the model's attempt to mark match done while fetch_payments is still open is refused, and after a crash the new process resumes at match from the file, not from a conversation that no longer exists.
  • Many customers on one system: partition the store by tenant and user, chosen by your authentication layer, never by the query. A query that quotes another tenant's secret word for word, with OR 1=1 appended, returns nothing from outside the caller's partition here, because there is no path to search it.
  • Whatever you pick: the model proposes, your code decides what is written. Arrows into the stores never come straight from the model.

What it costs. Short-term memory costs the summarizer (free and crude if extractive; a small model call, with less lost, in production) and the detail it drops. Long-term memory costs an embedding per record at write time, a similarity search per recall, and the engineering around it: a write policy, supersession with provenance, per-user export and deletion. The last two are not optional where privacy law applies; the GDPR's Article 17 gives a person the right to erasure "without undue delay". Task state costs a schema, a validator and a transaction per update, and buys runs that people and monitoring can query. What none of this costs is model quality: memory is retrieval over your own history, and it lives entirely in your code.

What breaks.

  • Context rot. Quality drops and cost climbs as a session grows. Budget the history, summarize the old turns, promote durable facts to long-term memory, or reset with a written hand-off.
  • Remembering everything. Small talk crowds out facts, and a stored secret is read back into every future prompt and every export. Skip chatter, refuse anything that looks like a credential, and strip sensitive data before a note is written.
  • Silent overwrites. The user said April, then corrected to July; a store that overwrites cannot explain why the agent ever said April. Supersede under the same key, recall only the current value, and keep the history with its source.
  • Isolation by filter. A global index filtered after the search is one bug away from a leak; in this lesson's figure the other tenant's secret scores highest against the adversarial query. Partition first; search inside the partition only; test with adversarial queries.
  • Progress kept in the conversation. It is lost on a crash, and it can be summarized, trimmed or misread. Keep it in a database and let code enforce the order of steps.
  • Paths that escape. A memory tool that maps names to files must reject ../ and its encodings, or a request for /memories/../secrets reads outside the store.

In the wild. Park et al. (2023), Generative Agents, gave twenty-five simulated agents a memory stream of natural-language observations, retrieved by relevance, recency and importance and periodically reflected into higher-level memories. Packer et al. (2023), MemGPT, treat the window as fast memory and external storage as slow memory, paging between them the way an operating system does. Claude's memory tool is the note-taking option as a product: the model issues view, create, replace, insert, delete and rename commands against a /memories directory that your application maps onto storage it controls, its system instruction tells the model to assume the window may be reset at any moment, and it pairs with server-side compaction. LangGraph names the same split: thread-scoped short-term memory held by a checkpointer, long-term memory in a store with namespaces, and the semantic, episodic and procedural kinds. SQLite's transactions are the all-or-nothing writes the task store relies on.

Go deeper. Level 2 builds each store in plain Python: a short-term memory whose summary only gets the room the recent turns leave, a long-term memory with three kinds of record recalled by cosine similarity, the write policy, supersession and per-tenant partitions with the leak that a filter would allow drawn as a figure, and a SQLite task store that refuses a skipped step and resumes after a crash. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

What follows builds the four pieces, each small enough to read in one sitting, and shows what each one prevents.

A language model remembers nothing between calls. Every call starts from a blank page plus whatever you put in the prompt. "Memory" is therefore entirely your system: what you keep, where you keep it, what you put back into the prompt, and who is allowed to see it. This lesson builds four pieces: short-term memory, long-term memory with three kinds of record, the safety rules around it (write policy, conflicts, isolation, deletion), and task state kept in a database instead of in the conversation.

flowchart LR subgraph Prompt["What the model sees on this call"] S[System prompt<br/>+ summary of older turns] R[Recalled long-term memories] M[Recent messages, word for word] end STM[(Short-term memory<br/>this conversation)] --> S STM --> M LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R DB[(Task state<br/>SQLite)] -->|current step| S Model[Model] -->|proposes updates| Code[Your code] Code -->|validated writes| LTM Code -->|validated writes| DB

Reading it: the box on the left is the only thing the model ever sees, and it's rebuilt from scratch on every call. The three stores on the right feed it. Arrows into the stores come from your code, never directly from the model. The model proposes and your code decides what gets written.

1. Short-term memory: the conversation, inside a budget

Everyday picture. A flip chart in a long meeting. When the page fills up, you don't find a bigger pad. You tear off the old pages and start a fresh one with one line at the top: "Earlier: agreed on the budget, rejected vendor B."

Tiny worked example. Six question-and-answer exchanges about invoices, 17 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the first sentence of each older message, but all eight of those would take 65 tokens, so the oldest drop out, one at a time, until the rest fit in 40:

summary  : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
           Question 4 is about invoices; Answer 4 lists the invoice              (36 tokens)
messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ...   (verbatim, 70 tokens)

Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 are gone from the prompt entirely. That's the price of a fixed budget, and it's why anything worth keeping forever belongs in long-term memory (section 2), not in the conversation.

Sending the whole history climbs without end, while the managed history levels off just under its 300-token budget

Reading it: the x-axis is the turn number in a long conversation and the y-axis is how many tokens of history are sent on that turn. Unmanaged, the line climbs forever, and so do cost and latency. The model also gets worse at using details buried in the middle. Managed, it climbs until the history first passes the 300-token budget (turn 9), drops as the older turns fold into a summary, then climbs back and stays flat just under 300 from about turn 17 on. It stays flat because the summary only gets the room the recent messages leave: each new turn folds one more exchange in, and the oldest one drops out of the summary to make space.

The code. ShortTermMemory.context() returns (summary, messages). The summary goes in the system prompt rather than as a message, so user and assistant turns still alternate as the API requires.

In code: ShortTermMemory.add appends a message and ShortTermMemory.tokens totals the history; first_sentences is the default summarizer, keeping the first sentence of each folded message. ShortTermMemory.context gives the summary only the budget the recent messages leave and drops the oldest folded messages until it fits. A production system often re-summarizes the summary with an LLM call instead, trading an extra call for losing less.

Why it matters. Context rot, where quality drops as a session grows, is one of the most common agent failures. Summarize, trim, or reset with a written hand-off.

2. Long-term memory: three kinds of record

Everyday picture. Three notebooks: a diary of what happened (episodic: "on 18 Sept Alice rejected Globex"), an encyclopedia of facts (semantic: "Alice's fiscal year starts in April"), and a habit, the way you've learned to do something (procedural: "to reconcile, match each invoice to its payment, then list mismatches").

Tiny worked example. Alice has one memory of each kind. The question "what happened with Globex?" is embedded and compared with each memory (cosine similarity, see primer.ml.embeddings.similarity), and the diary entry comes back first. Asking with kinds=("procedural",) searches only habits.

kind stored as typical recall trigger
episodic dated event "what happened with...", "last time..."
semantic keyed fact (fiscal_year_start) any question the fact answers
procedural how-to steps "how do I...", before starting a known task

Each question recalls the right kind: the Globex question scores the diary entry highest, the how-to question the procedure

Reading it: each group of bars is one question, and each bar is one of Alice's memories, coloured by kind. The tallest bar in each group is what gets recalled. "What happened with Globex?" lights up the diary entry, because only it mentions Globex. "How do I reconcile invoices?" lights up the procedure. This is RAG (retrieval-augmented generation, see primer.agents.rag) over the agent's own history.

In code: LongTermMemory.remember stores a MemoryRecord of one kind with its embedding and reports back in a WriteResult. LongTermMemory.recall returns the current records in a scope most similar to the query, optionally limited to some kinds.

3. The hard parts: what to write, conflicts, isolation, deletion

What to write. Everyday picture: a good assistant's notebook has "prefers morning meetings", not "said thanks at 3:02pm", and never your bank PIN. worth_remembering skips small talk and refuses anything that looks like a secret, since memory is read back into prompts and shown in exports.

Conflicts. Worked example: Alice said her fiscal year starts in April, then corrected it to July. Both are stored under the key fiscal_year_start. The April record is marked superseded by the July one, recall returns only July, and history() still shows both with their source, so "why did the agent think April?" has an answer.

Isolation. Everyday picture: separate locked filing cabinets per company, not one cabinet with a "please only read your own folder" sign. Multi-tenant means one system serves many customers (tenants). Every memory call names a Scope (tenant + user), and storage is partitioned by it, so a lookup starts inside the caller's own cabinet and can't reach another.

flowchart TD Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol] P --> S[Similarity search<br/>inside this partition only] S --> R[Results] A[(acme/alice partition)] -.-x S

Reading it: the query never runs against the whole store. It first opens exactly one partition, chosen from the scope that your authentication layer supplies, never from the query text. The crossed dotted line is the point: there is no path from Carol's search to Acme's records, so no clever query can create one.

Acme's secret scores a perfect match to Carol's quoted query, yet it sits outside her partition and is never searched

Reading it: Carol (tenant Globex) asks a question that quotes Acme's confidential memory word for word. Each bar is that question's similarity to one memory in the whole system. Acme's record scores highest, so a store that searched globally and filtered afterwards is one bug away from leaking it. Hatched bars are outside Carol's partition and are never scored in the real code path.

Deletion. Users may need to see, correct or delete what's remembered, and privacy law such as the GDPR (the EU's General Data Protection Regulation) can require it. export(scope) shows everything including superseded facts, and delete_user(scope) erases one user without touching colleagues.

In code: LongTermMemory.remember applies the write policy and marks an older record with the same key as superseded; LongTermMemory.history lists every value a key has had. LongTermMemory keeps one partition per Scope, and LongTermMemory.export, LongTermMemory.forget and LongTermMemory.delete_user show, remove one record, and erase a user.

4. Task state outside the model

Everyday picture. A checklist on a clipboard. The assistant can suggest ticking a box, but only the supervisor holds the pen. If the assistant goes home sick, the next person picks up the clipboard and starts at the first unticked box.

Tiny worked example.

steps                  : fetch_invoices, fetch_payments, match
model proposes         : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"}   -> written
model proposes         : {"step": "match", "status": "done"}  -> refused: the current step is fetch_payments
process crashes after fetch_payments is done
new process, same file : next_step() -> "match"
sequenceDiagram participant M as Model participant C as Your code participant D as SQLite M->>C: propose {"step": "fetch_invoices", "status": "done"} C->>D: read current step D-->>C: fetch_invoices C->>D: write status=done (transaction) M->>C: propose {"step": "match", "status": "done"} C-->>M: refused: current step is fetch_payments Note over C,D: crash, restart C->>D: next_step? D-->>C: match

Reading it: the model never talks to the database. Every proposal goes through your code, which checks it against the stored truth before writing it inside a transaction (all or nothing). After the crash, the new process learns where to resume from the database, not from a conversation that no longer exists.

In code: TaskStateStore keeps tasks and steps in SQLite. TaskStateStore.create_task writes a task and its steps in one transaction, TaskStateStore.next_step finds the first unfinished step, and TaskStateStore.apply validates a proposal, raising InvalidUpdate for a bad status or a skipped step, before writing it.

Why it matters. Authoritative state in a database makes runs inspectable, resumable and auditable. That's much of the difference between a demo and a production system.

In 20 seconds

  • The model is stateless; memory is what your code puts back into the prompt.
  • Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
  • Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
  • Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
  • Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
  • Keep task state in a database; the model proposes, code validates and writes.

Self-test questions

Q: How would you design long-term memory for a multi-tenant agent platform? A: Every read and write takes a scope (tenant, user) from the authenticated session, and storage is partitioned by it: separate namespaces, collections or row-level security, never a global index filtered after the search. Store typed records (episodic, semantic with keys, procedural) with embeddings, source and time. Apply a write policy (durable, non-secret), supersede conflicting facts instead of overwriting, retrieve the top few by similarity within scope, and support export and deletion per user. Test isolation with adversarial queries.

Q: A long support chat gets worse and more expensive over time. Why, and what helps? A: Every call re-sends the whole history, so cost rises, and models use details in the middle of long contexts less reliably. Keep a budget: recent turns verbatim, older ones summarized, important facts promoted to long-term memory, or reset with a written hand-off.

Q: The user corrects a fact the agent remembered. What should happen? A: Store the new value under the same key, mark the old one as superseded by the new (don't delete it silently), and recall only current records. Keep the source of each so the change is explainable.

Q: Why keep task state in a database when the model can "remember" progress in the conversation? A: The conversation is lost on a crash, and it can be summarized, trimmed or misread. A database is authoritative, survives restarts, can be queried by people and monitoring, and lets code enforce rules (no skipping steps) that a prompt can only request.

The papers behind this lesson

  • Park et al., Generative Agents: Interactive Simulacra of Human Behavior (2023). https://arxiv.org/abs/2304.03442. It introduced a memory stream of observations retrieved by relevance, recency and importance, plus periodic reflection into higher-level memories, a template for long-term agent memory.
  • Packer et al., MemGPT: Towards LLMs as Operating Systems (2023). https://arxiv.org/abs/2310.08560. It treats the context window like RAM and external storage like disk, with the agent paging information in and out, which is the short-term/long-term split made explicit.

Further reading

on GitHub
  1r"""
  2# Memory and state
  3
  4Run: `python -m primer.agents.memory`
  5
  6This lesson builds on the assembled prompt from `primer.agents.context`,
  7on similarity search from `primer.ml.embeddings.similarity`, and on the
  8agent loop from `primer.agents.agent_loop`.
  9
 10## Level 1: The practitioner's guide
 11
 12**In one sentence.** A model remembers nothing between calls, so "memory"
 13is the system you build around it: what you keep from a conversation, what
 14you store for months, what you put back into the prompt, who may see it,
 15and where the state of a long task lives so that a crash doesn't lose it.
 16
 17**When you need it.** Any product where the second conversation should
 18know about the first, or where one conversation runs long enough to
 19outgrow the window: assistants that remember preferences, agents that
 20resume a multi-step job, support bots with long threads, anything
 21multi-tenant. You don't need long-term memory for a one-shot task, and you
 22don't need a database for a job that finishes in one call. The tells: users
 23repeating themselves every session; a chat that gets slower, dearer and
 24vaguer as it goes (this lesson's unmanaged history climbs without end,
 25while the managed one levels off just under its 300-token budget from about
 26turn 17); an agent that "forgot" a fact the user corrected an hour ago; or
 27a job that has to start over because the process died halfway.
 28
 29**Your options.** From nothing to a full system, each layer added on top
 30of the previous one:
 31
 32| Option | What it does | What it guarantees | What it costs | Where it lives |
 33|---|---|---|---|---|
 34| The whole history in the prompt | Resends every turn on every call | Nothing is ever forgotten inside one session | Tokens that grow every turn; quality that fades with length | Your prompt |
 35| Short-term memory with a budget | Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave | The history never exceeds the budget, however long the session | Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call | Your code, per conversation |
 36| Notes the agent reads on demand | The model writes and reads files or notes outside the window through a tool | Facts and progress survive a reset of the window | Tool calls per read and write, and storage you control | A tool and a directory or table |
 37| Long-term memory, typed and recalled by similarity | Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few | The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased | An embedding per record, a write policy, a partitioned store, export and deletion paths | A store your code owns |
 38| Task state in a database | Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions | A run survives a crash and resumes at the first unfinished step; no step can be skipped | A schema and a validator per task type | A database |
 39
 40**How to choose.** Decide separately for the conversation, for facts that
 41outlive it, and for the state of a job, because they fail differently.
 42
 43- A chat that runs long: short-term memory with a budget, always. Keep the
 44  recent turns verbatim (the order number the user just typed) and
 45  summarize the rest.
 46- Anything the user would be annoyed to repeat (preferences, corrections,
 47  how they like a task done): long-term memory, written selectively. A good
 48  assistant's notebook holds "prefers morning meetings", not "said thanks
 49  at 3:02pm", and never a password.
 50- A job with steps that must happen in order and may outlive a process:
 51  task state in a database, with the model proposing and code holding the
 52  pen. In this lesson the model's attempt to mark `match` done while
 53  `fetch_payments` is still open is refused, and after a crash the new
 54  process resumes at `match` from the file, not from a conversation that
 55  no longer exists.
 56- Many customers on one system: partition the store by tenant and user,
 57  chosen by your authentication layer, never by the query. A query that
 58  quotes another tenant's secret word for word, with `OR 1=1` appended,
 59  returns nothing from outside the caller's partition here, because there
 60  is no path to search it.
 61- Whatever you pick: the model proposes, your code decides what is
 62  written. Arrows into the stores never come straight from the model.
 63
 64**What it costs.** Short-term memory costs the summarizer (free and crude
 65if extractive; a small model call, with less lost, in production) and the
 66detail it drops. Long-term memory costs an embedding per record at write
 67time, a similarity search per recall, and the engineering around it: a
 68write policy, supersession with provenance, per-user export and deletion.
 69The last two are not optional where privacy law applies; the GDPR's Article
 7017 gives a person the right to erasure "without undue delay". Task state
 71costs a schema, a validator and a transaction per update, and buys runs
 72that people and monitoring can query. What none of this costs is model
 73quality: memory is retrieval over your own history, and it lives entirely
 74in your code.
 75
 76**What breaks.**
 77
 78- **Context rot.** Quality drops and cost climbs as a session grows.
 79  Budget the history, summarize the old turns, promote durable facts to
 80  long-term memory, or reset with a written hand-off.
 81- **Remembering everything.** Small talk crowds out facts, and a stored
 82  secret is read back into every future prompt and every export. Skip
 83  chatter, refuse anything that looks like a credential, and strip
 84  sensitive data before a note is written.
 85- **Silent overwrites.** The user said April, then corrected to July; a
 86  store that overwrites cannot explain why the agent ever said April.
 87  Supersede under the same key, recall only the current value, and keep
 88  the history with its source.
 89- **Isolation by filter.** A global index filtered after the search is one
 90  bug away from a leak; in this lesson's figure the other tenant's secret
 91  scores highest against the adversarial query. Partition first; search
 92  inside the partition only; test with adversarial queries.
 93- **Progress kept in the conversation.** It is lost on a crash, and it can
 94  be summarized, trimmed or misread. Keep it in a database and let code
 95  enforce the order of steps.
 96- **Paths that escape.** A memory tool that maps names to files must
 97  reject `../` and its encodings, or a request for `/memories/../secrets`
 98  reads outside the store.
 99
100**In the wild.** Park et al. (2023), *Generative Agents*, gave twenty-five
101simulated agents a memory stream of natural-language observations,
102retrieved by relevance, recency and importance and periodically reflected
103into higher-level memories. Packer et al. (2023), *MemGPT*, treat the window
104as fast memory and external storage as slow memory, paging between them
105the way an operating system does. Claude's memory tool is the note-taking
106option as a product: the model issues view, create, replace, insert,
107delete and rename commands against a `/memories` directory that your
108application maps onto storage it controls, its system instruction tells the
109model to assume the window may be reset at any moment, and it pairs with
110server-side compaction. LangGraph names the same split: thread-scoped
111short-term memory held by a checkpointer, long-term memory in a store with
112namespaces, and the semantic, episodic and procedural kinds. SQLite's
113transactions are the all-or-nothing writes the task store relies on.
114
115**Go deeper.** Level 2 builds each store in plain Python: a short-term
116memory whose summary only gets the room the recent turns leave, a long-term
117memory with three kinds of record recalled by cosine similarity, the write
118policy, supersession and per-tenant partitions with the leak that a filter
119would allow drawn as a figure, and a SQLite task store that refuses a
120skipped step and resumes after a crash. If you only needed to choose, you
121are done.
122
123## Level 2: How it works, from scratch
124
125What follows builds the four pieces, each small enough to read in one
126sitting, and shows what each one prevents.
127
128A language model remembers nothing between calls. Every call starts from a
129blank page plus whatever you put in the prompt. "Memory" is therefore
130entirely *your* system: what you keep, where you keep it, what you put back
131into the prompt, and who is allowed to see it. This lesson builds four
132pieces: short-term memory, long-term memory with three kinds of record, the
133safety rules around it (write policy, conflicts, isolation, deletion), and
134task state kept in a database instead of in the conversation.
135
136```mermaid
137flowchart LR
138  subgraph Prompt["What the model sees on this call"]
139    S[System prompt<br/>+ summary of older turns]
140    R[Recalled long-term memories]
141    M[Recent messages, word for word]
142  end
143  STM[(Short-term memory<br/>this conversation)] --> S
144  STM --> M
145  LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R
146  DB[(Task state<br/>SQLite)] -->|current step| S
147  Model[Model] -->|proposes updates| Code[Your code]
148  Code -->|validated writes| LTM
149  Code -->|validated writes| DB
150```
151
152**Reading it:** the box on the left is the only thing the model ever sees,
153and it's rebuilt from scratch on every call. The three stores on the right
154feed it. Arrows *into* the stores come from your code, never directly from
155the model. The model proposes and your code decides what gets written.
156
157## 1. Short-term memory: the conversation, inside a budget
158
159**Everyday picture.** A flip chart in a long meeting. When the page fills up,
160you don't find a bigger pad. You tear off the old pages and start a fresh one
161with one line at the top: "Earlier: agreed on the budget, rejected vendor B."
162
163**Tiny worked example.** Six question-and-answer exchanges about invoices,
16417 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens.
165The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens,
166which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the
167first sentence of each older message, but all eight of those would take 65
168tokens, so the oldest drop out, one at a time, until the rest fit in 40:
169
170```text
171summary  : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
172           Question 4 is about invoices; Answer 4 lists the invoice              (36 tokens)
173messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ...   (verbatim, 70 tokens)
174```
175
176Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2
177are gone from the prompt entirely. That's the price of a fixed budget, and
178it's why anything worth keeping forever belongs in long-term memory
179(section 2), not in the conversation.
180
181![Sending the whole history climbs without end, while the managed history levels off just under its 300-token budget](figures/primer.agents.memory.context_tokens.svg)
182
183**Reading it:** the x-axis is the turn number in a long conversation and the
184y-axis is how many tokens of history are sent on that turn. Unmanaged, the
185line climbs forever, and so do cost and latency. The model also gets worse at
186using details buried in the middle. Managed, it climbs until the history
187first passes the 300-token budget (turn 9), drops as the older turns fold
188into a summary, then climbs back and stays flat just under 300 from about
189turn 17 on. It stays flat because the summary only gets the room the recent
190messages leave: each new turn folds one more exchange in, and the oldest
191one drops out of the summary to make space.
192
193**The code.** `ShortTermMemory.context()` returns `(summary, messages)`. The
194summary goes in the system prompt rather than as a message, so user and
195assistant turns still alternate as the API requires.
196
197**In code:** `ShortTermMemory.add` appends a message and
198`ShortTermMemory.tokens` totals the history; `first_sentences` is the
199default summarizer, keeping the first sentence of each folded message.
200`ShortTermMemory.context` gives the summary only the budget the recent messages leave
201and drops the oldest folded messages until it fits. A production system
202often re-summarizes the summary with an LLM call instead, trading an
203extra call for losing less.
204
205**Why it matters.** Context rot, where quality drops as a session grows, is
206one of the most common agent failures. Summarize, trim, or reset with a
207written hand-off.
208
209## 2. Long-term memory: three kinds of record
210
211**Everyday picture.** Three notebooks: a **diary** of what happened
212(*episodic*: "on 18 Sept Alice rejected Globex"), an **encyclopedia** of
213facts (*semantic*: "Alice's fiscal year starts in April"), and a **habit**,
214the way you've learned to do something (*procedural*: "to reconcile,
215match each invoice to its payment, then list mismatches").
216
217**Tiny worked example.** Alice has one memory of each kind. The question "what
218happened with Globex?" is embedded and compared with each memory (cosine
219similarity, see `primer.ml.embeddings.similarity`), and the diary entry
220comes back first. Asking with `kinds=("procedural",)` searches only habits.
221
222| kind | stored as | typical recall trigger |
223|---|---|---|
224| episodic | dated event | "what happened with...", "last time..." |
225| semantic | keyed fact (`fiscal_year_start`) | any question the fact answers |
226| procedural | how-to steps | "how do I...", before starting a known task |
227
228![Each question recalls the right kind: the Globex question scores the diary entry highest, the how-to question the procedure](figures/primer.agents.memory.recall_scores.svg)
229
230**Reading it:** each group of bars is one question, and each bar is one of
231Alice's memories, coloured by kind. The tallest bar in each group is what
232gets recalled. "What happened with Globex?" lights up the diary entry,
233because only it mentions Globex. "How do I reconcile invoices?" lights up
234the procedure. This is RAG (retrieval-augmented generation, see
235`primer.agents.rag`) over the agent's own history.
236
237**In code:** `LongTermMemory.remember` stores a `MemoryRecord` of one kind
238with its embedding and reports back in a `WriteResult`.
239`LongTermMemory.recall` returns the current records in a scope most similar
240to the query, optionally limited to some kinds.
241
242## 3. The hard parts: what to write, conflicts, isolation, deletion
243
244**What to write.** *Everyday picture:* a good assistant's notebook has
245"prefers morning meetings", not "said thanks at 3:02pm", and never your
246bank PIN. `worth_remembering` skips small talk and refuses anything that
247looks like a secret, since memory is read back into prompts and shown in
248exports.
249
250**Conflicts.** *Worked example:* Alice said her fiscal year starts in April,
251then corrected it to July. Both are stored under the key `fiscal_year_start`.
252The April record is marked *superseded by* the July one, recall returns only
253July, and `history()` still shows both with their source, so "why did the agent
254think April?" has an answer.
255
256**Isolation.** *Everyday picture:* separate locked filing cabinets per
257company, not one cabinet with a "please only read your own folder" sign.
258**Multi-tenant** means one system serves many customers (tenants). Every
259memory call names a `Scope` (tenant + user), and storage is *partitioned*
260by it, so a lookup starts inside the caller's own cabinet and can't reach
261another.
262
263```mermaid
264flowchart TD
265  Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol]
266  P --> S[Similarity search<br/>inside this partition only]
267  S --> R[Results]
268  A[(acme/alice partition)] -.-x S
269```
270
271**Reading it:** the query never runs against the whole store. It first opens
272exactly one partition, chosen from the scope that your authentication layer
273supplies, never from the query text. The crossed dotted line is the point:
274there is no path from Carol's search to Acme's records, so no clever query
275can create one.
276
277![Acme's secret scores a perfect match to Carol's quoted query, yet it sits outside her partition and is never searched](figures/primer.agents.memory.isolation.svg)
278
279**Reading it:** Carol (tenant Globex) asks a question that quotes Acme's
280confidential memory word for word. Each bar is that question's similarity to
281one memory in the *whole* system. Acme's record scores highest, so a
282store that searched globally and filtered afterwards is one bug away from
283leaking it. Hatched bars are outside Carol's partition and are never scored
284in the real code path.
285
286**Deletion.** Users may need to see, correct or delete what's remembered,
287and privacy law such as the GDPR (the EU's General Data Protection
288Regulation) can require it. `export(scope)` shows everything including
289superseded facts, and `delete_user(scope)` erases one user without touching
290colleagues.
291
292**In code:** `LongTermMemory.remember` applies the write policy and marks
293an older record with the same key as superseded; `LongTermMemory.history`
294lists every value a key has had. `LongTermMemory` keeps one partition per
295`Scope`, and `LongTermMemory.export`, `LongTermMemory.forget` and
296`LongTermMemory.delete_user` show, remove one record, and erase a user.
297
298## 4. Task state outside the model
299
300**Everyday picture.** A checklist on a clipboard. The assistant can *suggest*
301ticking a box, but only the supervisor holds the pen. If the assistant goes
302home sick, the next person picks up the clipboard and starts at the first
303unticked box.
304
305**Tiny worked example.**
306
307```text
308steps                  : fetch_invoices, fetch_payments, match
309model proposes         : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"}   -> written
310model proposes         : {"step": "match", "status": "done"}  -> refused: the current step is fetch_payments
311process crashes after fetch_payments is done
312new process, same file : next_step() -> "match"
313```
314
315```mermaid
316sequenceDiagram
317  participant M as Model
318  participant C as Your code
319  participant D as SQLite
320  M->>C: propose {"step": "fetch_invoices", "status": "done"}
321  C->>D: read current step
322  D-->>C: fetch_invoices
323  C->>D: write status=done (transaction)
324  M->>C: propose {"step": "match", "status": "done"}
325  C-->>M: refused: current step is fetch_payments
326  Note over C,D: crash, restart
327  C->>D: next_step?
328  D-->>C: match
329```
330
331**Reading it:** the model never talks to the database. Every proposal goes
332through your code, which checks it against the stored truth before writing
333it inside a transaction (all or nothing). After the crash, the new process
334learns where to resume from the database, not from a conversation that no
335longer exists.
336
337**In code:** `TaskStateStore` keeps tasks and steps in SQLite.
338`TaskStateStore.create_task` writes a task and its steps in one transaction,
339`TaskStateStore.next_step` finds the first unfinished step, and
340`TaskStateStore.apply` validates a proposal, raising `InvalidUpdate` for a
341bad status or a skipped step, before writing it.
342
343**Why it matters.** Authoritative state in a database makes runs
344inspectable, resumable and auditable. That's much of the difference between a
345demo and a production system.
346
347## In 20 seconds
348- The model is stateless; memory is what your code puts back into the prompt.
349- Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
350- Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
351- Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
352- Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
353- Keep task state in a database; the model proposes, code validates and writes.
354
355## Self-test questions
356
357**Q: How would you design long-term memory for a multi-tenant agent platform?**
358A: Every read and write takes a scope (tenant, user) from the authenticated
359session, and storage is partitioned by it: separate namespaces, collections
360or row-level security, never a global index filtered after the search.
361Store typed records (episodic, semantic with keys, procedural) with
362embeddings, source and time. Apply a write policy (durable, non-secret),
363supersede conflicting facts instead of overwriting, retrieve the top few by
364similarity within scope, and support export and deletion per user. Test
365isolation with adversarial queries.
366
367**Q: A long support chat gets worse and more expensive over time. Why, and what helps?**
368A: Every call re-sends the whole history, so cost rises, and models use
369details in the middle of long contexts less reliably. Keep a budget: recent
370turns verbatim, older ones summarized, important facts promoted to
371long-term memory, or reset with a written hand-off.
372
373**Q: The user corrects a fact the agent remembered. What should happen?**
374A: Store the new value under the same key, mark the old one as superseded
375by the new (don't delete it silently), and recall only current records.
376Keep the source of each so the change is explainable.
377
378**Q: Why keep task state in a database when the model can "remember" progress in the conversation?**
379A: The conversation is lost on a crash, and it can be summarized, trimmed or
380misread. A database is authoritative, survives restarts, can be queried by
381people and monitoring, and lets code enforce rules (no skipping steps)
382that a prompt can only request.
383
384## The papers behind this lesson
385
386- **Park et al., *Generative Agents: Interactive Simulacra of Human Behavior* (2023).**
387  https://arxiv.org/abs/2304.03442. It introduced a *memory stream* of
388  observations retrieved by relevance, recency and importance, plus periodic
389  reflection into higher-level memories, a template for long-term agent memory.
390- **Packer et al., *MemGPT: Towards LLMs as Operating Systems* (2023).**
391  https://arxiv.org/abs/2310.08560. It treats the context window like RAM and
392  external storage like disk, with the agent paging information in and out,
393  which is the short-term/long-term split made explicit.
394
395## Further reading
396- Lilian Weng, *LLM Powered Autonomous Agents* (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/
397- LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/
398- SQLite, transactions: https://www.sqlite.org/lang_transaction.html
399- GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/
400"""
401
402from __future__ import annotations
403
404import re
405import sqlite3
406from dataclasses import dataclass, field
407from pathlib import Path
408from typing import Any, Callable
409
410import numpy as np
411
412from primer.agents.llm import estimate_tokens
413from primer.common.embedder import ConceptEmbedder
414
415
416@dataclass
417class ShortTermMemory:
418    """The conversation so far, kept inside a token budget."""
419
420    budget_tokens: int
421    keep_last: int = 4
422    messages: list[dict[str, Any]] = field(default_factory=list)
423    # How older messages are compressed. The default keeps the first sentence of
424    # each (cheap and deterministic); in production this is often an LLM call.
425    summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs))
426
427    def add(self, role: str, text: str) -> None:
428        self.messages.append({"role": role, "content": text})
429
430    def tokens(self) -> int:
431        return sum(estimate_tokens(m["content"]) for m in self.messages)
432
433    def context(self) -> tuple[str, list[dict[str, Any]]]:
434        """(summary for the system prompt, messages to send).
435
436        Under budget: everything, no summary. Over budget: fold all but the last
437        `keep_last` messages into a summary. The summary goes in the system
438        prompt rather than as a message, so user/assistant turns still alternate.
439
440        The summary only gets the room the recent messages leave. Without that
441        cap, a summary that gains a sentence per message would itself outgrow
442        the budget in a long conversation. When it doesn't fit, the oldest
443        folded messages drop out first. If even the recent messages overflow
444        the budget, there is no room left and the summary is empty.
445        """
446        if self.tokens() <= self.budget_tokens:
447            return "", list(self.messages)
448        old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :]
449        room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent)
450        while old:
451            summary = "Earlier in this conversation: " + self.summarize(old)
452            if estimate_tokens(summary) <= room:
453                return summary, list(recent)
454            old = old[1:]  # the oldest detail is the cheapest one to lose
455        return "", list(recent)
456
457
458def first_sentences(messages: list[dict[str, Any]]) -> str:
459    """Extractive summary: the first sentence of each message, joined with '; '."""
460    return "; ".join(m["content"].split(". ")[0].rstrip(".") for m in messages)
461
462
463# ---------------------------------------------------------------------------
464# Long-term memory
465# ---------------------------------------------------------------------------
466
467KINDS = ("episodic", "semantic", "procedural")
468
469_SMALL_TALK = re.compile(r"^\s*(hi|hello|hey|thanks|thank you|ok|okay|great|cool|bye)\b", re.I)
470# Deliberately broad: a false alarm costs one unsaved note; a stored secret leaks
471# into every future prompt and every data export.
472_SECRET = re.compile(r"password|passcode|api[ _-]?key|secret|token|\b\d{3}-\d{2}-\d{4}\b", re.I)
473
474
475@dataclass(frozen=True)
476class Scope:
477    """Who a memory belongs to. Every read and write names one; there is no global view."""
478
479    tenant_id: str
480    user_id: str
481
482    def __post_init__(self) -> None:
483        if not self.tenant_id or not self.user_id:
484            raise ValueError("a memory scope needs both a tenant_id and a user_id")
485
486
487@dataclass
488class MemoryRecord:
489    id: str
490    scope: Scope
491    kind: str
492    content: str
493    key: str | None
494    source: str
495    seq: int  # write order; stands in for a timestamp so examples are deterministic
496    superseded_by: str | None = None
497
498
499@dataclass
500class WriteResult:
501    stored: bool
502    reason: str
503    record: MemoryRecord | None = None
504
505
506def worth_remembering(content: str) -> tuple[bool, str]:
507    """The write policy: store durable, safe facts; skip chatter and refuse secrets."""
508    if _SECRET.search(content):
509        return False, "looks like a secret: credentials never go into memory"
510    if _SMALL_TALK.match(content) and len(content.split()) <= 5:
511        return False, "small talk: nothing durable to remember"
512    return True, "stored"
513
514
515class LongTermMemory:
516    """Memories that outlive a conversation, partitioned by tenant then user."""
517
518    def __init__(self) -> None:
519        # tenant -> user -> records. Partitioning (not filtering) is the isolation:
520        # a lookup starts from the caller's own partition and can't reach another.
521        self._store: dict[str, dict[str, list[MemoryRecord]]] = {}
522        self._vectors: dict[str, np.ndarray] = {}
523        self._seq = 0
524        self.embedder = ConceptEmbedder()
525
526    def _records(self, scope: Scope) -> list[MemoryRecord]:
527        return self._store.setdefault(scope.tenant_id, {}).setdefault(scope.user_id, [])
528
529    def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult:
530        if kind not in KINDS:
531            raise ValueError(f"kind must be one of {KINDS}")
532        ok, reason = worth_remembering(content)
533        if not ok:
534            return WriteResult(False, reason)
535        self._seq += 1
536        record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq)
537        # A keyed fact replaces the older value for the same key, but the old
538        # record stays (marked) so the change is explainable.
539        if key is not None:
540            for old in self._records(scope):
541                if old.key == key and old.superseded_by is None:
542                    old.superseded_by = record.id
543        self._records(scope).append(record)
544        self._vectors[record.id] = self.embedder.encode(content)
545        return WriteResult(True, reason, record)
546
547    def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]:
548        """The k current memories in this scope most similar to the query."""
549        candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds]
550        q = self.embedder.encode(query)
551        return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k]
552
553    def history(self, scope: Scope, key: str) -> list[MemoryRecord]:
554        """Every value a keyed fact has had, oldest first, including superseded ones."""
555        return [r for r in self._records(scope) if r.key == key]
556
557    def forget(self, scope: Scope, record_id: str) -> None:
558        """Delete one memory. Only ids inside the caller's own scope can be found at all."""
559        records = self._records(scope)
560        for i, r in enumerate(records):
561            if r.id == record_id:
562                del records[i]
563                self._vectors.pop(record_id, None)
564                return
565        raise KeyError(f"no memory {record_id} in this scope")
566
567    def export(self, scope: Scope) -> list[dict[str, Any]]:
568        """Everything stored about this user, for them to see and correct (right of access)."""
569        return [
570            {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None}
571            for r in self._records(scope)
572        ]
573
574    def delete_user(self, scope: Scope) -> int:
575        """Erase every memory of one user (right to erasure). Returns how many were deleted."""
576        records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, [])
577        for r in records:
578            self._vectors.pop(r.id, None)
579        return len(records)
580
581
582# ---------------------------------------------------------------------------
583# Task state outside the model (SQLite)
584# ---------------------------------------------------------------------------
585
586STATUSES = ("pending", "running", "done", "failed")
587
588
589class InvalidUpdate(ValueError):
590    """A state change the model proposed that the code refuses to write."""
591
592
593class TaskStateStore:
594    """The authoritative record of a multi-step task, in a database, not in the prompt.
595
596    The model reads it and *proposes* updates as JSON; `apply` checks each
597    proposal against the rules and only then writes it. Because the state is on
598    disk, a crashed run resumes from the first unfinished step.
599    """
600
601    def __init__(self, path: str | Path):
602        self.db = sqlite3.connect(str(path))
603        self.db.executescript(
604            """
605            CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL);
606            CREATE TABLE IF NOT EXISTS steps (
607                task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL,
608                status TEXT NOT NULL DEFAULT 'pending', output TEXT,
609                PRIMARY KEY (task_id, idx)
610            );
611            """
612        )
613
614    def create_task(self, task_id: str, goal: str, steps: list[str]) -> None:
615        with self.db:  # one transaction: the task and all its steps, or nothing
616            self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal))
617            self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)])
618
619    def steps(self, task_id: str) -> list[tuple[str, str, str | None]]:
620        return self.db.execute("SELECT name, status, output FROM steps WHERE task_id = ? ORDER BY idx", (task_id,)).fetchall()
621
622    def next_step(self, task_id: str) -> str | None:
623        """The first step that isn't done, or None when the task is finished."""
624        return next((name for name, status, _ in self.steps(task_id) if status != "done"), None)
625
626    def apply(self, task_id: str, proposal: dict[str, Any]) -> None:
627        """Validate a proposed update from the model, then write it."""
628        step, status = proposal.get("step"), proposal.get("status")
629        if status not in STATUSES:
630            raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}")
631        current = self.next_step(task_id)
632        if step != current:
633            raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}")
634        with self.db:
635            self.db.execute(
636                "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?",
637                (status, proposal.get("output"), task_id, step),
638            )
639
640
641# ---------------------------------------------------------------------------
642# Figures and walkthrough
643# ---------------------------------------------------------------------------
644
645
646def _demo_store() -> tuple[LongTermMemory, Scope, Scope]:
647    mem = LongTermMemory()
648    alice, carol = Scope("acme", "alice"), Scope("globex", "carol")
649    mem.remember(alice, "semantic", "Alice's fiscal year starts in April.", key="fiscal_year_start")
650    mem.remember(alice, "episodic", "On 2026-09-18 Alice rejected the vendor Globex over late invoices.")
651    mem.remember(alice, "procedural", "To reconcile invoices, match each vendor invoice to its payment, then list mismatches.")
652    mem.remember(alice, "semantic", "Acme plans to acquire Initech in Q4.", key="acquisition")
653    mem.remember(carol, "semantic", "Globex prefers invoices in euros.", key="currency")
654    return mem, alice, carol
655
656
657def figures() -> dict:
658    """Plots computed from this lesson's own code (matplotlib imported here)."""
659    import matplotlib
660
661    matplotlib.use("Agg")
662    import matplotlib.pyplot as plt
663
664    figs = {}
665
666    managed = ShortTermMemory(budget_tokens=300, keep_last=6)
667    raw, kept = [], []
668    for i in range(1, 41):
669        managed.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.")
670        managed.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.")
671        summary, msgs = managed.context()
672        raw.append(managed.tokens())
673        kept.append((estimate_tokens(summary) if summary else 0) + sum(estimate_tokens(m["content"]) for m in msgs))
674    fig, ax = plt.subplots(figsize=(7, 4))
675    ax.plot(range(1, 41), raw, label="send the whole history")
676    ax.plot(range(1, 41), kept, label="recent turns + first-sentence summary")
677    ax.axhline(300, ls="--", color="gray", label="budget before summarizing (300)")
678    ax.set(xlabel="turn", ylabel="tokens of history sent", title="Short-term memory under a budget")
679    ax.legend()
680    figs["context_tokens"] = fig
681
682    mem, alice, carol = _demo_store()
683    records = [r for r in mem._records(alice) if r.key != "acquisition"]
684    questions = ["what happened with Globex?", "how do I reconcile invoices?"]
685    colors = {"episodic": "#e0a458", "semantic": "#5b8fd6", "procedural": "#6bb36b"}
686    fig, ax = plt.subplots(figsize=(7, 4))
687    width = 0.25
688    for qi, q in enumerate(questions):
689        qv = mem.embedder.encode(q)
690        for ri, r in enumerate(records):
691            ax.bar(qi + (ri - 1) * width, float(mem._vectors[r.id] @ qv), width, color=colors[r.kind],
692                   label=r.kind if qi == 0 else None)
693    ax.set_xticks(range(len(questions)), questions)
694    ax.set(ylabel="cosine similarity", title="Recall: which of Alice's memories each question finds")
695    ax.legend(title="memory kind")
696    figs["recall_scores"] = fig
697
698    crafted = "Acme plans to acquire Initech in Q4."
699    everything = [r for users in mem._store.values() for recs in users.values() for r in recs]
700    qv = mem.embedder.encode(crafted)
701    scores = [float(mem._vectors[r.id] @ qv) for r in everything]
702    fig, ax = plt.subplots(figsize=(8, 4))
703    for i, (r, sc) in enumerate(zip(everything, scores)):
704        visible = r.scope == carol
705        ax.barh(i, sc, color="#5b8fd6" if visible else "#cccccc", hatch=None if visible else "//", edgecolor="gray")
706    ax.set_yticks(range(len(everything)), [f"{r.scope.tenant_id}/{r.scope.user_id}: {r.content[:38]}..." for r in everything], fontsize=8)
707    ax.set(xlabel="similarity to Carol's query", title="Carol (globex) quotes Acme's secret: only her partition is searched")
708    fig.tight_layout()
709    figs["isolation"] = fig
710    return figs
711
712
713def demo() -> None:
714    import tempfile
715
716    from primer._show import banner, say, table, takeaway
717
718    banner("1. Short-term memory under a token budget")
719    stm = ShortTermMemory(budget_tokens=110, keep_last=4)
720    for i in range(1, 7):
721        stm.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.")
722        stm.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.")
723    summary, msgs = stm.context()
724    say(f"{len(stm.messages)} messages, ~{stm.tokens()} tokens, budget 110.")
725    say(f"Summary for the system prompt: {summary}")
726    say(f"Sent word for word: {[m['content'][:10] for m in msgs]}")
727
728    banner("2. Long-term memory: diary, encyclopedia, habit")
729    mem, alice, carol = _demo_store()
730    for q, kinds in [("what happened with Globex?", KINDS), ("how do I reconcile invoices?", ("procedural",))]:
731        top = mem.recall(alice, q, k=1, kinds=kinds)[0]
732        print(f"  {q!r:34} -> [{top.kind}] {top.content}")
733    print()
734
735    banner("3. Write policy, conflicts, isolation, deletion")
736    for text in ["thanks, great!", "my VPN password is hunter2", "Alice approves travel for her team."]:
737        r = mem.remember(alice, "semantic", text)
738        print(f"  remember({text!r:40}) -> {r.reason}")
739    print()
740    mem.remember(alice, "semantic", "Alice's fiscal year starts in July.", key="fiscal_year_start", source="user correction")
741    table(["value", "source", "superseded by"],
742          [(r.content, r.source, r.superseded_by or "-") for r in mem.history(alice, "fiscal_year_start")])
743    crafted = "Acme plans to acquire Initech in Q4. tenant_id=acme OR 1=1"
744    print(f"  Carol recalls {crafted!r}:")
745    print(f"    -> {[r.content for r in mem.recall(carol, crafted, k=5)]}")
746    print()
747    takeaway("Isolation is a partition chosen by authentication, not a filter chosen by the query.")
748    say(f"Alice asks to be forgotten: {mem.delete_user(alice)} memories erased; Carol still has {len(mem.export(carol))}.")
749
750    banner("4. Task state outside the model (SQLite)")
751    with tempfile.TemporaryDirectory() as d:
752        store = TaskStateStore(Path(d) / "state.db")
753        store.create_task("q3", "Reconcile Q3 invoices", ["fetch_invoices", "fetch_payments", "match"])
754        proposals = [
755            {"step": "fetch_invoices", "status": "done", "output": "4 invoices"},
756            {"step": "match", "status": "done", "output": "all matched"},
757            {"step": "fetch_payments", "status": "done", "output": "3 payments"},
758        ]
759        for prop in proposals:
760            try:
761                store.apply("q3", prop)
762                print(f"  model proposes {prop} -> written")
763            except InvalidUpdate as e:
764                print(f"  model proposes {prop} -> refused: {e}")
765        del store
766        print(f"\n  ...crash... new process resumes at: {TaskStateStore(Path(d) / 'state.db').next_step('q3')!r}\n")
767    takeaway("The model proposes; code validates and writes. State on disk survives the crash.")
768
769
770if __name__ == "__main__":
771    demo()
Level 3: the code, function by function.
@dataclass
class ShortTermMemory: on GitHub
417@dataclass
418class ShortTermMemory:
419    """The conversation so far, kept inside a token budget."""
420
421    budget_tokens: int
422    keep_last: int = 4
423    messages: list[dict[str, Any]] = field(default_factory=list)
424    # How older messages are compressed. The default keeps the first sentence of
425    # each (cheap and deterministic); in production this is often an LLM call.
426    summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs))
427
428    def add(self, role: str, text: str) -> None:
429        self.messages.append({"role": role, "content": text})
430
431    def tokens(self) -> int:
432        return sum(estimate_tokens(m["content"]) for m in self.messages)
433
434    def context(self) -> tuple[str, list[dict[str, Any]]]:
435        """(summary for the system prompt, messages to send).
436
437        Under budget: everything, no summary. Over budget: fold all but the last
438        `keep_last` messages into a summary. The summary goes in the system
439        prompt rather than as a message, so user/assistant turns still alternate.
440
441        The summary only gets the room the recent messages leave. Without that
442        cap, a summary that gains a sentence per message would itself outgrow
443        the budget in a long conversation. When it doesn't fit, the oldest
444        folded messages drop out first. If even the recent messages overflow
445        the budget, there is no room left and the summary is empty.
446        """
447        if self.tokens() <= self.budget_tokens:
448            return "", list(self.messages)
449        old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :]
450        room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent)
451        while old:
452            summary = "Earlier in this conversation: " + self.summarize(old)
453            if estimate_tokens(summary) <= room:
454                return summary, list(recent)
455            old = old[1:]  # the oldest detail is the cheapest one to lose
456        return "", list(recent)

The conversation so far, kept inside a token budget.

ShortTermMemory( budget_tokens: int, keep_last: int = 4, messages: list[dict[str, typing.Any]] = <factory>, summarize: Callable[[list[dict[str, Any]]], str] = <function ShortTermMemory.<lambda>>)
keep_last: int = 4
messages: list[dict[str, typing.Any]]
def summarize(msgs): on GitHub
426    summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs))
def add(self, role: str, text: str) -> None: on GitHub
428    def add(self, role: str, text: str) -> None:
429        self.messages.append({"role": role, "content": text})
def tokens(self) -> int: on GitHub
431    def tokens(self) -> int:
432        return sum(estimate_tokens(m["content"]) for m in self.messages)
def context(self) -> tuple[str, list[dict[str, typing.Any]]]: on GitHub
434    def context(self) -> tuple[str, list[dict[str, Any]]]:
435        """(summary for the system prompt, messages to send).
436
437        Under budget: everything, no summary. Over budget: fold all but the last
438        `keep_last` messages into a summary. The summary goes in the system
439        prompt rather than as a message, so user/assistant turns still alternate.
440
441        The summary only gets the room the recent messages leave. Without that
442        cap, a summary that gains a sentence per message would itself outgrow
443        the budget in a long conversation. When it doesn't fit, the oldest
444        folded messages drop out first. If even the recent messages overflow
445        the budget, there is no room left and the summary is empty.
446        """
447        if self.tokens() <= self.budget_tokens:
448            return "", list(self.messages)
449        old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :]
450        room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent)
451        while old:
452            summary = "Earlier in this conversation: " + self.summarize(old)
453            if estimate_tokens(summary) <= room:
454                return summary, list(recent)
455            old = old[1:]  # the oldest detail is the cheapest one to lose
456        return "", list(recent)

(summary for the system prompt, messages to send).

Under budget: everything, no summary. Over budget: fold all but the last keep_last messages into a summary. The summary goes in the system prompt rather than as a message, so user/assistant turns still alternate.

The summary only gets the room the recent messages leave. Without that cap, a summary that gains a sentence per message would itself outgrow the budget in a long conversation. When it doesn't fit, the oldest folded messages drop out first. If even the recent messages overflow the budget, there is no room left and the summary is empty.

def first_sentences(messages: list[dict[str, typing.Any]]) -> str: on GitHub
459def first_sentences(messages: list[dict[str, Any]]) -> str:
460    """Extractive summary: the first sentence of each message, joined with '; '."""
461    return "; ".join(m["content"].split(". ")[0].rstrip(".") for m in messages)

Extractive summary: the first sentence of each message, joined with '; '.

KINDS = ('episodic', 'semantic', 'procedural')
@dataclass(frozen=True)
class Scope: on GitHub
476@dataclass(frozen=True)
477class Scope:
478    """Who a memory belongs to. Every read and write names one; there is no global view."""
479
480    tenant_id: str
481    user_id: str
482
483    def __post_init__(self) -> None:
484        if not self.tenant_id or not self.user_id:
485            raise ValueError("a memory scope needs both a tenant_id and a user_id")

Who a memory belongs to. Every read and write names one; there is no global view.

Scope(tenant_id: str, user_id: str)
tenant_id: str
user_id: str
@dataclass
class MemoryRecord: on GitHub
488@dataclass
489class MemoryRecord:
490    id: str
491    scope: Scope
492    kind: str
493    content: str
494    key: str | None
495    source: str
496    seq: int  # write order; stands in for a timestamp so examples are deterministic
497    superseded_by: str | None = None
MemoryRecord( id: str, scope: Scope, kind: str, content: str, key: str | None, source: str, seq: int, superseded_by: str | None = None)
id: str
scope: Scope
kind: str
content: str
key: str | None
source: str
seq: int
superseded_by: str | None = None
@dataclass
class WriteResult: on GitHub
500@dataclass
501class WriteResult:
502    stored: bool
503    reason: str
504    record: MemoryRecord | None = None
WriteResult( stored: bool, reason: str, record: MemoryRecord | None = None)
stored: bool
reason: str
record: MemoryRecord | None = None
def worth_remembering(content: str) -> tuple[bool, str]: on GitHub
507def worth_remembering(content: str) -> tuple[bool, str]:
508    """The write policy: store durable, safe facts; skip chatter and refuse secrets."""
509    if _SECRET.search(content):
510        return False, "looks like a secret: credentials never go into memory"
511    if _SMALL_TALK.match(content) and len(content.split()) <= 5:
512        return False, "small talk: nothing durable to remember"
513    return True, "stored"

The write policy: store durable, safe facts; skip chatter and refuse secrets.

class LongTermMemory: on GitHub
516class LongTermMemory:
517    """Memories that outlive a conversation, partitioned by tenant then user."""
518
519    def __init__(self) -> None:
520        # tenant -> user -> records. Partitioning (not filtering) is the isolation:
521        # a lookup starts from the caller's own partition and can't reach another.
522        self._store: dict[str, dict[str, list[MemoryRecord]]] = {}
523        self._vectors: dict[str, np.ndarray] = {}
524        self._seq = 0
525        self.embedder = ConceptEmbedder()
526
527    def _records(self, scope: Scope) -> list[MemoryRecord]:
528        return self._store.setdefault(scope.tenant_id, {}).setdefault(scope.user_id, [])
529
530    def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult:
531        if kind not in KINDS:
532            raise ValueError(f"kind must be one of {KINDS}")
533        ok, reason = worth_remembering(content)
534        if not ok:
535            return WriteResult(False, reason)
536        self._seq += 1
537        record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq)
538        # A keyed fact replaces the older value for the same key, but the old
539        # record stays (marked) so the change is explainable.
540        if key is not None:
541            for old in self._records(scope):
542                if old.key == key and old.superseded_by is None:
543                    old.superseded_by = record.id
544        self._records(scope).append(record)
545        self._vectors[record.id] = self.embedder.encode(content)
546        return WriteResult(True, reason, record)
547
548    def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]:
549        """The k current memories in this scope most similar to the query."""
550        candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds]
551        q = self.embedder.encode(query)
552        return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k]
553
554    def history(self, scope: Scope, key: str) -> list[MemoryRecord]:
555        """Every value a keyed fact has had, oldest first, including superseded ones."""
556        return [r for r in self._records(scope) if r.key == key]
557
558    def forget(self, scope: Scope, record_id: str) -> None:
559        """Delete one memory. Only ids inside the caller's own scope can be found at all."""
560        records = self._records(scope)
561        for i, r in enumerate(records):
562            if r.id == record_id:
563                del records[i]
564                self._vectors.pop(record_id, None)
565                return
566        raise KeyError(f"no memory {record_id} in this scope")
567
568    def export(self, scope: Scope) -> list[dict[str, Any]]:
569        """Everything stored about this user, for them to see and correct (right of access)."""
570        return [
571            {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None}
572            for r in self._records(scope)
573        ]
574
575    def delete_user(self, scope: Scope) -> int:
576        """Erase every memory of one user (right to erasure). Returns how many were deleted."""
577        records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, [])
578        for r in records:
579            self._vectors.pop(r.id, None)
580        return len(records)

Memories that outlive a conversation, partitioned by tenant then user.

embedder
def remember( self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = 'conversation') -> WriteResult: on GitHub
530    def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult:
531        if kind not in KINDS:
532            raise ValueError(f"kind must be one of {KINDS}")
533        ok, reason = worth_remembering(content)
534        if not ok:
535            return WriteResult(False, reason)
536        self._seq += 1
537        record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq)
538        # A keyed fact replaces the older value for the same key, but the old
539        # record stays (marked) so the change is explainable.
540        if key is not None:
541            for old in self._records(scope):
542                if old.key == key and old.superseded_by is None:
543                    old.superseded_by = record.id
544        self._records(scope).append(record)
545        self._vectors[record.id] = self.embedder.encode(content)
546        return WriteResult(True, reason, record)
def recall( self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = ('episodic', 'semantic', 'procedural')) -> list[MemoryRecord]: on GitHub
548    def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]:
549        """The k current memories in this scope most similar to the query."""
550        candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds]
551        q = self.embedder.encode(query)
552        return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k]

The k current memories in this scope most similar to the query.

def history( self, scope: Scope, key: str) -> list[MemoryRecord]: on GitHub
554    def history(self, scope: Scope, key: str) -> list[MemoryRecord]:
555        """Every value a keyed fact has had, oldest first, including superseded ones."""
556        return [r for r in self._records(scope) if r.key == key]

Every value a keyed fact has had, oldest first, including superseded ones.

def forget(self, scope: Scope, record_id: str) -> None: on GitHub
558    def forget(self, scope: Scope, record_id: str) -> None:
559        """Delete one memory. Only ids inside the caller's own scope can be found at all."""
560        records = self._records(scope)
561        for i, r in enumerate(records):
562            if r.id == record_id:
563                del records[i]
564                self._vectors.pop(record_id, None)
565                return
566        raise KeyError(f"no memory {record_id} in this scope")

Delete one memory. Only ids inside the caller's own scope can be found at all.

def export(self, scope: Scope) -> list[dict[str, typing.Any]]: on GitHub
568    def export(self, scope: Scope) -> list[dict[str, Any]]:
569        """Everything stored about this user, for them to see and correct (right of access)."""
570        return [
571            {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None}
572            for r in self._records(scope)
573        ]

Everything stored about this user, for them to see and correct (right of access).

def delete_user(self, scope: Scope) -> int: on GitHub
575    def delete_user(self, scope: Scope) -> int:
576        """Erase every memory of one user (right to erasure). Returns how many were deleted."""
577        records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, [])
578        for r in records:
579            self._vectors.pop(r.id, None)
580        return len(records)

Erase every memory of one user (right to erasure). Returns how many were deleted.

STATUSES = ('pending', 'running', 'done', 'failed')
class InvalidUpdate(builtins.ValueError): on GitHub
590class InvalidUpdate(ValueError):
591    """A state change the model proposed that the code refuses to write."""

A state change the model proposed that the code refuses to write.

class TaskStateStore: on GitHub
594class TaskStateStore:
595    """The authoritative record of a multi-step task, in a database, not in the prompt.
596
597    The model reads it and *proposes* updates as JSON; `apply` checks each
598    proposal against the rules and only then writes it. Because the state is on
599    disk, a crashed run resumes from the first unfinished step.
600    """
601
602    def __init__(self, path: str | Path):
603        self.db = sqlite3.connect(str(path))
604        self.db.executescript(
605            """
606            CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL);
607            CREATE TABLE IF NOT EXISTS steps (
608                task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL,
609                status TEXT NOT NULL DEFAULT 'pending', output TEXT,
610                PRIMARY KEY (task_id, idx)
611            );
612            """
613        )
614
615    def create_task(self, task_id: str, goal: str, steps: list[str]) -> None:
616        with self.db:  # one transaction: the task and all its steps, or nothing
617            self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal))
618            self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)])
619
620    def steps(self, task_id: str) -> list[tuple[str, str, str | None]]:
621        return self.db.execute("SELECT name, status, output FROM steps WHERE task_id = ? ORDER BY idx", (task_id,)).fetchall()
622
623    def next_step(self, task_id: str) -> str | None:
624        """The first step that isn't done, or None when the task is finished."""
625        return next((name for name, status, _ in self.steps(task_id) if status != "done"), None)
626
627    def apply(self, task_id: str, proposal: dict[str, Any]) -> None:
628        """Validate a proposed update from the model, then write it."""
629        step, status = proposal.get("step"), proposal.get("status")
630        if status not in STATUSES:
631            raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}")
632        current = self.next_step(task_id)
633        if step != current:
634            raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}")
635        with self.db:
636            self.db.execute(
637                "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?",
638                (status, proposal.get("output"), task_id, step),
639            )

The authoritative record of a multi-step task, in a database, not in the prompt.

The model reads it and proposes updates as JSON; apply checks each proposal against the rules and only then writes it. Because the state is on disk, a crashed run resumes from the first unfinished step.

TaskStateStore(path: str | pathlib.Path) on GitHub
602    def __init__(self, path: str | Path):
603        self.db = sqlite3.connect(str(path))
604        self.db.executescript(
605            """
606            CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL);
607            CREATE TABLE IF NOT EXISTS steps (
608                task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL,
609                status TEXT NOT NULL DEFAULT 'pending', output TEXT,
610                PRIMARY KEY (task_id, idx)
611            );
612            """
613        )
db
def create_task(self, task_id: str, goal: str, steps: list[str]) -> None: on GitHub
615    def create_task(self, task_id: str, goal: str, steps: list[str]) -> None:
616        with self.db:  # one transaction: the task and all its steps, or nothing
617            self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal))
618            self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)])
def steps(self, task_id: str) -> list[tuple[str, str, str | None]]: on GitHub
620    def steps(self, task_id: str) -> list[tuple[str, str, str | None]]:
621        return self.db.execute("SELECT name, status, output FROM steps WHERE task_id = ? ORDER BY idx", (task_id,)).fetchall()
def next_step(self, task_id: str) -> str | None: on GitHub
623    def next_step(self, task_id: str) -> str | None:
624        """The first step that isn't done, or None when the task is finished."""
625        return next((name for name, status, _ in self.steps(task_id) if status != "done"), None)

The first step that isn't done, or None when the task is finished.

def apply(self, task_id: str, proposal: dict[str, typing.Any]) -> None: on GitHub
627    def apply(self, task_id: str, proposal: dict[str, Any]) -> None:
628        """Validate a proposed update from the model, then write it."""
629        step, status = proposal.get("step"), proposal.get("status")
630        if status not in STATUSES:
631            raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}")
632        current = self.next_step(task_id)
633        if step != current:
634            raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}")
635        with self.db:
636            self.db.execute(
637                "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?",
638                (status, proposal.get("output"), task_id, step),
639            )

Validate a proposed update from the model, then write it.

def figures() -> dict: on GitHub
658def figures() -> dict:
659    """Plots computed from this lesson's own code (matplotlib imported here)."""
660    import matplotlib
661
662    matplotlib.use("Agg")
663    import matplotlib.pyplot as plt
664
665    figs = {}
666
667    managed = ShortTermMemory(budget_tokens=300, keep_last=6)
668    raw, kept = [], []
669    for i in range(1, 41):
670        managed.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.")
671        managed.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.")
672        summary, msgs = managed.context()
673        raw.append(managed.tokens())
674        kept.append((estimate_tokens(summary) if summary else 0) + sum(estimate_tokens(m["content"]) for m in msgs))
675    fig, ax = plt.subplots(figsize=(7, 4))
676    ax.plot(range(1, 41), raw, label="send the whole history")
677    ax.plot(range(1, 41), kept, label="recent turns + first-sentence summary")
678    ax.axhline(300, ls="--", color="gray", label="budget before summarizing (300)")
679    ax.set(xlabel="turn", ylabel="tokens of history sent", title="Short-term memory under a budget")
680    ax.legend()
681    figs["context_tokens"] = fig
682
683    mem, alice, carol = _demo_store()
684    records = [r for r in mem._records(alice) if r.key != "acquisition"]
685    questions = ["what happened with Globex?", "how do I reconcile invoices?"]
686    colors = {"episodic": "#e0a458", "semantic": "#5b8fd6", "procedural": "#6bb36b"}
687    fig, ax = plt.subplots(figsize=(7, 4))
688    width = 0.25
689    for qi, q in enumerate(questions):
690        qv = mem.embedder.encode(q)
691        for ri, r in enumerate(records):
692            ax.bar(qi + (ri - 1) * width, float(mem._vectors[r.id] @ qv), width, color=colors[r.kind],
693                   label=r.kind if qi == 0 else None)
694    ax.set_xticks(range(len(questions)), questions)
695    ax.set(ylabel="cosine similarity", title="Recall: which of Alice's memories each question finds")
696    ax.legend(title="memory kind")
697    figs["recall_scores"] = fig
698
699    crafted = "Acme plans to acquire Initech in Q4."
700    everything = [r for users in mem._store.values() for recs in users.values() for r in recs]
701    qv = mem.embedder.encode(crafted)
702    scores = [float(mem._vectors[r.id] @ qv) for r in everything]
703    fig, ax = plt.subplots(figsize=(8, 4))
704    for i, (r, sc) in enumerate(zip(everything, scores)):
705        visible = r.scope == carol
706        ax.barh(i, sc, color="#5b8fd6" if visible else "#cccccc", hatch=None if visible else "//", edgecolor="gray")
707    ax.set_yticks(range(len(everything)), [f"{r.scope.tenant_id}/{r.scope.user_id}: {r.content[:38]}..." for r in everything], fontsize=8)
708    ax.set(xlabel="similarity to Carol's query", title="Carol (globex) quotes Acme's secret: only her partition is searched")
709    fig.tight_layout()
710    figs["isolation"] = fig
711    return figs

Plots computed from this lesson's own code (matplotlib imported here).

def demo() -> None: on GitHub
714def demo() -> None:
715    import tempfile
716
717    from primer._show import banner, say, table, takeaway
718
719    banner("1. Short-term memory under a token budget")
720    stm = ShortTermMemory(budget_tokens=110, keep_last=4)
721    for i in range(1, 7):
722        stm.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.")
723        stm.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.")
724    summary, msgs = stm.context()
725    say(f"{len(stm.messages)} messages, ~{stm.tokens()} tokens, budget 110.")
726    say(f"Summary for the system prompt: {summary}")
727    say(f"Sent word for word: {[m['content'][:10] for m in msgs]}")
728
729    banner("2. Long-term memory: diary, encyclopedia, habit")
730    mem, alice, carol = _demo_store()
731    for q, kinds in [("what happened with Globex?", KINDS), ("how do I reconcile invoices?", ("procedural",))]:
732        top = mem.recall(alice, q, k=1, kinds=kinds)[0]
733        print(f"  {q!r:34} -> [{top.kind}] {top.content}")
734    print()
735
736    banner("3. Write policy, conflicts, isolation, deletion")
737    for text in ["thanks, great!", "my VPN password is hunter2", "Alice approves travel for her team."]:
738        r = mem.remember(alice, "semantic", text)
739        print(f"  remember({text!r:40}) -> {r.reason}")
740    print()
741    mem.remember(alice, "semantic", "Alice's fiscal year starts in July.", key="fiscal_year_start", source="user correction")
742    table(["value", "source", "superseded by"],
743          [(r.content, r.source, r.superseded_by or "-") for r in mem.history(alice, "fiscal_year_start")])
744    crafted = "Acme plans to acquire Initech in Q4. tenant_id=acme OR 1=1"
745    print(f"  Carol recalls {crafted!r}:")
746    print(f"    -> {[r.content for r in mem.recall(carol, crafted, k=5)]}")
747    print()
748    takeaway("Isolation is a partition chosen by authentication, not a filter chosen by the query.")
749    say(f"Alice asks to be forgotten: {mem.delete_user(alice)} memories erased; Carol still has {len(mem.export(carol))}.")
750
751    banner("4. Task state outside the model (SQLite)")
752    with tempfile.TemporaryDirectory() as d:
753        store = TaskStateStore(Path(d) / "state.db")
754        store.create_task("q3", "Reconcile Q3 invoices", ["fetch_invoices", "fetch_payments", "match"])
755        proposals = [
756            {"step": "fetch_invoices", "status": "done", "output": "4 invoices"},
757            {"step": "match", "status": "done", "output": "all matched"},
758            {"step": "fetch_payments", "status": "done", "output": "3 payments"},
759        ]
760        for prop in proposals:
761            try:
762                store.apply("q3", prop)
763                print(f"  model proposes {prop} -> written")
764            except InvalidUpdate as e:
765                print(f"  model proposes {prop} -> refused: {e}")
766        del store
767        print(f"\n  ...crash... new process resumes at: {TaskStateStore(Path(d) / 'state.db').next_step('q3')!r}\n")
768    takeaway("The model proposes; code validates and writes. State on disk survives the crash.")