primer.agents.memory
Memory and state
Run: python -m primer.agents.memory
This lesson builds on the assembled prompt from primer.agents.context,
on similarity search from primer.ml.embeddings.similarity, and on the
agent loop from primer.agents.agent_loop.
Level 1: The practitioner's guide
In one sentence. A model remembers nothing between calls, so "memory" is the system you build around it: what you keep from a conversation, what you store for months, what you put back into the prompt, who may see it, and where the state of a long task lives so that a crash doesn't lose it.
When you need it. Any product where the second conversation should know about the first, or where one conversation runs long enough to outgrow the window: assistants that remember preferences, agents that resume a multi-step job, support bots with long threads, anything multi-tenant. You don't need long-term memory for a one-shot task, and you don't need a database for a job that finishes in one call. The tells: users repeating themselves every session; a chat that gets slower, dearer and vaguer as it goes (this lesson's unmanaged history climbs without end, while the managed one levels off just under its 300-token budget from about turn 17); an agent that "forgot" a fact the user corrected an hour ago; or a job that has to start over because the process died halfway.
Your options. From nothing to a full system, each layer added on top of the previous one:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| The whole history in the prompt | Resends every turn on every call | Nothing is ever forgotten inside one session | Tokens that grow every turn; quality that fades with length | Your prompt |
| Short-term memory with a budget | Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave | The history never exceeds the budget, however long the session | Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call | Your code, per conversation |
| Notes the agent reads on demand | The model writes and reads files or notes outside the window through a tool | Facts and progress survive a reset of the window | Tool calls per read and write, and storage you control | A tool and a directory or table |
| Long-term memory, typed and recalled by similarity | Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few | The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased | An embedding per record, a write policy, a partitioned store, export and deletion paths | A store your code owns |
| Task state in a database | Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions | A run survives a crash and resumes at the first unfinished step; no step can be skipped | A schema and a validator per task type | A database |
How to choose. Decide separately for the conversation, for facts that outlive it, and for the state of a job, because they fail differently.
- A chat that runs long: short-term memory with a budget, always. Keep the recent turns verbatim (the order number the user just typed) and summarize the rest.
- Anything the user would be annoyed to repeat (preferences, corrections, how they like a task done): long-term memory, written selectively. A good assistant's notebook holds "prefers morning meetings", not "said thanks at 3:02pm", and never a password.
- A job with steps that must happen in order and may outlive a process:
task state in a database, with the model proposing and code holding the
pen. In this lesson the model's attempt to mark
matchdone whilefetch_paymentsis still open is refused, and after a crash the new process resumes atmatchfrom the file, not from a conversation that no longer exists. - Many customers on one system: partition the store by tenant and user,
chosen by your authentication layer, never by the query. A query that
quotes another tenant's secret word for word, with
OR 1=1appended, returns nothing from outside the caller's partition here, because there is no path to search it. - Whatever you pick: the model proposes, your code decides what is written. Arrows into the stores never come straight from the model.
What it costs. Short-term memory costs the summarizer (free and crude if extractive; a small model call, with less lost, in production) and the detail it drops. Long-term memory costs an embedding per record at write time, a similarity search per recall, and the engineering around it: a write policy, supersession with provenance, per-user export and deletion. The last two are not optional where privacy law applies; the GDPR's Article 17 gives a person the right to erasure "without undue delay". Task state costs a schema, a validator and a transaction per update, and buys runs that people and monitoring can query. What none of this costs is model quality: memory is retrieval over your own history, and it lives entirely in your code.
What breaks.
- Context rot. Quality drops and cost climbs as a session grows. Budget the history, summarize the old turns, promote durable facts to long-term memory, or reset with a written hand-off.
- Remembering everything. Small talk crowds out facts, and a stored secret is read back into every future prompt and every export. Skip chatter, refuse anything that looks like a credential, and strip sensitive data before a note is written.
- Silent overwrites. The user said April, then corrected to July; a store that overwrites cannot explain why the agent ever said April. Supersede under the same key, recall only the current value, and keep the history with its source.
- Isolation by filter. A global index filtered after the search is one bug away from a leak; in this lesson's figure the other tenant's secret scores highest against the adversarial query. Partition first; search inside the partition only; test with adversarial queries.
- Progress kept in the conversation. It is lost on a crash, and it can be summarized, trimmed or misread. Keep it in a database and let code enforce the order of steps.
- Paths that escape. A memory tool that maps names to files must
reject
../and its encodings, or a request for/memories/../secretsreads outside the store.
In the wild. Park et al. (2023), Generative Agents, gave twenty-five
simulated agents a memory stream of natural-language observations,
retrieved by relevance, recency and importance and periodically reflected
into higher-level memories. Packer et al. (2023), MemGPT, treat the window
as fast memory and external storage as slow memory, paging between them
the way an operating system does. Claude's memory tool is the note-taking
option as a product: the model issues view, create, replace, insert,
delete and rename commands against a /memories directory that your
application maps onto storage it controls, its system instruction tells the
model to assume the window may be reset at any moment, and it pairs with
server-side compaction. LangGraph names the same split: thread-scoped
short-term memory held by a checkpointer, long-term memory in a store with
namespaces, and the semantic, episodic and procedural kinds. SQLite's
transactions are the all-or-nothing writes the task store relies on.
Go deeper. Level 2 builds each store in plain Python: a short-term memory whose summary only gets the room the recent turns leave, a long-term memory with three kinds of record recalled by cosine similarity, the write policy, supersession and per-tenant partitions with the leak that a filter would allow drawn as a figure, and a SQLite task store that refuses a skipped step and resumes after a crash. If you only needed to choose, you are done.
Level 2: How it works, from scratch
What follows builds the four pieces, each small enough to read in one sitting, and shows what each one prevents.
A language model remembers nothing between calls. Every call starts from a blank page plus whatever you put in the prompt. "Memory" is therefore entirely your system: what you keep, where you keep it, what you put back into the prompt, and who is allowed to see it. This lesson builds four pieces: short-term memory, long-term memory with three kinds of record, the safety rules around it (write policy, conflicts, isolation, deletion), and task state kept in a database instead of in the conversation.
flowchart LR subgraph Prompt["What the model sees on this call"] S[System prompt<br/>+ summary of older turns] R[Recalled long-term memories] M[Recent messages, word for word] end STM[(Short-term memory<br/>this conversation)] --> S STM --> M LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R DB[(Task state<br/>SQLite)] -->|current step| S Model[Model] -->|proposes updates| Code[Your code] Code -->|validated writes| LTM Code -->|validated writes| DB
Reading it: the box on the left is the only thing the model ever sees, and it's rebuilt from scratch on every call. The three stores on the right feed it. Arrows into the stores come from your code, never directly from the model. The model proposes and your code decides what gets written.
1. Short-term memory: the conversation, inside a budget
Everyday picture. A flip chart in a long meeting. When the page fills up, you don't find a bigger pad. You tear off the old pages and start a fresh one with one line at the top: "Earlier: agreed on the budget, rejected vendor B."
Tiny worked example. Six question-and-answer exchanges about invoices, 17 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the first sentence of each older message, but all eight of those would take 65 tokens, so the oldest drop out, one at a time, until the rest fit in 40:
summary : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
Question 4 is about invoices; Answer 4 lists the invoice (36 tokens)
messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ... (verbatim, 70 tokens)
Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 are gone from the prompt entirely. That's the price of a fixed budget, and it's why anything worth keeping forever belongs in long-term memory (section 2), not in the conversation.
Reading it: the x-axis is the turn number in a long conversation and the y-axis is how many tokens of history are sent on that turn. Unmanaged, the line climbs forever, and so do cost and latency. The model also gets worse at using details buried in the middle. Managed, it climbs until the history first passes the 300-token budget (turn 9), drops as the older turns fold into a summary, then climbs back and stays flat just under 300 from about turn 17 on. It stays flat because the summary only gets the room the recent messages leave: each new turn folds one more exchange in, and the oldest one drops out of the summary to make space.
The code. ShortTermMemory.context() returns (summary, messages). The
summary goes in the system prompt rather than as a message, so user and
assistant turns still alternate as the API requires.
In code: ShortTermMemory.add appends a message and
ShortTermMemory.tokens totals the history; first_sentences is the
default summarizer, keeping the first sentence of each folded message.
ShortTermMemory.context gives the summary only the budget the recent messages leave
and drops the oldest folded messages until it fits. A production system
often re-summarizes the summary with an LLM call instead, trading an
extra call for losing less.
Why it matters. Context rot, where quality drops as a session grows, is one of the most common agent failures. Summarize, trim, or reset with a written hand-off.
2. Long-term memory: three kinds of record
Everyday picture. Three notebooks: a diary of what happened (episodic: "on 18 Sept Alice rejected Globex"), an encyclopedia of facts (semantic: "Alice's fiscal year starts in April"), and a habit, the way you've learned to do something (procedural: "to reconcile, match each invoice to its payment, then list mismatches").
Tiny worked example. Alice has one memory of each kind. The question "what
happened with Globex?" is embedded and compared with each memory (cosine
similarity, see primer.ml.embeddings.similarity), and the diary entry
comes back first. Asking with kinds=("procedural",) searches only habits.
| kind | stored as | typical recall trigger |
|---|---|---|
| episodic | dated event | "what happened with...", "last time..." |
| semantic | keyed fact (fiscal_year_start) |
any question the fact answers |
| procedural | how-to steps | "how do I...", before starting a known task |
Reading it: each group of bars is one question, and each bar is one of
Alice's memories, coloured by kind. The tallest bar in each group is what
gets recalled. "What happened with Globex?" lights up the diary entry,
because only it mentions Globex. "How do I reconcile invoices?" lights up
the procedure. This is RAG (retrieval-augmented generation, see
primer.agents.rag) over the agent's own history.
In code: LongTermMemory.remember stores a MemoryRecord of one kind
with its embedding and reports back in a WriteResult.
LongTermMemory.recall returns the current records in a scope most similar
to the query, optionally limited to some kinds.
3. The hard parts: what to write, conflicts, isolation, deletion
What to write. Everyday picture: a good assistant's notebook has
"prefers morning meetings", not "said thanks at 3:02pm", and never your
bank PIN. worth_remembering skips small talk and refuses anything that
looks like a secret, since memory is read back into prompts and shown in
exports.
Conflicts. Worked example: Alice said her fiscal year starts in April,
then corrected it to July. Both are stored under the key fiscal_year_start.
The April record is marked superseded by the July one, recall returns only
July, and history() still shows both with their source, so "why did the agent
think April?" has an answer.
Isolation. Everyday picture: separate locked filing cabinets per
company, not one cabinet with a "please only read your own folder" sign.
Multi-tenant means one system serves many customers (tenants). Every
memory call names a Scope (tenant + user), and storage is partitioned
by it, so a lookup starts inside the caller's own cabinet and can't reach
another.
flowchart TD Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol] P --> S[Similarity search<br/>inside this partition only] S --> R[Results] A[(acme/alice partition)] -.-x S
Reading it: the query never runs against the whole store. It first opens exactly one partition, chosen from the scope that your authentication layer supplies, never from the query text. The crossed dotted line is the point: there is no path from Carol's search to Acme's records, so no clever query can create one.
Reading it: Carol (tenant Globex) asks a question that quotes Acme's confidential memory word for word. Each bar is that question's similarity to one memory in the whole system. Acme's record scores highest, so a store that searched globally and filtered afterwards is one bug away from leaking it. Hatched bars are outside Carol's partition and are never scored in the real code path.
Deletion. Users may need to see, correct or delete what's remembered,
and privacy law such as the GDPR (the EU's General Data Protection
Regulation) can require it. export(scope) shows everything including
superseded facts, and delete_user(scope) erases one user without touching
colleagues.
In code: LongTermMemory.remember applies the write policy and marks
an older record with the same key as superseded; LongTermMemory.history
lists every value a key has had. LongTermMemory keeps one partition per
Scope, and LongTermMemory.export, LongTermMemory.forget and
LongTermMemory.delete_user show, remove one record, and erase a user.
4. Task state outside the model
Everyday picture. A checklist on a clipboard. The assistant can suggest ticking a box, but only the supervisor holds the pen. If the assistant goes home sick, the next person picks up the clipboard and starts at the first unticked box.
Tiny worked example.
steps : fetch_invoices, fetch_payments, match
model proposes : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"} -> written
model proposes : {"step": "match", "status": "done"} -> refused: the current step is fetch_payments
process crashes after fetch_payments is done
new process, same file : next_step() -> "match"
sequenceDiagram participant M as Model participant C as Your code participant D as SQLite M->>C: propose {"step": "fetch_invoices", "status": "done"} C->>D: read current step D-->>C: fetch_invoices C->>D: write status=done (transaction) M->>C: propose {"step": "match", "status": "done"} C-->>M: refused: current step is fetch_payments Note over C,D: crash, restart C->>D: next_step? D-->>C: match
Reading it: the model never talks to the database. Every proposal goes through your code, which checks it against the stored truth before writing it inside a transaction (all or nothing). After the crash, the new process learns where to resume from the database, not from a conversation that no longer exists.
In code: TaskStateStore keeps tasks and steps in SQLite.
TaskStateStore.create_task writes a task and its steps in one transaction,
TaskStateStore.next_step finds the first unfinished step, and
TaskStateStore.apply validates a proposal, raising InvalidUpdate for a
bad status or a skipped step, before writing it.
Why it matters. Authoritative state in a database makes runs inspectable, resumable and auditable. That's much of the difference between a demo and a production system.
In 20 seconds
- The model is stateless; memory is what your code puts back into the prompt.
- Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
- Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
- Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
- Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
- Keep task state in a database; the model proposes, code validates and writes.
Self-test questions
Q: How would you design long-term memory for a multi-tenant agent platform? A: Every read and write takes a scope (tenant, user) from the authenticated session, and storage is partitioned by it: separate namespaces, collections or row-level security, never a global index filtered after the search. Store typed records (episodic, semantic with keys, procedural) with embeddings, source and time. Apply a write policy (durable, non-secret), supersede conflicting facts instead of overwriting, retrieve the top few by similarity within scope, and support export and deletion per user. Test isolation with adversarial queries.
Q: A long support chat gets worse and more expensive over time. Why, and what helps? A: Every call re-sends the whole history, so cost rises, and models use details in the middle of long contexts less reliably. Keep a budget: recent turns verbatim, older ones summarized, important facts promoted to long-term memory, or reset with a written hand-off.
Q: The user corrects a fact the agent remembered. What should happen? A: Store the new value under the same key, mark the old one as superseded by the new (don't delete it silently), and recall only current records. Keep the source of each so the change is explainable.
Q: Why keep task state in a database when the model can "remember" progress in the conversation? A: The conversation is lost on a crash, and it can be summarized, trimmed or misread. A database is authoritative, survives restarts, can be queried by people and monitoring, and lets code enforce rules (no skipping steps) that a prompt can only request.
The papers behind this lesson
- Park et al., Generative Agents: Interactive Simulacra of Human Behavior (2023). https://arxiv.org/abs/2304.03442. It introduced a memory stream of observations retrieved by relevance, recency and importance, plus periodic reflection into higher-level memories, a template for long-term agent memory.
- Packer et al., MemGPT: Towards LLMs as Operating Systems (2023). https://arxiv.org/abs/2310.08560. It treats the context window like RAM and external storage like disk, with the agent paging information in and out, which is the short-term/long-term split made explicit.
Further reading
- Lilian Weng, LLM Powered Autonomous Agents (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/
- LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/
- SQLite, transactions: https://www.sqlite.org/lang_transaction.html
- GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/
1r""" 2# Memory and state 3 4Run: `python -m primer.agents.memory` 5 6This lesson builds on the assembled prompt from `primer.agents.context`, 7on similarity search from `primer.ml.embeddings.similarity`, and on the 8agent loop from `primer.agents.agent_loop`. 9 10## Level 1: The practitioner's guide 11 12**In one sentence.** A model remembers nothing between calls, so "memory" 13is the system you build around it: what you keep from a conversation, what 14you store for months, what you put back into the prompt, who may see it, 15and where the state of a long task lives so that a crash doesn't lose it. 16 17**When you need it.** Any product where the second conversation should 18know about the first, or where one conversation runs long enough to 19outgrow the window: assistants that remember preferences, agents that 20resume a multi-step job, support bots with long threads, anything 21multi-tenant. You don't need long-term memory for a one-shot task, and you 22don't need a database for a job that finishes in one call. The tells: users 23repeating themselves every session; a chat that gets slower, dearer and 24vaguer as it goes (this lesson's unmanaged history climbs without end, 25while the managed one levels off just under its 300-token budget from about 26turn 17); an agent that "forgot" a fact the user corrected an hour ago; or 27a job that has to start over because the process died halfway. 28 29**Your options.** From nothing to a full system, each layer added on top 30of the previous one: 31 32| Option | What it does | What it guarantees | What it costs | Where it lives | 33|---|---|---|---|---| 34| The whole history in the prompt | Resends every turn on every call | Nothing is ever forgotten inside one session | Tokens that grow every turn; quality that fades with length | Your prompt | 35| Short-term memory with a budget | Keeps the last few turns word for word and folds older ones into a summary that only gets the room they leave | The history never exceeds the budget, however long the session | Older detail (in this lesson exchanges 1 and 2 vanish entirely); a summarizer, extractive or a cheap model call | Your code, per conversation | 36| Notes the agent reads on demand | The model writes and reads files or notes outside the window through a tool | Facts and progress survive a reset of the window | Tool calls per read and write, and storage you control | A tool and a directory or table | 37| Long-term memory, typed and recalled by similarity | Stores episodic (events), semantic (keyed facts) and procedural (how-to) records with embeddings, partitioned per tenant and user, and recalls the closest few | The right kind of record comes back for the right question; conflicts are superseded, not overwritten; one user can be exported or erased | An embedding per record, a write policy, a partitioned store, export and deletion paths | A store your code owns | 38| Task state in a database | Keeps the steps of a job in SQLite (or any database); the model proposes updates, code validates and writes them in transactions | A run survives a crash and resumes at the first unfinished step; no step can be skipped | A schema and a validator per task type | A database | 39 40**How to choose.** Decide separately for the conversation, for facts that 41outlive it, and for the state of a job, because they fail differently. 42 43- A chat that runs long: short-term memory with a budget, always. Keep the 44 recent turns verbatim (the order number the user just typed) and 45 summarize the rest. 46- Anything the user would be annoyed to repeat (preferences, corrections, 47 how they like a task done): long-term memory, written selectively. A good 48 assistant's notebook holds "prefers morning meetings", not "said thanks 49 at 3:02pm", and never a password. 50- A job with steps that must happen in order and may outlive a process: 51 task state in a database, with the model proposing and code holding the 52 pen. In this lesson the model's attempt to mark `match` done while 53 `fetch_payments` is still open is refused, and after a crash the new 54 process resumes at `match` from the file, not from a conversation that 55 no longer exists. 56- Many customers on one system: partition the store by tenant and user, 57 chosen by your authentication layer, never by the query. A query that 58 quotes another tenant's secret word for word, with `OR 1=1` appended, 59 returns nothing from outside the caller's partition here, because there 60 is no path to search it. 61- Whatever you pick: the model proposes, your code decides what is 62 written. Arrows into the stores never come straight from the model. 63 64**What it costs.** Short-term memory costs the summarizer (free and crude 65if extractive; a small model call, with less lost, in production) and the 66detail it drops. Long-term memory costs an embedding per record at write 67time, a similarity search per recall, and the engineering around it: a 68write policy, supersession with provenance, per-user export and deletion. 69The last two are not optional where privacy law applies; the GDPR's Article 7017 gives a person the right to erasure "without undue delay". Task state 71costs a schema, a validator and a transaction per update, and buys runs 72that people and monitoring can query. What none of this costs is model 73quality: memory is retrieval over your own history, and it lives entirely 74in your code. 75 76**What breaks.** 77 78- **Context rot.** Quality drops and cost climbs as a session grows. 79 Budget the history, summarize the old turns, promote durable facts to 80 long-term memory, or reset with a written hand-off. 81- **Remembering everything.** Small talk crowds out facts, and a stored 82 secret is read back into every future prompt and every export. Skip 83 chatter, refuse anything that looks like a credential, and strip 84 sensitive data before a note is written. 85- **Silent overwrites.** The user said April, then corrected to July; a 86 store that overwrites cannot explain why the agent ever said April. 87 Supersede under the same key, recall only the current value, and keep 88 the history with its source. 89- **Isolation by filter.** A global index filtered after the search is one 90 bug away from a leak; in this lesson's figure the other tenant's secret 91 scores highest against the adversarial query. Partition first; search 92 inside the partition only; test with adversarial queries. 93- **Progress kept in the conversation.** It is lost on a crash, and it can 94 be summarized, trimmed or misread. Keep it in a database and let code 95 enforce the order of steps. 96- **Paths that escape.** A memory tool that maps names to files must 97 reject `../` and its encodings, or a request for `/memories/../secrets` 98 reads outside the store. 99 100**In the wild.** Park et al. (2023), *Generative Agents*, gave twenty-five 101simulated agents a memory stream of natural-language observations, 102retrieved by relevance, recency and importance and periodically reflected 103into higher-level memories. Packer et al. (2023), *MemGPT*, treat the window 104as fast memory and external storage as slow memory, paging between them 105the way an operating system does. Claude's memory tool is the note-taking 106option as a product: the model issues view, create, replace, insert, 107delete and rename commands against a `/memories` directory that your 108application maps onto storage it controls, its system instruction tells the 109model to assume the window may be reset at any moment, and it pairs with 110server-side compaction. LangGraph names the same split: thread-scoped 111short-term memory held by a checkpointer, long-term memory in a store with 112namespaces, and the semantic, episodic and procedural kinds. SQLite's 113transactions are the all-or-nothing writes the task store relies on. 114 115**Go deeper.** Level 2 builds each store in plain Python: a short-term 116memory whose summary only gets the room the recent turns leave, a long-term 117memory with three kinds of record recalled by cosine similarity, the write 118policy, supersession and per-tenant partitions with the leak that a filter 119would allow drawn as a figure, and a SQLite task store that refuses a 120skipped step and resumes after a crash. If you only needed to choose, you 121are done. 122 123## Level 2: How it works, from scratch 124 125What follows builds the four pieces, each small enough to read in one 126sitting, and shows what each one prevents. 127 128A language model remembers nothing between calls. Every call starts from a 129blank page plus whatever you put in the prompt. "Memory" is therefore 130entirely *your* system: what you keep, where you keep it, what you put back 131into the prompt, and who is allowed to see it. This lesson builds four 132pieces: short-term memory, long-term memory with three kinds of record, the 133safety rules around it (write policy, conflicts, isolation, deletion), and 134task state kept in a database instead of in the conversation. 135 136```mermaid 137flowchart LR 138 subgraph Prompt["What the model sees on this call"] 139 S[System prompt<br/>+ summary of older turns] 140 R[Recalled long-term memories] 141 M[Recent messages, word for word] 142 end 143 STM[(Short-term memory<br/>this conversation)] --> S 144 STM --> M 145 LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R 146 DB[(Task state<br/>SQLite)] -->|current step| S 147 Model[Model] -->|proposes updates| Code[Your code] 148 Code -->|validated writes| LTM 149 Code -->|validated writes| DB 150``` 151 152**Reading it:** the box on the left is the only thing the model ever sees, 153and it's rebuilt from scratch on every call. The three stores on the right 154feed it. Arrows *into* the stores come from your code, never directly from 155the model. The model proposes and your code decides what gets written. 156 157## 1. Short-term memory: the conversation, inside a budget 158 159**Everyday picture.** A flip chart in a long meeting. When the page fills up, 160you don't find a bigger pad. You tear off the old pages and start a fresh one 161with one line at the top: "Earlier: agreed on the budget, rejected vendor B." 162 163**Tiny worked example.** Six question-and-answer exchanges about invoices, 16417 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. 165The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, 166which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the 167first sentence of each older message, but all eight of those would take 65 168tokens, so the oldest drop out, one at a time, until the rest fit in 40: 169 170```text 171summary : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice; 172 Question 4 is about invoices; Answer 4 lists the invoice (36 tokens) 173messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ... (verbatim, 70 tokens) 174``` 175 176Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 177are gone from the prompt entirely. That's the price of a fixed budget, and 178it's why anything worth keeping forever belongs in long-term memory 179(section 2), not in the conversation. 180 181 182 183**Reading it:** the x-axis is the turn number in a long conversation and the 184y-axis is how many tokens of history are sent on that turn. Unmanaged, the 185line climbs forever, and so do cost and latency. The model also gets worse at 186using details buried in the middle. Managed, it climbs until the history 187first passes the 300-token budget (turn 9), drops as the older turns fold 188into a summary, then climbs back and stays flat just under 300 from about 189turn 17 on. It stays flat because the summary only gets the room the recent 190messages leave: each new turn folds one more exchange in, and the oldest 191one drops out of the summary to make space. 192 193**The code.** `ShortTermMemory.context()` returns `(summary, messages)`. The 194summary goes in the system prompt rather than as a message, so user and 195assistant turns still alternate as the API requires. 196 197**In code:** `ShortTermMemory.add` appends a message and 198`ShortTermMemory.tokens` totals the history; `first_sentences` is the 199default summarizer, keeping the first sentence of each folded message. 200`ShortTermMemory.context` gives the summary only the budget the recent messages leave 201and drops the oldest folded messages until it fits. A production system 202often re-summarizes the summary with an LLM call instead, trading an 203extra call for losing less. 204 205**Why it matters.** Context rot, where quality drops as a session grows, is 206one of the most common agent failures. Summarize, trim, or reset with a 207written hand-off. 208 209## 2. Long-term memory: three kinds of record 210 211**Everyday picture.** Three notebooks: a **diary** of what happened 212(*episodic*: "on 18 Sept Alice rejected Globex"), an **encyclopedia** of 213facts (*semantic*: "Alice's fiscal year starts in April"), and a **habit**, 214the way you've learned to do something (*procedural*: "to reconcile, 215match each invoice to its payment, then list mismatches"). 216 217**Tiny worked example.** Alice has one memory of each kind. The question "what 218happened with Globex?" is embedded and compared with each memory (cosine 219similarity, see `primer.ml.embeddings.similarity`), and the diary entry 220comes back first. Asking with `kinds=("procedural",)` searches only habits. 221 222| kind | stored as | typical recall trigger | 223|---|---|---| 224| episodic | dated event | "what happened with...", "last time..." | 225| semantic | keyed fact (`fiscal_year_start`) | any question the fact answers | 226| procedural | how-to steps | "how do I...", before starting a known task | 227 228 229 230**Reading it:** each group of bars is one question, and each bar is one of 231Alice's memories, coloured by kind. The tallest bar in each group is what 232gets recalled. "What happened with Globex?" lights up the diary entry, 233because only it mentions Globex. "How do I reconcile invoices?" lights up 234the procedure. This is RAG (retrieval-augmented generation, see 235`primer.agents.rag`) over the agent's own history. 236 237**In code:** `LongTermMemory.remember` stores a `MemoryRecord` of one kind 238with its embedding and reports back in a `WriteResult`. 239`LongTermMemory.recall` returns the current records in a scope most similar 240to the query, optionally limited to some kinds. 241 242## 3. The hard parts: what to write, conflicts, isolation, deletion 243 244**What to write.** *Everyday picture:* a good assistant's notebook has 245"prefers morning meetings", not "said thanks at 3:02pm", and never your 246bank PIN. `worth_remembering` skips small talk and refuses anything that 247looks like a secret, since memory is read back into prompts and shown in 248exports. 249 250**Conflicts.** *Worked example:* Alice said her fiscal year starts in April, 251then corrected it to July. Both are stored under the key `fiscal_year_start`. 252The April record is marked *superseded by* the July one, recall returns only 253July, and `history()` still shows both with their source, so "why did the agent 254think April?" has an answer. 255 256**Isolation.** *Everyday picture:* separate locked filing cabinets per 257company, not one cabinet with a "please only read your own folder" sign. 258**Multi-tenant** means one system serves many customers (tenants). Every 259memory call names a `Scope` (tenant + user), and storage is *partitioned* 260by it, so a lookup starts inside the caller's own cabinet and can't reach 261another. 262 263```mermaid 264flowchart TD 265 Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol] 266 P --> S[Similarity search<br/>inside this partition only] 267 S --> R[Results] 268 A[(acme/alice partition)] -.-x S 269``` 270 271**Reading it:** the query never runs against the whole store. It first opens 272exactly one partition, chosen from the scope that your authentication layer 273supplies, never from the query text. The crossed dotted line is the point: 274there is no path from Carol's search to Acme's records, so no clever query 275can create one. 276 277 278 279**Reading it:** Carol (tenant Globex) asks a question that quotes Acme's 280confidential memory word for word. Each bar is that question's similarity to 281one memory in the *whole* system. Acme's record scores highest, so a 282store that searched globally and filtered afterwards is one bug away from 283leaking it. Hatched bars are outside Carol's partition and are never scored 284in the real code path. 285 286**Deletion.** Users may need to see, correct or delete what's remembered, 287and privacy law such as the GDPR (the EU's General Data Protection 288Regulation) can require it. `export(scope)` shows everything including 289superseded facts, and `delete_user(scope)` erases one user without touching 290colleagues. 291 292**In code:** `LongTermMemory.remember` applies the write policy and marks 293an older record with the same key as superseded; `LongTermMemory.history` 294lists every value a key has had. `LongTermMemory` keeps one partition per 295`Scope`, and `LongTermMemory.export`, `LongTermMemory.forget` and 296`LongTermMemory.delete_user` show, remove one record, and erase a user. 297 298## 4. Task state outside the model 299 300**Everyday picture.** A checklist on a clipboard. The assistant can *suggest* 301ticking a box, but only the supervisor holds the pen. If the assistant goes 302home sick, the next person picks up the clipboard and starts at the first 303unticked box. 304 305**Tiny worked example.** 306 307```text 308steps : fetch_invoices, fetch_payments, match 309model proposes : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"} -> written 310model proposes : {"step": "match", "status": "done"} -> refused: the current step is fetch_payments 311process crashes after fetch_payments is done 312new process, same file : next_step() -> "match" 313``` 314 315```mermaid 316sequenceDiagram 317 participant M as Model 318 participant C as Your code 319 participant D as SQLite 320 M->>C: propose {"step": "fetch_invoices", "status": "done"} 321 C->>D: read current step 322 D-->>C: fetch_invoices 323 C->>D: write status=done (transaction) 324 M->>C: propose {"step": "match", "status": "done"} 325 C-->>M: refused: current step is fetch_payments 326 Note over C,D: crash, restart 327 C->>D: next_step? 328 D-->>C: match 329``` 330 331**Reading it:** the model never talks to the database. Every proposal goes 332through your code, which checks it against the stored truth before writing 333it inside a transaction (all or nothing). After the crash, the new process 334learns where to resume from the database, not from a conversation that no 335longer exists. 336 337**In code:** `TaskStateStore` keeps tasks and steps in SQLite. 338`TaskStateStore.create_task` writes a task and its steps in one transaction, 339`TaskStateStore.next_step` finds the first unfinished step, and 340`TaskStateStore.apply` validates a proposal, raising `InvalidUpdate` for a 341bad status or a skipped step, before writing it. 342 343**Why it matters.** Authoritative state in a database makes runs 344inspectable, resumable and auditable. That's much of the difference between a 345demo and a production system. 346 347## In 20 seconds 348- The model is stateless; memory is what your code puts back into the prompt. 349- Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget. 350- Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity. 351- Write selectively, never store secrets, supersede rather than overwrite, and keep provenance. 352- Isolate by partitioning on tenant and user, supplied by authentication, never by the query. 353- Keep task state in a database; the model proposes, code validates and writes. 354 355## Self-test questions 356 357**Q: How would you design long-term memory for a multi-tenant agent platform?** 358A: Every read and write takes a scope (tenant, user) from the authenticated 359session, and storage is partitioned by it: separate namespaces, collections 360or row-level security, never a global index filtered after the search. 361Store typed records (episodic, semantic with keys, procedural) with 362embeddings, source and time. Apply a write policy (durable, non-secret), 363supersede conflicting facts instead of overwriting, retrieve the top few by 364similarity within scope, and support export and deletion per user. Test 365isolation with adversarial queries. 366 367**Q: A long support chat gets worse and more expensive over time. Why, and what helps?** 368A: Every call re-sends the whole history, so cost rises, and models use 369details in the middle of long contexts less reliably. Keep a budget: recent 370turns verbatim, older ones summarized, important facts promoted to 371long-term memory, or reset with a written hand-off. 372 373**Q: The user corrects a fact the agent remembered. What should happen?** 374A: Store the new value under the same key, mark the old one as superseded 375by the new (don't delete it silently), and recall only current records. 376Keep the source of each so the change is explainable. 377 378**Q: Why keep task state in a database when the model can "remember" progress in the conversation?** 379A: The conversation is lost on a crash, and it can be summarized, trimmed or 380misread. A database is authoritative, survives restarts, can be queried by 381people and monitoring, and lets code enforce rules (no skipping steps) 382that a prompt can only request. 383 384## The papers behind this lesson 385 386- **Park et al., *Generative Agents: Interactive Simulacra of Human Behavior* (2023).** 387 https://arxiv.org/abs/2304.03442. It introduced a *memory stream* of 388 observations retrieved by relevance, recency and importance, plus periodic 389 reflection into higher-level memories, a template for long-term agent memory. 390- **Packer et al., *MemGPT: Towards LLMs as Operating Systems* (2023).** 391 https://arxiv.org/abs/2310.08560. It treats the context window like RAM and 392 external storage like disk, with the agent paging information in and out, 393 which is the short-term/long-term split made explicit. 394 395## Further reading 396- Lilian Weng, *LLM Powered Autonomous Agents* (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/ 397- LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/ 398- SQLite, transactions: https://www.sqlite.org/lang_transaction.html 399- GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/ 400""" 401 402from __future__ import annotations 403 404import re 405import sqlite3 406from dataclasses import dataclass, field 407from pathlib import Path 408from typing import Any, Callable 409 410import numpy as np 411 412from primer.agents.llm import estimate_tokens 413from primer.common.embedder import ConceptEmbedder 414 415 416@dataclass 417class ShortTermMemory: 418 """The conversation so far, kept inside a token budget.""" 419 420 budget_tokens: int 421 keep_last: int = 4 422 messages: list[dict[str, Any]] = field(default_factory=list) 423 # How older messages are compressed. The default keeps the first sentence of 424 # each (cheap and deterministic); in production this is often an LLM call. 425 summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs)) 426 427 def add(self, role: str, text: str) -> None: 428 self.messages.append({"role": role, "content": text}) 429 430 def tokens(self) -> int: 431 return sum(estimate_tokens(m["content"]) for m in self.messages) 432 433 def context(self) -> tuple[str, list[dict[str, Any]]]: 434 """(summary for the system prompt, messages to send). 435 436 Under budget: everything, no summary. Over budget: fold all but the last 437 `keep_last` messages into a summary. The summary goes in the system 438 prompt rather than as a message, so user/assistant turns still alternate. 439 440 The summary only gets the room the recent messages leave. Without that 441 cap, a summary that gains a sentence per message would itself outgrow 442 the budget in a long conversation. When it doesn't fit, the oldest 443 folded messages drop out first. If even the recent messages overflow 444 the budget, there is no room left and the summary is empty. 445 """ 446 if self.tokens() <= self.budget_tokens: 447 return "", list(self.messages) 448 old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :] 449 room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent) 450 while old: 451 summary = "Earlier in this conversation: " + self.summarize(old) 452 if estimate_tokens(summary) <= room: 453 return summary, list(recent) 454 old = old[1:] # the oldest detail is the cheapest one to lose 455 return "", list(recent) 456 457 458def first_sentences(messages: list[dict[str, Any]]) -> str: 459 """Extractive summary: the first sentence of each message, joined with '; '.""" 460 return "; ".join(m["content"].split(". ")[0].rstrip(".") for m in messages) 461 462 463# --------------------------------------------------------------------------- 464# Long-term memory 465# --------------------------------------------------------------------------- 466 467KINDS = ("episodic", "semantic", "procedural") 468 469_SMALL_TALK = re.compile(r"^\s*(hi|hello|hey|thanks|thank you|ok|okay|great|cool|bye)\b", re.I) 470# Deliberately broad: a false alarm costs one unsaved note; a stored secret leaks 471# into every future prompt and every data export. 472_SECRET = re.compile(r"password|passcode|api[ _-]?key|secret|token|\b\d{3}-\d{2}-\d{4}\b", re.I) 473 474 475@dataclass(frozen=True) 476class Scope: 477 """Who a memory belongs to. Every read and write names one; there is no global view.""" 478 479 tenant_id: str 480 user_id: str 481 482 def __post_init__(self) -> None: 483 if not self.tenant_id or not self.user_id: 484 raise ValueError("a memory scope needs both a tenant_id and a user_id") 485 486 487@dataclass 488class MemoryRecord: 489 id: str 490 scope: Scope 491 kind: str 492 content: str 493 key: str | None 494 source: str 495 seq: int # write order; stands in for a timestamp so examples are deterministic 496 superseded_by: str | None = None 497 498 499@dataclass 500class WriteResult: 501 stored: bool 502 reason: str 503 record: MemoryRecord | None = None 504 505 506def worth_remembering(content: str) -> tuple[bool, str]: 507 """The write policy: store durable, safe facts; skip chatter and refuse secrets.""" 508 if _SECRET.search(content): 509 return False, "looks like a secret: credentials never go into memory" 510 if _SMALL_TALK.match(content) and len(content.split()) <= 5: 511 return False, "small talk: nothing durable to remember" 512 return True, "stored" 513 514 515class LongTermMemory: 516 """Memories that outlive a conversation, partitioned by tenant then user.""" 517 518 def __init__(self) -> None: 519 # tenant -> user -> records. Partitioning (not filtering) is the isolation: 520 # a lookup starts from the caller's own partition and can't reach another. 521 self._store: dict[str, dict[str, list[MemoryRecord]]] = {} 522 self._vectors: dict[str, np.ndarray] = {} 523 self._seq = 0 524 self.embedder = ConceptEmbedder() 525 526 def _records(self, scope: Scope) -> list[MemoryRecord]: 527 return self._store.setdefault(scope.tenant_id, {}).setdefault(scope.user_id, []) 528 529 def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult: 530 if kind not in KINDS: 531 raise ValueError(f"kind must be one of {KINDS}") 532 ok, reason = worth_remembering(content) 533 if not ok: 534 return WriteResult(False, reason) 535 self._seq += 1 536 record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq) 537 # A keyed fact replaces the older value for the same key, but the old 538 # record stays (marked) so the change is explainable. 539 if key is not None: 540 for old in self._records(scope): 541 if old.key == key and old.superseded_by is None: 542 old.superseded_by = record.id 543 self._records(scope).append(record) 544 self._vectors[record.id] = self.embedder.encode(content) 545 return WriteResult(True, reason, record) 546 547 def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]: 548 """The k current memories in this scope most similar to the query.""" 549 candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds] 550 q = self.embedder.encode(query) 551 return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k] 552 553 def history(self, scope: Scope, key: str) -> list[MemoryRecord]: 554 """Every value a keyed fact has had, oldest first, including superseded ones.""" 555 return [r for r in self._records(scope) if r.key == key] 556 557 def forget(self, scope: Scope, record_id: str) -> None: 558 """Delete one memory. Only ids inside the caller's own scope can be found at all.""" 559 records = self._records(scope) 560 for i, r in enumerate(records): 561 if r.id == record_id: 562 del records[i] 563 self._vectors.pop(record_id, None) 564 return 565 raise KeyError(f"no memory {record_id} in this scope") 566 567 def export(self, scope: Scope) -> list[dict[str, Any]]: 568 """Everything stored about this user, for them to see and correct (right of access).""" 569 return [ 570 {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None} 571 for r in self._records(scope) 572 ] 573 574 def delete_user(self, scope: Scope) -> int: 575 """Erase every memory of one user (right to erasure). Returns how many were deleted.""" 576 records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, []) 577 for r in records: 578 self._vectors.pop(r.id, None) 579 return len(records) 580 581 582# --------------------------------------------------------------------------- 583# Task state outside the model (SQLite) 584# --------------------------------------------------------------------------- 585 586STATUSES = ("pending", "running", "done", "failed") 587 588 589class InvalidUpdate(ValueError): 590 """A state change the model proposed that the code refuses to write.""" 591 592 593class TaskStateStore: 594 """The authoritative record of a multi-step task, in a database, not in the prompt. 595 596 The model reads it and *proposes* updates as JSON; `apply` checks each 597 proposal against the rules and only then writes it. Because the state is on 598 disk, a crashed run resumes from the first unfinished step. 599 """ 600 601 def __init__(self, path: str | Path): 602 self.db = sqlite3.connect(str(path)) 603 self.db.executescript( 604 """ 605 CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL); 606 CREATE TABLE IF NOT EXISTS steps ( 607 task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL, 608 status TEXT NOT NULL DEFAULT 'pending', output TEXT, 609 PRIMARY KEY (task_id, idx) 610 ); 611 """ 612 ) 613 614 def create_task(self, task_id: str, goal: str, steps: list[str]) -> None: 615 with self.db: # one transaction: the task and all its steps, or nothing 616 self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal)) 617 self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)]) 618 619 def steps(self, task_id: str) -> list[tuple[str, str, str | None]]: 620 return self.db.execute("SELECT name, status, output FROM steps WHERE task_id = ? ORDER BY idx", (task_id,)).fetchall() 621 622 def next_step(self, task_id: str) -> str | None: 623 """The first step that isn't done, or None when the task is finished.""" 624 return next((name for name, status, _ in self.steps(task_id) if status != "done"), None) 625 626 def apply(self, task_id: str, proposal: dict[str, Any]) -> None: 627 """Validate a proposed update from the model, then write it.""" 628 step, status = proposal.get("step"), proposal.get("status") 629 if status not in STATUSES: 630 raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}") 631 current = self.next_step(task_id) 632 if step != current: 633 raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}") 634 with self.db: 635 self.db.execute( 636 "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?", 637 (status, proposal.get("output"), task_id, step), 638 ) 639 640 641# --------------------------------------------------------------------------- 642# Figures and walkthrough 643# --------------------------------------------------------------------------- 644 645 646def _demo_store() -> tuple[LongTermMemory, Scope, Scope]: 647 mem = LongTermMemory() 648 alice, carol = Scope("acme", "alice"), Scope("globex", "carol") 649 mem.remember(alice, "semantic", "Alice's fiscal year starts in April.", key="fiscal_year_start") 650 mem.remember(alice, "episodic", "On 2026-09-18 Alice rejected the vendor Globex over late invoices.") 651 mem.remember(alice, "procedural", "To reconcile invoices, match each vendor invoice to its payment, then list mismatches.") 652 mem.remember(alice, "semantic", "Acme plans to acquire Initech in Q4.", key="acquisition") 653 mem.remember(carol, "semantic", "Globex prefers invoices in euros.", key="currency") 654 return mem, alice, carol 655 656 657def figures() -> dict: 658 """Plots computed from this lesson's own code (matplotlib imported here).""" 659 import matplotlib 660 661 matplotlib.use("Agg") 662 import matplotlib.pyplot as plt 663 664 figs = {} 665 666 managed = ShortTermMemory(budget_tokens=300, keep_last=6) 667 raw, kept = [], [] 668 for i in range(1, 41): 669 managed.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.") 670 managed.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.") 671 summary, msgs = managed.context() 672 raw.append(managed.tokens()) 673 kept.append((estimate_tokens(summary) if summary else 0) + sum(estimate_tokens(m["content"]) for m in msgs)) 674 fig, ax = plt.subplots(figsize=(7, 4)) 675 ax.plot(range(1, 41), raw, label="send the whole history") 676 ax.plot(range(1, 41), kept, label="recent turns + first-sentence summary") 677 ax.axhline(300, ls="--", color="gray", label="budget before summarizing (300)") 678 ax.set(xlabel="turn", ylabel="tokens of history sent", title="Short-term memory under a budget") 679 ax.legend() 680 figs["context_tokens"] = fig 681 682 mem, alice, carol = _demo_store() 683 records = [r for r in mem._records(alice) if r.key != "acquisition"] 684 questions = ["what happened with Globex?", "how do I reconcile invoices?"] 685 colors = {"episodic": "#e0a458", "semantic": "#5b8fd6", "procedural": "#6bb36b"} 686 fig, ax = plt.subplots(figsize=(7, 4)) 687 width = 0.25 688 for qi, q in enumerate(questions): 689 qv = mem.embedder.encode(q) 690 for ri, r in enumerate(records): 691 ax.bar(qi + (ri - 1) * width, float(mem._vectors[r.id] @ qv), width, color=colors[r.kind], 692 label=r.kind if qi == 0 else None) 693 ax.set_xticks(range(len(questions)), questions) 694 ax.set(ylabel="cosine similarity", title="Recall: which of Alice's memories each question finds") 695 ax.legend(title="memory kind") 696 figs["recall_scores"] = fig 697 698 crafted = "Acme plans to acquire Initech in Q4." 699 everything = [r for users in mem._store.values() for recs in users.values() for r in recs] 700 qv = mem.embedder.encode(crafted) 701 scores = [float(mem._vectors[r.id] @ qv) for r in everything] 702 fig, ax = plt.subplots(figsize=(8, 4)) 703 for i, (r, sc) in enumerate(zip(everything, scores)): 704 visible = r.scope == carol 705 ax.barh(i, sc, color="#5b8fd6" if visible else "#cccccc", hatch=None if visible else "//", edgecolor="gray") 706 ax.set_yticks(range(len(everything)), [f"{r.scope.tenant_id}/{r.scope.user_id}: {r.content[:38]}..." for r in everything], fontsize=8) 707 ax.set(xlabel="similarity to Carol's query", title="Carol (globex) quotes Acme's secret: only her partition is searched") 708 fig.tight_layout() 709 figs["isolation"] = fig 710 return figs 711 712 713def demo() -> None: 714 import tempfile 715 716 from primer._show import banner, say, table, takeaway 717 718 banner("1. Short-term memory under a token budget") 719 stm = ShortTermMemory(budget_tokens=110, keep_last=4) 720 for i in range(1, 7): 721 stm.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.") 722 stm.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.") 723 summary, msgs = stm.context() 724 say(f"{len(stm.messages)} messages, ~{stm.tokens()} tokens, budget 110.") 725 say(f"Summary for the system prompt: {summary}") 726 say(f"Sent word for word: {[m['content'][:10] for m in msgs]}") 727 728 banner("2. Long-term memory: diary, encyclopedia, habit") 729 mem, alice, carol = _demo_store() 730 for q, kinds in [("what happened with Globex?", KINDS), ("how do I reconcile invoices?", ("procedural",))]: 731 top = mem.recall(alice, q, k=1, kinds=kinds)[0] 732 print(f" {q!r:34} -> [{top.kind}] {top.content}") 733 print() 734 735 banner("3. Write policy, conflicts, isolation, deletion") 736 for text in ["thanks, great!", "my VPN password is hunter2", "Alice approves travel for her team."]: 737 r = mem.remember(alice, "semantic", text) 738 print(f" remember({text!r:40}) -> {r.reason}") 739 print() 740 mem.remember(alice, "semantic", "Alice's fiscal year starts in July.", key="fiscal_year_start", source="user correction") 741 table(["value", "source", "superseded by"], 742 [(r.content, r.source, r.superseded_by or "-") for r in mem.history(alice, "fiscal_year_start")]) 743 crafted = "Acme plans to acquire Initech in Q4. tenant_id=acme OR 1=1" 744 print(f" Carol recalls {crafted!r}:") 745 print(f" -> {[r.content for r in mem.recall(carol, crafted, k=5)]}") 746 print() 747 takeaway("Isolation is a partition chosen by authentication, not a filter chosen by the query.") 748 say(f"Alice asks to be forgotten: {mem.delete_user(alice)} memories erased; Carol still has {len(mem.export(carol))}.") 749 750 banner("4. Task state outside the model (SQLite)") 751 with tempfile.TemporaryDirectory() as d: 752 store = TaskStateStore(Path(d) / "state.db") 753 store.create_task("q3", "Reconcile Q3 invoices", ["fetch_invoices", "fetch_payments", "match"]) 754 proposals = [ 755 {"step": "fetch_invoices", "status": "done", "output": "4 invoices"}, 756 {"step": "match", "status": "done", "output": "all matched"}, 757 {"step": "fetch_payments", "status": "done", "output": "3 payments"}, 758 ] 759 for prop in proposals: 760 try: 761 store.apply("q3", prop) 762 print(f" model proposes {prop} -> written") 763 except InvalidUpdate as e: 764 print(f" model proposes {prop} -> refused: {e}") 765 del store 766 print(f"\n ...crash... new process resumes at: {TaskStateStore(Path(d) / 'state.db').next_step('q3')!r}\n") 767 takeaway("The model proposes; code validates and writes. State on disk survives the crash.") 768 769 770if __name__ == "__main__": 771 demo()
417@dataclass 418class ShortTermMemory: 419 """The conversation so far, kept inside a token budget.""" 420 421 budget_tokens: int 422 keep_last: int = 4 423 messages: list[dict[str, Any]] = field(default_factory=list) 424 # How older messages are compressed. The default keeps the first sentence of 425 # each (cheap and deterministic); in production this is often an LLM call. 426 summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs)) 427 428 def add(self, role: str, text: str) -> None: 429 self.messages.append({"role": role, "content": text}) 430 431 def tokens(self) -> int: 432 return sum(estimate_tokens(m["content"]) for m in self.messages) 433 434 def context(self) -> tuple[str, list[dict[str, Any]]]: 435 """(summary for the system prompt, messages to send). 436 437 Under budget: everything, no summary. Over budget: fold all but the last 438 `keep_last` messages into a summary. The summary goes in the system 439 prompt rather than as a message, so user/assistant turns still alternate. 440 441 The summary only gets the room the recent messages leave. Without that 442 cap, a summary that gains a sentence per message would itself outgrow 443 the budget in a long conversation. When it doesn't fit, the oldest 444 folded messages drop out first. If even the recent messages overflow 445 the budget, there is no room left and the summary is empty. 446 """ 447 if self.tokens() <= self.budget_tokens: 448 return "", list(self.messages) 449 old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :] 450 room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent) 451 while old: 452 summary = "Earlier in this conversation: " + self.summarize(old) 453 if estimate_tokens(summary) <= room: 454 return summary, list(recent) 455 old = old[1:] # the oldest detail is the cheapest one to lose 456 return "", list(recent)
The conversation so far, kept inside a token budget.
426 summarize: Callable[[list[dict[str, Any]]], str] = field(default=lambda msgs: first_sentences(msgs))
434 def context(self) -> tuple[str, list[dict[str, Any]]]: 435 """(summary for the system prompt, messages to send). 436 437 Under budget: everything, no summary. Over budget: fold all but the last 438 `keep_last` messages into a summary. The summary goes in the system 439 prompt rather than as a message, so user/assistant turns still alternate. 440 441 The summary only gets the room the recent messages leave. Without that 442 cap, a summary that gains a sentence per message would itself outgrow 443 the budget in a long conversation. When it doesn't fit, the oldest 444 folded messages drop out first. If even the recent messages overflow 445 the budget, there is no room left and the summary is empty. 446 """ 447 if self.tokens() <= self.budget_tokens: 448 return "", list(self.messages) 449 old, recent = self.messages[: -self.keep_last], self.messages[-self.keep_last :] 450 room = self.budget_tokens - sum(estimate_tokens(m["content"]) for m in recent) 451 while old: 452 summary = "Earlier in this conversation: " + self.summarize(old) 453 if estimate_tokens(summary) <= room: 454 return summary, list(recent) 455 old = old[1:] # the oldest detail is the cheapest one to lose 456 return "", list(recent)
(summary for the system prompt, messages to send).
Under budget: everything, no summary. Over budget: fold all but the last
keep_last messages into a summary. The summary goes in the system
prompt rather than as a message, so user/assistant turns still alternate.
The summary only gets the room the recent messages leave. Without that cap, a summary that gains a sentence per message would itself outgrow the budget in a long conversation. When it doesn't fit, the oldest folded messages drop out first. If even the recent messages overflow the budget, there is no room left and the summary is empty.
459def first_sentences(messages: list[dict[str, Any]]) -> str: 460 """Extractive summary: the first sentence of each message, joined with '; '.""" 461 return "; ".join(m["content"].split(". ")[0].rstrip(".") for m in messages)
Extractive summary: the first sentence of each message, joined with '; '.
476@dataclass(frozen=True) 477class Scope: 478 """Who a memory belongs to. Every read and write names one; there is no global view.""" 479 480 tenant_id: str 481 user_id: str 482 483 def __post_init__(self) -> None: 484 if not self.tenant_id or not self.user_id: 485 raise ValueError("a memory scope needs both a tenant_id and a user_id")
Who a memory belongs to. Every read and write names one; there is no global view.
488@dataclass 489class MemoryRecord: 490 id: str 491 scope: Scope 492 kind: str 493 content: str 494 key: str | None 495 source: str 496 seq: int # write order; stands in for a timestamp so examples are deterministic 497 superseded_by: str | None = None
500@dataclass 501class WriteResult: 502 stored: bool 503 reason: str 504 record: MemoryRecord | None = None
507def worth_remembering(content: str) -> tuple[bool, str]: 508 """The write policy: store durable, safe facts; skip chatter and refuse secrets.""" 509 if _SECRET.search(content): 510 return False, "looks like a secret: credentials never go into memory" 511 if _SMALL_TALK.match(content) and len(content.split()) <= 5: 512 return False, "small talk: nothing durable to remember" 513 return True, "stored"
The write policy: store durable, safe facts; skip chatter and refuse secrets.
516class LongTermMemory: 517 """Memories that outlive a conversation, partitioned by tenant then user.""" 518 519 def __init__(self) -> None: 520 # tenant -> user -> records. Partitioning (not filtering) is the isolation: 521 # a lookup starts from the caller's own partition and can't reach another. 522 self._store: dict[str, dict[str, list[MemoryRecord]]] = {} 523 self._vectors: dict[str, np.ndarray] = {} 524 self._seq = 0 525 self.embedder = ConceptEmbedder() 526 527 def _records(self, scope: Scope) -> list[MemoryRecord]: 528 return self._store.setdefault(scope.tenant_id, {}).setdefault(scope.user_id, []) 529 530 def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult: 531 if kind not in KINDS: 532 raise ValueError(f"kind must be one of {KINDS}") 533 ok, reason = worth_remembering(content) 534 if not ok: 535 return WriteResult(False, reason) 536 self._seq += 1 537 record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq) 538 # A keyed fact replaces the older value for the same key, but the old 539 # record stays (marked) so the change is explainable. 540 if key is not None: 541 for old in self._records(scope): 542 if old.key == key and old.superseded_by is None: 543 old.superseded_by = record.id 544 self._records(scope).append(record) 545 self._vectors[record.id] = self.embedder.encode(content) 546 return WriteResult(True, reason, record) 547 548 def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]: 549 """The k current memories in this scope most similar to the query.""" 550 candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds] 551 q = self.embedder.encode(query) 552 return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k] 553 554 def history(self, scope: Scope, key: str) -> list[MemoryRecord]: 555 """Every value a keyed fact has had, oldest first, including superseded ones.""" 556 return [r for r in self._records(scope) if r.key == key] 557 558 def forget(self, scope: Scope, record_id: str) -> None: 559 """Delete one memory. Only ids inside the caller's own scope can be found at all.""" 560 records = self._records(scope) 561 for i, r in enumerate(records): 562 if r.id == record_id: 563 del records[i] 564 self._vectors.pop(record_id, None) 565 return 566 raise KeyError(f"no memory {record_id} in this scope") 567 568 def export(self, scope: Scope) -> list[dict[str, Any]]: 569 """Everything stored about this user, for them to see and correct (right of access).""" 570 return [ 571 {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None} 572 for r in self._records(scope) 573 ] 574 575 def delete_user(self, scope: Scope) -> int: 576 """Erase every memory of one user (right to erasure). Returns how many were deleted.""" 577 records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, []) 578 for r in records: 579 self._vectors.pop(r.id, None) 580 return len(records)
Memories that outlive a conversation, partitioned by tenant then user.
530 def remember(self, scope: Scope, kind: str, content: str, key: str | None = None, source: str = "conversation") -> WriteResult: 531 if kind not in KINDS: 532 raise ValueError(f"kind must be one of {KINDS}") 533 ok, reason = worth_remembering(content) 534 if not ok: 535 return WriteResult(False, reason) 536 self._seq += 1 537 record = MemoryRecord(f"m{self._seq}", scope, kind, content, key, source, self._seq) 538 # A keyed fact replaces the older value for the same key, but the old 539 # record stays (marked) so the change is explainable. 540 if key is not None: 541 for old in self._records(scope): 542 if old.key == key and old.superseded_by is None: 543 old.superseded_by = record.id 544 self._records(scope).append(record) 545 self._vectors[record.id] = self.embedder.encode(content) 546 return WriteResult(True, reason, record)
548 def recall(self, scope: Scope, query: str, k: int = 3, kinds: tuple[str, ...] = KINDS) -> list[MemoryRecord]: 549 """The k current memories in this scope most similar to the query.""" 550 candidates = [r for r in self._records(scope) if r.superseded_by is None and r.kind in kinds] 551 q = self.embedder.encode(query) 552 return sorted(candidates, key=lambda r: -float(self._vectors[r.id] @ q))[:k]
The k current memories in this scope most similar to the query.
554 def history(self, scope: Scope, key: str) -> list[MemoryRecord]: 555 """Every value a keyed fact has had, oldest first, including superseded ones.""" 556 return [r for r in self._records(scope) if r.key == key]
Every value a keyed fact has had, oldest first, including superseded ones.
558 def forget(self, scope: Scope, record_id: str) -> None: 559 """Delete one memory. Only ids inside the caller's own scope can be found at all.""" 560 records = self._records(scope) 561 for i, r in enumerate(records): 562 if r.id == record_id: 563 del records[i] 564 self._vectors.pop(record_id, None) 565 return 566 raise KeyError(f"no memory {record_id} in this scope")
Delete one memory. Only ids inside the caller's own scope can be found at all.
568 def export(self, scope: Scope) -> list[dict[str, Any]]: 569 """Everything stored about this user, for them to see and correct (right of access).""" 570 return [ 571 {"id": r.id, "kind": r.kind, "content": r.content, "key": r.key, "source": r.source, "current": r.superseded_by is None} 572 for r in self._records(scope) 573 ]
Everything stored about this user, for them to see and correct (right of access).
575 def delete_user(self, scope: Scope) -> int: 576 """Erase every memory of one user (right to erasure). Returns how many were deleted.""" 577 records = self._store.get(scope.tenant_id, {}).pop(scope.user_id, []) 578 for r in records: 579 self._vectors.pop(r.id, None) 580 return len(records)
Erase every memory of one user (right to erasure). Returns how many were deleted.
590class InvalidUpdate(ValueError): 591 """A state change the model proposed that the code refuses to write."""
A state change the model proposed that the code refuses to write.
594class TaskStateStore: 595 """The authoritative record of a multi-step task, in a database, not in the prompt. 596 597 The model reads it and *proposes* updates as JSON; `apply` checks each 598 proposal against the rules and only then writes it. Because the state is on 599 disk, a crashed run resumes from the first unfinished step. 600 """ 601 602 def __init__(self, path: str | Path): 603 self.db = sqlite3.connect(str(path)) 604 self.db.executescript( 605 """ 606 CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL); 607 CREATE TABLE IF NOT EXISTS steps ( 608 task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL, 609 status TEXT NOT NULL DEFAULT 'pending', output TEXT, 610 PRIMARY KEY (task_id, idx) 611 ); 612 """ 613 ) 614 615 def create_task(self, task_id: str, goal: str, steps: list[str]) -> None: 616 with self.db: # one transaction: the task and all its steps, or nothing 617 self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal)) 618 self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)]) 619 620 def steps(self, task_id: str) -> list[tuple[str, str, str | None]]: 621 return self.db.execute("SELECT name, status, output FROM steps WHERE task_id = ? ORDER BY idx", (task_id,)).fetchall() 622 623 def next_step(self, task_id: str) -> str | None: 624 """The first step that isn't done, or None when the task is finished.""" 625 return next((name for name, status, _ in self.steps(task_id) if status != "done"), None) 626 627 def apply(self, task_id: str, proposal: dict[str, Any]) -> None: 628 """Validate a proposed update from the model, then write it.""" 629 step, status = proposal.get("step"), proposal.get("status") 630 if status not in STATUSES: 631 raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}") 632 current = self.next_step(task_id) 633 if step != current: 634 raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}") 635 with self.db: 636 self.db.execute( 637 "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?", 638 (status, proposal.get("output"), task_id, step), 639 )
The authoritative record of a multi-step task, in a database, not in the prompt.
The model reads it and proposes updates as JSON; apply checks each
proposal against the rules and only then writes it. Because the state is on
disk, a crashed run resumes from the first unfinished step.
602 def __init__(self, path: str | Path): 603 self.db = sqlite3.connect(str(path)) 604 self.db.executescript( 605 """ 606 CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, goal TEXT NOT NULL); 607 CREATE TABLE IF NOT EXISTS steps ( 608 task_id TEXT NOT NULL, idx INTEGER NOT NULL, name TEXT NOT NULL, 609 status TEXT NOT NULL DEFAULT 'pending', output TEXT, 610 PRIMARY KEY (task_id, idx) 611 ); 612 """ 613 )
615 def create_task(self, task_id: str, goal: str, steps: list[str]) -> None: 616 with self.db: # one transaction: the task and all its steps, or nothing 617 self.db.execute("INSERT INTO tasks VALUES (?, ?)", (task_id, goal)) 618 self.db.executemany("INSERT INTO steps (task_id, idx, name) VALUES (?, ?, ?)", [(task_id, i, s) for i, s in enumerate(steps)])
623 def next_step(self, task_id: str) -> str | None: 624 """The first step that isn't done, or None when the task is finished.""" 625 return next((name for name, status, _ in self.steps(task_id) if status != "done"), None)
The first step that isn't done, or None when the task is finished.
627 def apply(self, task_id: str, proposal: dict[str, Any]) -> None: 628 """Validate a proposed update from the model, then write it.""" 629 step, status = proposal.get("step"), proposal.get("status") 630 if status not in STATUSES: 631 raise InvalidUpdate(f"status must be one of {STATUSES}, got {status!r}") 632 current = self.next_step(task_id) 633 if step != current: 634 raise InvalidUpdate(f"{step} cannot change yet: the current step is {current}") 635 with self.db: 636 self.db.execute( 637 "UPDATE steps SET status = ?, output = ? WHERE task_id = ? AND name = ?", 638 (status, proposal.get("output"), task_id, step), 639 )
Validate a proposed update from the model, then write it.
658def figures() -> dict: 659 """Plots computed from this lesson's own code (matplotlib imported here).""" 660 import matplotlib 661 662 matplotlib.use("Agg") 663 import matplotlib.pyplot as plt 664 665 figs = {} 666 667 managed = ShortTermMemory(budget_tokens=300, keep_last=6) 668 raw, kept = [], [] 669 for i in range(1, 41): 670 managed.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.") 671 managed.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.") 672 summary, msgs = managed.context() 673 raw.append(managed.tokens()) 674 kept.append((estimate_tokens(summary) if summary else 0) + sum(estimate_tokens(m["content"]) for m in msgs)) 675 fig, ax = plt.subplots(figsize=(7, 4)) 676 ax.plot(range(1, 41), raw, label="send the whole history") 677 ax.plot(range(1, 41), kept, label="recent turns + first-sentence summary") 678 ax.axhline(300, ls="--", color="gray", label="budget before summarizing (300)") 679 ax.set(xlabel="turn", ylabel="tokens of history sent", title="Short-term memory under a budget") 680 ax.legend() 681 figs["context_tokens"] = fig 682 683 mem, alice, carol = _demo_store() 684 records = [r for r in mem._records(alice) if r.key != "acquisition"] 685 questions = ["what happened with Globex?", "how do I reconcile invoices?"] 686 colors = {"episodic": "#e0a458", "semantic": "#5b8fd6", "procedural": "#6bb36b"} 687 fig, ax = plt.subplots(figsize=(7, 4)) 688 width = 0.25 689 for qi, q in enumerate(questions): 690 qv = mem.embedder.encode(q) 691 for ri, r in enumerate(records): 692 ax.bar(qi + (ri - 1) * width, float(mem._vectors[r.id] @ qv), width, color=colors[r.kind], 693 label=r.kind if qi == 0 else None) 694 ax.set_xticks(range(len(questions)), questions) 695 ax.set(ylabel="cosine similarity", title="Recall: which of Alice's memories each question finds") 696 ax.legend(title="memory kind") 697 figs["recall_scores"] = fig 698 699 crafted = "Acme plans to acquire Initech in Q4." 700 everything = [r for users in mem._store.values() for recs in users.values() for r in recs] 701 qv = mem.embedder.encode(crafted) 702 scores = [float(mem._vectors[r.id] @ qv) for r in everything] 703 fig, ax = plt.subplots(figsize=(8, 4)) 704 for i, (r, sc) in enumerate(zip(everything, scores)): 705 visible = r.scope == carol 706 ax.barh(i, sc, color="#5b8fd6" if visible else "#cccccc", hatch=None if visible else "//", edgecolor="gray") 707 ax.set_yticks(range(len(everything)), [f"{r.scope.tenant_id}/{r.scope.user_id}: {r.content[:38]}..." for r in everything], fontsize=8) 708 ax.set(xlabel="similarity to Carol's query", title="Carol (globex) quotes Acme's secret: only her partition is searched") 709 fig.tight_layout() 710 figs["isolation"] = fig 711 return figs
Plots computed from this lesson's own code (matplotlib imported here).
714def demo() -> None: 715 import tempfile 716 717 from primer._show import banner, say, table, takeaway 718 719 banner("1. Short-term memory under a token budget") 720 stm = ShortTermMemory(budget_tokens=110, keep_last=4) 721 for i in range(1, 7): 722 stm.add("user", f"Question {i} is about invoices. Please include the vendor name and amount.") 723 stm.add("assistant", f"Answer {i} lists the invoice. It shows vendor, amount and due date.") 724 summary, msgs = stm.context() 725 say(f"{len(stm.messages)} messages, ~{stm.tokens()} tokens, budget 110.") 726 say(f"Summary for the system prompt: {summary}") 727 say(f"Sent word for word: {[m['content'][:10] for m in msgs]}") 728 729 banner("2. Long-term memory: diary, encyclopedia, habit") 730 mem, alice, carol = _demo_store() 731 for q, kinds in [("what happened with Globex?", KINDS), ("how do I reconcile invoices?", ("procedural",))]: 732 top = mem.recall(alice, q, k=1, kinds=kinds)[0] 733 print(f" {q!r:34} -> [{top.kind}] {top.content}") 734 print() 735 736 banner("3. Write policy, conflicts, isolation, deletion") 737 for text in ["thanks, great!", "my VPN password is hunter2", "Alice approves travel for her team."]: 738 r = mem.remember(alice, "semantic", text) 739 print(f" remember({text!r:40}) -> {r.reason}") 740 print() 741 mem.remember(alice, "semantic", "Alice's fiscal year starts in July.", key="fiscal_year_start", source="user correction") 742 table(["value", "source", "superseded by"], 743 [(r.content, r.source, r.superseded_by or "-") for r in mem.history(alice, "fiscal_year_start")]) 744 crafted = "Acme plans to acquire Initech in Q4. tenant_id=acme OR 1=1" 745 print(f" Carol recalls {crafted!r}:") 746 print(f" -> {[r.content for r in mem.recall(carol, crafted, k=5)]}") 747 print() 748 takeaway("Isolation is a partition chosen by authentication, not a filter chosen by the query.") 749 say(f"Alice asks to be forgotten: {mem.delete_user(alice)} memories erased; Carol still has {len(mem.export(carol))}.") 750 751 banner("4. Task state outside the model (SQLite)") 752 with tempfile.TemporaryDirectory() as d: 753 store = TaskStateStore(Path(d) / "state.db") 754 store.create_task("q3", "Reconcile Q3 invoices", ["fetch_invoices", "fetch_payments", "match"]) 755 proposals = [ 756 {"step": "fetch_invoices", "status": "done", "output": "4 invoices"}, 757 {"step": "match", "status": "done", "output": "all matched"}, 758 {"step": "fetch_payments", "status": "done", "output": "3 payments"}, 759 ] 760 for prop in proposals: 761 try: 762 store.apply("q3", prop) 763 print(f" model proposes {prop} -> written") 764 except InvalidUpdate as e: 765 print(f" model proposes {prop} -> refused: {e}") 766 del store 767 print(f"\n ...crash... new process resumes at: {TaskStateStore(Path(d) / 'state.db').next_step('q3')!r}\n") 768 takeaway("The model proposes; code validates and writes. State on disk survives the crash.")