primer.agents.deployment
Safe deployment: letting an agent act in the real world, one earned step at a time
Run: python -m primer.agents.deployment
This lesson builds on action guardrails from primer.agents.guardrails, on
the release gate from primer.agents.evals and on traces from
primer.agents.observability.
Level 1: The practitioner's guide
In one sentence. Safe deployment means letting an agent act in the real world one earned step at a time: it proves itself in shadow, then proposes, then acts alone on low-risk work, while every release goes to a few users first, every action can be stopped and rate-limited, and every action is written to a log nobody can quietly edit.
When you need it. The moment an agent's output stops being a draft and
becomes an action: a refund issued, an email sent, a record changed in the
system a company runs on. Offline evals (primer.agents.evals) gate what
ships; this lesson is about limiting the damage from whatever the evals
missed, because real traffic always holds inputs the golden set didn't.
The tell: someone asks "what's the worst this could do in the next five
minutes?" and the answer is "we'd notice eventually". This lesson's worked
example shows why the first step is to watch rather than act: in shadow
mode over ten support tickets the agent agrees with staff 80% of the time
overall, but the split is 100% on replies and 60% on refunds, so the right
first grant of autonomy is replies, not refunds, and only the breakdown
shows it. You don't need canaries and hash chains for an internal tool
that drafts text a person edits; you need all of it for anything that
moves money or changes records, and regulated work needs the audit log
whatever the risk.
Your options. Seven mechanisms, from the ones that cost nothing to run to the ones that hold up under audit. They stack rather than compete:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Shadow mode | The agent records what it would do on real inputs; people keep doing the work | Zero user impact and a clean comparison against what staff actually did, at full volume | Staff carry the whole load meanwhile; you need to log proposals and compare | Your agent loop, plus a comparison job |
| Human approval of each action | The agent proposes; a person approves before anything executes | Nothing executes unseen | Review load, and reviewers who start rubber-stamping | The reviewers' existing workflow |
| Graduated autonomy | Acts alone on low-risk actions once a full window of recent decisions clears a bar; high-risk actions still go to a person; any incident demotes at once | Trust is earned on evidence and taken back on one bad event | Classifying every action's risk; keeping the window's statistics | Your action gate |
| Prompts as code | Every prompt version is named by a hash of its text, reviewed, evaluated, and one step from rollback | Any behaviour change ties to the exact edit that caused it | A registry and the discipline to use it | Your code and configuration |
| Canary release | A new version goes to 1% of users, then 5, 25, 50 and 100, growing only while its success rate stays within a margin of the current version's | At worst a few percent of users saw the bad version for one step | Enough traffic per step to judge (100 canary tasks in this lesson), and metrics split by version | The router and the metrics store |
| Kill switch and rate limit | Stop the agent instantly, per tenant or for everyone; cap actions per minute with a token bucket | A bounded blast radius even when everything else has failed | Someone on call to flip the switch; a burst above the cap is refused | In front of every action |
| Tamper-evident audit log | Each entry stores the hash of the one before, so any edit or deletion breaks every later link | Tampering is detectable; with write-once storage, provable | Storage that can't be modified, and a link from each entry to its trace | Storage |
How to choose. Go by what the action can do if it's wrong, and how fast you'd know.
- Reversible and low value (a reply, a draft, a label): shadow briefly, then approval, then autonomy, with the eval gate and a canary on every prompt or model change.
- Irreversible or high value (a payment, a deletion, an external email): approval stays, whatever the agent's record, above a value threshold you set on purpose. Autonomy applies below it.
- Many customers on one platform: kill switches and rate limits per tenant, so one customer's bad day never becomes everyone's.
- Regulated work: the audit log from day one, on write-once storage, linked to traces, with the agent acting under its own service identity and the narrowest permissions it can do the job with.
- Whatever you pick, promote on a sustained record over a window of decisions (this lesson uses the last 50), never a lucky streak, and demote on a single incident.
What it costs. Shadow mode costs the humans nothing extra and costs you patience: in the lesson's simulation the rolling agreement first clears the 95% bar after 98 decisions, and a window of 50 is the least you can judge on. Approval mode costs reviewer time on every action. A canary costs traffic and time per step; the lesson's controller refuses to judge on fewer than 100 canary tasks, and the SRE Workbook's chapter on canarying makes the trade explicit: a 5% canary with a 20% error rate hurts 1% of requests overall, which is the whole point, but a daily release cycle can't afford a week-long canary. A rate limit costs refused actions during a burst (a bucket of 5 refilling at one per second lets 5 of a burst of 8 through and refuses 3). The audit log costs a hash per entry and storage you can't reclaim. Rollback of a prompt costs nothing, because the previous version is already in the registry. Against all of that, the cost of not having them is the incident: a refund issued on a bad prompt that no switch could stop and no log can show.
What breaks.
- Promoting on a streak. Twenty good decisions in a row is luck at 95%. Require a full window.
- A canary judged too early. The lesson's bad version, truly 91% successful, passes the 1% step by luck on 100 tasks and is caught at 5%. Small steps are cheap to get wrong; that's why there are several.
- Users flipping between versions. Assign users by hashing their id into a bucket, so nobody changes version mid-conversation.
- Reviewers who stop reading. Approval mode can bias people toward accepting the agent's suggestion. Sample approvals for a second look, and keep incident demotion automatic.
- Retried writes that double-post. A canary or a retry that re-sends "create order" makes two orders. Use idempotency keys before granting any write.
- An ordinary log table. Anyone with write access can edit a row. Chain the hashes and anchor the latest one somewhere separate.
- A working agent nobody uses. Adoption is part of the job: build with the people whose work it touches, show sources, and make handing a case to a person one click.
In the wild. Canarying comes from site reliability engineering: the Google SRE Workbook defines it as "a partial and time-limited deployment of a change in a service and its evaluation", and its advice to rank metrics by how well they show user-visible problems is why this lesson compares success rates rather than latency alone. Feature flags are the mechanism behind both canaries and kill switches; Martin Fowler's taxonomy separates release, experiment, ops and permissioning toggles, with the kill switch as a long-lived ops toggle, and the pattern is what feature-flag services and OpenFeature, a vendor-neutral specification for feature flagging, exist to implement. The token bucket is a textbook rate-limiting algorithm, linked in Further reading. The hash chain is Haber and Stornetta's 1991 construction for time-stamping documents, the ancestor of every tamper-evident log and of blockchains. The NIST AI Risk Management Framework (released in January 2023) organises this whole lesson's concerns into four functions, Govern, Map, Measure and Manage, and is the vocabulary compliance teams will use for it.
Go deeper. Level 2 builds each gate in plain code: agreement counted
decision by decision, the autonomy state machine with its promotion window,
a content-addressed prompt registry, a canary controller with the rollback
rule worked in numbers, the token bucket formula, and a hash chain you can
break and watch verify() find the break. If you only needed to plan a
rollout, you are done.
Level 2: How it works, from scratch
A new pilot doesn't get the captain's seat on day one. First they sit in the right seat and call out every move they would make while the captain flies. Then they fly with the captain's hands hovering over the controls. Then they fly routine legs alone, while the captain still handles storms. Trust is earned in steps, with evidence at each one, and it can be taken back.
Deploying an agent that takes real actions (refunds, emails, changes to a company's records) works the same way. This lesson builds each mechanism: graduated autonomy, prompt versioning, canary releases, kill switches, rate limits, and a tamper-evident audit log.
flowchart LR P[Proposed action] --> K{Kill switch<br/>on?} K -->|yes| X[Blocked] K -->|no| R{Rate limit<br/>allows it?} R -->|no| X R -->|yes| A{Autonomy level<br/>and risk allow<br/>acting alone?} A -->|no| H[Queue for<br/>human approval] A -->|yes| E[Execute] H -->|approved| E E --> L[Append to<br/>audit log]
Reading it: every action the agent proposes passes the same gates in the same order, cheapest and most absolute first. The kill switch stops everything instantly; the rate limit bounds how much damage even a misbehaving agent can do per minute; the autonomy level decides whether a person must approve. Whatever executes is written to the audit log.
In code: each gate is one call: KillSwitch.allowed, then
TokenBucket.allow, then AutonomyController.may_act_alone, and finally
AuditLog.append for whatever executes.
Graduated autonomy: shadow mode, approval, then autonomy
Everyday picture. The trainee pilot again: call out moves (shadow), fly with the captain ready to take over (approval), fly routine legs alone (autonomy for low-risk actions), and hand back the controls after any incident.
In shadow mode the agent runs on real inputs and records what it would do, but takes no action; people keep doing the work. You compare its decisions with theirs. When agreement is high enough, it proposes actions and a person approves each one. When approvals are nearly always "yes", it may act alone on low-risk actions, while high-risk ones keep needing a person. Any incident or drop in quality sends it back a level.
Worked example. Ten support tickets in shadow mode:
| Agent proposed | Human did | Agree? |
|---|---|---|
| refund | refund | yes |
| refund | escalate | no |
| reply ×5 | reply ×5 | yes ×5 |
| refund | refund | yes |
| refund | escalate | no |
| refund | refund | yes |
Overall agreement is 8/10 = 80%. Broken down by what the agent proposed, replies agree 5/5 = 100% but refunds only 3/5 = 60%. The breakdown tells you what to let the agent do first: replies, not refunds.
Level 3: the formula and its symbols
$$ \text{agreement} = \frac{1}{N} \sum_{i=1}^{N} \mathbb{1}[a_i = h_i] $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $N$ | number of shadow decisions compared | 10 |
| $a_i$ | what the agent proposed for case $i$ | refund |
| $h_i$ | what the human actually did for case $i$ | escalate |
| $\mathbb{1}[\ldots]$ | 1 if the condition is true, 0 if not (an indicator) | 0 for case 2 |
| $\sum_{i=1}^{N}$ | add up over all cases |
In words: agreement is the share of cases where the agent proposed exactly what the human did.
On the worked example: eight of the ten indicators are 1, so agreement = 8/10 = 0.8.
Level 3: in Python
In Python:
a = ["refund", "refund"] + ["reply"] * 5 + ["refund", "refund", "refund"]
h = ["refund", "escalate"] + ["reply"] * 5 + ["refund", "escalate", "refund"]
# 𝟙[a_i = h_i]
indicators = [1 if a_i == h_i else 0 for a_i, h_i in zip(a, h)]
indicators # → [1, 0, 1, 1, 1, 1, 1, 1, 0, 1]
# (1/N) Σ over i
sum(indicators) / len(indicators) # → 0.8
stateDiagram-v2 state "Shadow mode: record, don't act" as Shadow state "Human approves each action" as Approve state "Autonomous for low-risk actions" as Auto [*] --> Shadow Shadow --> Approve: agreement ≥ 95% over the last 50 decisions Approve --> Auto: ≥ 98% of the last 50 approved, no incidents Auto --> Approve: incident or quality drift Approve --> Shadow: incident
Reading it: each box is a level of trust; each arrow is labelled with the evidence needed to move. Promotions need a sustained record over a window of recent decisions, not a lucky streak; demotions happen at once, on a single incident. Even at the top level, high-risk actions (large refunds, deletions) keep needing a person.
Reading it: the x-axis counts shadow decisions; the blue line is agreement with humans over the last 50 of them. The dashed line is the 95% bar. Early on the window is too short to judge (grey region); the marker shows where the rolling agreement first clears the bar and the agent is promoted to proposing actions for approval.
In code: shadow_agreement computes agreement overall and per proposed
action. AutonomyController is the state diagram: AutonomyController.record_shadow
and AutonomyController.record_approval promote once a full window clears
the bar, while AutonomyController.record_incident and
AutonomyController.record_quality demote at once.
Prompts are code
Everyday picture. A restaurant's recipe binder where every recipe change gets a new revision number, the old version stays in the binder, and the kitchen can switch back in one step if customers complain.
A change to a prompt can break behaviour as badly as a code change, so treat
it the same way: version it, review it, run the evals before release
(primer.agents.evals), and keep the previous version one step away.
PromptRegistry is content-addressed: a version's name includes a
hash of its text. A hash (here SHA-256) is a fixed-length fingerprint:
the same text always gives the same fingerprint, and any change, even one
character, gives a completely different one.
Worked example. "Be concise." hashes to ab018bdb…, so its version is
support@ab018bdb. Registering the same text again returns the same
version; "Be concise and friendly." gets a new one. Every trace records
which version produced each model call (primer.agents.observability), so
a behaviour change can be tied to the exact prompt edit that caused it.
In code: PromptRegistry.register hashes the text into a version name
and makes it active; PromptRegistry.rollback steps back one version;
PromptRegistry.active_text returns the prompt currently in use.
Canary releases: try it on a few users first
Everyday picture. Miners once carried a canary into the mine: if the air turned bad, the canary showed it before the miners were harmed. A canary release sends a small share of traffic to the new version, compares its metrics with the current version (the control), and grows the share only while the canary stays healthy.
Worked example. Steps 1% → 5% → 25% → 50% → 100%, allowed drop 2 points, at least 100 canary tasks before judging:
| Control success | Canary success | Canary tasks | Decision |
|---|---|---|---|
| 95% | 95% | 100 | advance from 1% to 5% |
| 95% | 90% | 100 | 90 < 95 − 2: roll back to 0% |
| 95% | 90% | 10 | too few samples: stay at 1% |
Level 3: the formula and its symbols
$$ \text{roll back if } \hat{p}_{\text{canary}} < \hat{p}_{\text{control}} - \delta $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $\hat{p}_{\text{canary}}$ | measured success rate on canary traffic (the hat means "measured from samples") | 0.90 |
| $\hat{p}_{\text{control}}$ | measured success rate on the current version | 0.95 |
| $\delta$ | the largest drop you'll tolerate (Greek delta) | 0.02 |
In words: roll back when the new version's success rate falls more than the allowed margin below the current version's.
On the worked example: 0.90 < 0.95 − 0.02 = 0.93, so roll back.
Level 3: in Python
In Python:
p_canary, p_control, delta = 0.90, 0.95, 0.02
round(p_control - delta, 2) # → 0.93
# roll back?
p_canary < p_control - delta # → True
Users are assigned to the canary by hashing their id into one of 100 buckets: the same user always lands in the same bucket, so nobody flips between versions mid-conversation, and growing from 5% to 25% only adds users.
sequenceDiagram participant D as Deploy system participant R as Router participant M as Metrics D->>R: send 1% of users to v2 R->>M: success rates, v1 vs v2 M-->>D: v2 95%, v1 95%: healthy D->>R: grow to 5% R->>M: success rates, v1 vs v2 M-->>D: v2 91%, v1 95%: worse by 4 points D->>R: roll back: 0% to v2 D-->>D: alert the owner with traces
Reading it: the deploy system only ever acts on measurements. Each growth step waits for enough traffic to judge, and a single bad reading sends everyone back to the old version automatically. At worst, a few percent of users saw the bad version for one step.
Reading it: each panel is one rollout. The x-axis is the rollout step (with the share of traffic on the canary); the lines are measured success rates for control and canary. The good version (left) tracks the control and reaches 100%. The bad version (right), truly 91% successful, gets lucky on the 100 tasks of the 1% step and advances, then falls below the tolerance band at 5% and is rolled back to 0%, marked by the red cross. Small steps are cheap to get wrong; that's why there are several.
In code: in_canary hashes a user id into one of 100 buckets.
CanaryController.observe accumulates success counts, and
CanaryController.step applies the rollback rule above: advance, wait for
more samples, or roll back. simulate_rollout drives a whole rollout for the
figure.
A feature flag is the same idea as an on/off switch in configuration: it turns a capability on for chosen tenants or users without redeploying, and off again just as fast.
Kill switches and rate limits: bound the blast radius
Everyday picture. The red emergency-stop button on a treadmill, and the turnstile at a stadium that lets people through only so fast however hard the crowd pushes.
A kill switch stops the agent instantly, for one tenant (one customer organisation) or for everyone. A rate limit caps actions per unit time, so even a malfunctioning agent can only send so many emails or change so many records per minute. The blast radius is how much damage a failure can do before someone notices; these two mechanisms bound it.
Worked example: the token bucket. Picture a jar that holds at most 5 tokens and gains 1 token per second. Each action takes a token; with no token, the action is refused. Starting full: 5 actions in a burst succeed, the 6th is refused. After 2 seconds the jar has 2 tokens again: 2 more succeed, the next is refused.
Level 3: the formula and its symbols
$$ b(t) = \min\bigl(C,\; b(t_0) + r\,(t - t_0)\bigr) $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $b(t)$ | tokens in the bucket at time $t$ | 2 at t = 2 s |
| $t_0$ | when the bucket was last updated | 0 s |
| $b(t_0)$ | tokens left then | 0 after the burst |
| $r$ | refill rate, tokens per second | 1 |
| $C$ | capacity: the largest burst allowed | 5 |
| $\min$ | the smaller of the two: the jar can't overflow |
In words: the tokens now are what was left plus what has dripped in since, but never more than the jar holds.
On the worked example: min(5, 0 + 1 × (2 − 0)) = 2 tokens, so two more actions are allowed.
Level 3: in Python
In Python:
# capacity, tokens per second
C, r = 5, 1
def b(t, t_0, b_t0):
# what was left plus the drip, capped at C
return min(C, b_t0 + r * (t - t_0))
# 2 s after the burst emptied the jar
b(2, 0, 0) # → 2
# a long wait refills only to capacity
b(60, 0, 0) # → 5
Reading it: the blue line is the number of tokens in the jar over time; green dots are allowed actions and red crosses refused ones. The opening burst drains the jar, refusals follow, and then actions trickle through at the refill rate: bursts are allowed, sustained floods are not.
In code: KillSwitch holds the global and per-tenant off switches.
TokenBucket.level is the formula for $b(t)$, and TokenBucket.allow
spends a token or refuses.
A tamper-evident audit log
Everyday picture. A ledger where every page ends with a wax seal pressed from the previous page's seal plus this page's contents. Change one number on page 12 and its seal no longer matches, and neither does any seal after it.
An audit log records which agent took which action, for which user, with what inputs. For regulated work, it must be tamper-evident: an alteration can't go unnoticed. A hash chain does this: each entry stores the hash of the previous entry, and its own hash covers its contents and that previous hash.
Level 3: the formula and its symbols
$$ h_i = \operatorname{SHA256}(h_{i-1} \,\Vert\, \text{entry}_i), \qquad h_0 = 00\ldots0 $$
(Entries are numbered from 1 here; verify() reports Python list
positions, which start at 0, so "entry 2" is position 1.)
Symbols
| Symbol | Meaning here |
|---|---|
| $h_i$ | the fingerprint (hash) stored with entry $i$ |
| $h_{i-1}$ | the previous entry's fingerprint |
| $\operatorname{SHA256}$ | the hash function: any input to a 64-hex-digit fingerprint |
| $\Vert$ | "joined together with" (concatenation) |
| $h_0$ | a fixed starting value for the first entry |
In words: each entry's fingerprint is the fingerprint of the previous fingerprint joined to this entry's contents.
On the worked example: three entries: refund A200 for 40, refund A201
for 15, escalate A300. Change entry 2's amount from 15 to 1,500 and
recomputing its fingerprint no longer gives the stored $h_2$, so verify()
reports position 1. Delete entry 2 instead, and entry 3 moves up to position
1; its stored previous fingerprint is $h_2$, the deleted entry's, which
doesn't match $h_1$, so the break is again reported at position 1.
Level 3: in Python
In Python:
import hashlib
def seal(prev, entry):
# SHA256(h_{i-1} ‖ entry_i)
return hashlib.sha256((prev + entry).encode()).hexdigest()
# h_0
log, prev = [], "0" * 64
for entry in ["refund A200 40", "refund A201 15", "escalate A300"]:
prev = seal(prev, entry)
# each entry is stored with its h_i
log.append((entry, prev))
def verify(log):
prev = "0" * 64
for position, (entry, h_i) in enumerate(log):
if seal(prev, entry) != h_i:
# the first broken link
return position
prev = h_i
verify(log) is None # → True
# entry 2's amount changed
verify([log[0], ("refund A201 1500", log[1][1]), log[2]]) # → 1
# entry 2 deleted
verify([log[0], log[2]]) # → 1
flowchart LR G["h0 = 000…"] --> E1["entry 1: refund A200, 40<br/>prev = h0<br/>h1 = SHA256(h0 ‖ entry 1)"] E1 --> E2["entry 2: refund A201, 15<br/>prev = h1<br/>h2 = SHA256(h1 ‖ entry 2)"] E2 --> E3["entry 3: escalate A300<br/>prev = h2<br/>h3 = SHA256(h2 ‖ entry 3)"]
Reading it: each box carries the fingerprint of the box before it, so the boxes are chained. Editing any box changes its fingerprint, which breaks the link to the next box, so verification walks the chain and stops at the first broken link. In production, also write the log to storage that can't be modified afterwards (write-once storage) and link each entry to its full trace.
In code: AuditLog.append stores the previous entry's hash in the new
entry and seals it with its own; AuditLog.verify walks the chain and
returns the position of the first broken link, or None.
Change management: adoption is part of the job
A technically working agent still fails if people don't trust or use it. Involve the people whose work it touches from shadow mode onwards (their corrections are your best eval data), design around their actual workflow, show sources and confidence so they can check answers, and make handing a case to a person one click.
In 20 seconds
- Earn autonomy in steps: shadow mode, then human approval of each action, then autonomy for low-risk actions only; demote on any incident or drift.
- Treat prompts as code: versioned (content hashes), reviewed, evaluated, and one step from rollback.
- Release to a small canary share first and grow it only while its metrics match the current version; roll back automatically otherwise.
- Bound the blast radius: kill switches per tenant and globally, and rate limits on actions.
- Keep a tamper-evident audit log (a hash chain) linking every action to its user, authority and trace.
Self-test questions
How would you roll out an agent that takes real actions in a company's ERP system (its finance and operations system of record)? Start read-only in shadow mode on real transactions, comparing proposals with what staff actually did, broken down by action type. Give the agent a dedicated service identity with the narrowest permissions, a sandbox or dry-run mode for writes, and idempotency keys so retries can't double-post. Move to proposing actions that a person approves in their existing workflow, starting with the low-value, reversible action types that scored best in shadow. Grant autonomy only for those, with value thresholds that still require approval, per-tenant kill switches, rate limits on writes, a canary rollout for every prompt or model change, eval gates in CI, and a tamper-evident audit log tied to traces. Demote automatically on incidents or drift, and keep the finance team involved throughout.
Why start in shadow mode instead of with human approval? Shadow mode costs the humans nothing extra and produces a clean comparison against what they actually did, at full volume, before the agent can influence anything. Approval mode adds review load and can bias reviewers toward accepting the agent's suggestion.
What does a canary protect against that offline evals don't? Real traffic: inputs, integrations and user behaviour the golden set didn't anticipate. Evals gate the release; the canary limits the damage from whatever the evals missed.
Why a hash chain rather than an ordinary log table? Anyone with write access can quietly edit an ordinary row. With a hash chain, any edit or deletion breaks every later link, so tampering is detectable, and anchoring the latest hash somewhere separate (or using write-once storage) makes it provable.
The papers behind this lesson
- Haber & Stornetta, How to Time-Stamp a Digital Document (Journal of Cryptology, 1991), https://doi.org/10.1007/BF00196791. Introduced chaining each record to the hash of the one before, the idea behind tamper-evident logs (and, later, blockchains).
Further reading
- Google SRE workbook, canarying releases: https://sre.google/workbook/canarying-releases/
- Martin Fowler, feature toggles: https://martinfowler.com/articles/feature-toggles.html
- Token bucket algorithm: https://en.wikipedia.org/wiki/Token_bucket
- NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework
1r""" 2# Safe deployment: letting an agent act in the real world, one earned step at a time 3 4Run: `python -m primer.agents.deployment` 5 6This lesson builds on action guardrails from `primer.agents.guardrails`, on 7the release gate from `primer.agents.evals` and on traces from 8`primer.agents.observability`. 9 10## Level 1: The practitioner's guide 11 12**In one sentence.** Safe deployment means letting an agent act in the real 13world one earned step at a time: it proves itself in shadow, then proposes, 14then acts alone on low-risk work, while every release goes to a few users 15first, every action can be stopped and rate-limited, and every action is 16written to a log nobody can quietly edit. 17 18**When you need it.** The moment an agent's output stops being a draft and 19becomes an action: a refund issued, an email sent, a record changed in the 20system a company runs on. Offline evals (`primer.agents.evals`) gate what 21ships; this lesson is about limiting the damage from whatever the evals 22missed, because real traffic always holds inputs the golden set didn't. 23The tell: someone asks "what's the worst this could do in the next five 24minutes?" and the answer is "we'd notice eventually". This lesson's worked 25example shows why the first step is to watch rather than act: in shadow 26mode over ten support tickets the agent agrees with staff 80% of the time 27overall, but the split is 100% on replies and 60% on refunds, so the right 28first grant of autonomy is replies, not refunds, and only the breakdown 29shows it. You don't need canaries and hash chains for an internal tool 30that drafts text a person edits; you need all of it for anything that 31moves money or changes records, and regulated work needs the audit log 32whatever the risk. 33 34**Your options.** Seven mechanisms, from the ones that cost nothing to run 35to the ones that hold up under audit. They stack rather than compete: 36 37| Option | What it does | What it guarantees | What it costs | Where it lives | 38|---|---|---|---|---| 39| Shadow mode | The agent records what it would do on real inputs; people keep doing the work | Zero user impact and a clean comparison against what staff actually did, at full volume | Staff carry the whole load meanwhile; you need to log proposals and compare | Your agent loop, plus a comparison job | 40| Human approval of each action | The agent proposes; a person approves before anything executes | Nothing executes unseen | Review load, and reviewers who start rubber-stamping | The reviewers' existing workflow | 41| Graduated autonomy | Acts alone on low-risk actions once a full window of recent decisions clears a bar; high-risk actions still go to a person; any incident demotes at once | Trust is earned on evidence and taken back on one bad event | Classifying every action's risk; keeping the window's statistics | Your action gate | 42| Prompts as code | Every prompt version is named by a hash of its text, reviewed, evaluated, and one step from rollback | Any behaviour change ties to the exact edit that caused it | A registry and the discipline to use it | Your code and configuration | 43| Canary release | A new version goes to 1% of users, then 5, 25, 50 and 100, growing only while its success rate stays within a margin of the current version's | At worst a few percent of users saw the bad version for one step | Enough traffic per step to judge (100 canary tasks in this lesson), and metrics split by version | The router and the metrics store | 44| Kill switch and rate limit | Stop the agent instantly, per tenant or for everyone; cap actions per minute with a token bucket | A bounded blast radius even when everything else has failed | Someone on call to flip the switch; a burst above the cap is refused | In front of every action | 45| Tamper-evident audit log | Each entry stores the hash of the one before, so any edit or deletion breaks every later link | Tampering is detectable; with write-once storage, provable | Storage that can't be modified, and a link from each entry to its trace | Storage | 46 47**How to choose.** Go by what the action can do if it's wrong, and how 48fast you'd know. 49 50- Reversible and low value (a reply, a draft, a label): shadow briefly, 51 then approval, then autonomy, with the eval gate and a canary on every 52 prompt or model change. 53- Irreversible or high value (a payment, a deletion, an external email): 54 approval stays, whatever the agent's record, above a value threshold you 55 set on purpose. Autonomy applies below it. 56- Many customers on one platform: kill switches and rate limits per 57 tenant, so one customer's bad day never becomes everyone's. 58- Regulated work: the audit log from day one, on write-once storage, 59 linked to traces, with the agent acting under its own service identity 60 and the narrowest permissions it can do the job with. 61- Whatever you pick, promote on a sustained record over a window of 62 decisions (this lesson uses the last 50), never a lucky streak, and demote 63 on a single incident. 64 65**What it costs.** Shadow mode costs the humans nothing extra and costs you 66patience: in the lesson's simulation the rolling agreement first clears the 6795% bar after 98 decisions, and a window of 50 is the least you can judge 68on. Approval mode costs reviewer time on every action. A canary costs 69traffic and time per step; the lesson's controller refuses to judge on 70fewer than 100 canary tasks, and the SRE Workbook's chapter on canarying 71makes the trade explicit: a 5% canary with a 20% error rate hurts 1% of 72requests overall, which is the whole point, but a daily release cycle 73can't afford a week-long canary. A rate limit costs refused actions during 74a burst (a bucket of 5 refilling at one per second lets 5 of a burst of 8 75through and refuses 3). The audit log costs a hash per entry and storage 76you can't reclaim. Rollback of a prompt costs nothing, because the previous 77version is already in the registry. Against all of that, the cost of not 78having them is the incident: a refund issued on a bad prompt that no 79switch could stop and no log can show. 80 81**What breaks.** 82 83- **Promoting on a streak.** Twenty good decisions in a row is luck at 95%. 84 Require a full window. 85- **A canary judged too early.** The lesson's bad version, truly 91% 86 successful, passes the 1% step by luck on 100 tasks and is caught at 5%. 87 Small steps are cheap to get wrong; that's why there are several. 88- **Users flipping between versions.** Assign users by hashing their id 89 into a bucket, so nobody changes version mid-conversation. 90- **Reviewers who stop reading.** Approval mode can bias people toward 91 accepting the agent's suggestion. Sample approvals for a second look, and 92 keep incident demotion automatic. 93- **Retried writes that double-post.** A canary or a retry that re-sends 94 "create order" makes two orders. Use idempotency keys before granting any 95 write. 96- **An ordinary log table.** Anyone with write access can edit a row. Chain 97 the hashes and anchor the latest one somewhere separate. 98- **A working agent nobody uses.** Adoption is part of the job: build with 99 the people whose work it touches, show sources, and make handing a case 100 to a person one click. 101 102**In the wild.** Canarying comes from site reliability engineering: the 103Google SRE Workbook defines it as "a partial and time-limited deployment of 104a change in a service and its evaluation", and its advice to rank metrics by 105how well they show user-visible problems is why this lesson compares 106success rates rather than latency alone. Feature flags are the mechanism 107behind both canaries and kill switches; Martin Fowler's taxonomy separates 108release, experiment, ops and permissioning toggles, with the kill switch as 109a long-lived ops toggle, and the pattern is what feature-flag services and 110OpenFeature, a vendor-neutral specification for feature flagging, exist to 111implement. The token bucket is a textbook rate-limiting algorithm, linked 112in Further reading. The hash chain is Haber and Stornetta's 1991 113construction for time-stamping documents, the ancestor of every 114tamper-evident log and of blockchains. The NIST AI Risk Management 115Framework (released in January 2023) organises this whole lesson's 116concerns into four functions, Govern, Map, Measure and Manage, and is the 117vocabulary compliance teams will use for it. 118 119**Go deeper.** Level 2 builds each gate in plain code: agreement counted 120decision by decision, the autonomy state machine with its promotion window, 121a content-addressed prompt registry, a canary controller with the rollback 122rule worked in numbers, the token bucket formula, and a hash chain you can 123break and watch `verify()` find the break. If you only needed to plan a 124rollout, you are done. 125 126## Level 2: How it works, from scratch 127 128A new pilot doesn't get the captain's seat on day one. First they sit in the 129right seat and call out every move they *would* make while the captain 130flies. Then they fly with the captain's hands hovering over the controls. 131Then they fly routine legs alone, while the captain still handles storms. 132Trust is earned in steps, with evidence at each one, and it can be taken 133back. 134 135Deploying an agent that takes real actions (refunds, emails, changes to a 136company's records) works the same way. This lesson builds each mechanism: 137graduated autonomy, prompt versioning, canary releases, kill switches, rate 138limits, and a tamper-evident audit log. 139 140```mermaid 141flowchart LR 142 P[Proposed action] --> K{Kill switch<br/>on?} 143 K -->|yes| X[Blocked] 144 K -->|no| R{Rate limit<br/>allows it?} 145 R -->|no| X 146 R -->|yes| A{Autonomy level<br/>and risk allow<br/>acting alone?} 147 A -->|no| H[Queue for<br/>human approval] 148 A -->|yes| E[Execute] 149 H -->|approved| E 150 E --> L[Append to<br/>audit log] 151``` 152 153**Reading it:** every action the agent proposes passes the same gates in 154the same order, cheapest and most absolute first. The kill switch stops 155everything instantly; the rate limit bounds how much damage even a 156misbehaving agent can do per minute; the autonomy level decides whether a 157person must approve. Whatever executes is written to the audit log. 158 159**In code:** each gate is one call: `KillSwitch.allowed`, then 160`TokenBucket.allow`, then `AutonomyController.may_act_alone`, and finally 161`AuditLog.append` for whatever executes. 162 163## Graduated autonomy: shadow mode, approval, then autonomy 164 165**Everyday picture.** The trainee pilot again: call out moves (shadow), fly 166with the captain ready to take over (approval), fly routine legs alone 167(autonomy for low-risk actions), and hand back the controls after any 168incident. 169 170In **shadow mode** the agent runs on real inputs and records what it 171*would* do, but takes no action; people keep doing the work. You compare its 172decisions with theirs. When agreement is high enough, it **proposes** 173actions and a person approves each one. When approvals are nearly always 174"yes", it may act alone on low-risk actions, while high-risk ones keep 175needing a person. Any incident or drop in quality sends it back a level. 176 177**Worked example.** Ten support tickets in shadow mode: 178 179| Agent proposed | Human did | Agree? | 180|---|---|---| 181| refund | refund | yes | 182| refund | escalate | **no** | 183| reply ×5 | reply ×5 | yes ×5 | 184| refund | refund | yes | 185| refund | escalate | **no** | 186| refund | refund | yes | 187 188Overall agreement is 8/10 = 80%. Broken down by what the agent proposed, 189replies agree 5/5 = 100% but refunds only 3/5 = 60%. The breakdown tells 190you *what* to let the agent do first: replies, not refunds. 191 192$$ 193\text{agreement} = \frac{1}{N} \sum_{i=1}^{N} \mathbb{1}[a_i = h_i] 194$$ 195 196**Symbols** 197 198| Symbol | Meaning here | Worked example | 199|---|---|---| 200| $N$ | number of shadow decisions compared | 10 | 201| $a_i$ | what the agent proposed for case $i$ | refund | 202| $h_i$ | what the human actually did for case $i$ | escalate | 203| $\mathbb{1}[\ldots]$ | 1 if the condition is true, 0 if not (an *indicator*) | 0 for case 2 | 204| $\sum_{i=1}^{N}$ | add up over all cases | | 205 206**In words:** agreement is the share of cases where the agent proposed 207exactly what the human did. 208 209**On the worked example:** eight of the ten indicators are 1, so agreement 210= 8/10 = 0.8. 211 212**In Python:** 213 214```python 215a = ["refund", "refund"] + ["reply"] * 5 + ["refund", "refund", "refund"] 216h = ["refund", "escalate"] + ["reply"] * 5 + ["refund", "escalate", "refund"] 217# 𝟙[a_i = h_i] 218indicators = [1 if a_i == h_i else 0 for a_i, h_i in zip(a, h)] 219indicators # → [1, 0, 1, 1, 1, 1, 1, 1, 0, 1] 220# (1/N) Σ over i 221sum(indicators) / len(indicators) # → 0.8 222``` 223 224```mermaid 225stateDiagram-v2 226 state "Shadow mode: record, don't act" as Shadow 227 state "Human approves each action" as Approve 228 state "Autonomous for low-risk actions" as Auto 229 [*] --> Shadow 230 Shadow --> Approve: agreement ≥ 95% over the last 50 decisions 231 Approve --> Auto: ≥ 98% of the last 50 approved, no incidents 232 Auto --> Approve: incident or quality drift 233 Approve --> Shadow: incident 234``` 235 236**Reading it:** each box is a level of trust; each arrow is labelled with 237the evidence needed to move. Promotions need a sustained record over a 238window of recent decisions, not a lucky streak; demotions happen at once, 239on a single incident. Even at the top level, high-risk actions (large 240refunds, deletions) keep needing a person. 241 242 243 244**Reading it:** the x-axis counts shadow decisions; the blue line is 245agreement with humans over the last 50 of them. The dashed line is the 95% 246bar. Early on the window is too short to judge (grey region); the marker 247shows where the rolling agreement first clears the bar and the agent is 248promoted to proposing actions for approval. 249 250**In code:** `shadow_agreement` computes agreement overall and per proposed 251action. `AutonomyController` is the state diagram: `AutonomyController.record_shadow` 252and `AutonomyController.record_approval` promote once a full window clears 253the bar, while `AutonomyController.record_incident` and 254`AutonomyController.record_quality` demote at once. 255 256## Prompts are code 257 258**Everyday picture.** A restaurant's recipe binder where every recipe 259change gets a new revision number, the old version stays in the binder, and 260the kitchen can switch back in one step if customers complain. 261 262A change to a prompt can break behaviour as badly as a code change, so treat 263it the same way: version it, review it, run the evals before release 264(`primer.agents.evals`), and keep the previous version one step away. 265`PromptRegistry` is **content-addressed**: a version's name includes a 266**hash** of its text. A hash (here SHA-256) is a fixed-length fingerprint: 267the same text always gives the same fingerprint, and any change, even one 268character, gives a completely different one. 269 270**Worked example.** `"Be concise."` hashes to `ab018bdb…`, so its version is 271`support@ab018bdb`. Registering the same text again returns the same 272version; `"Be concise and friendly."` gets a new one. Every trace records 273which version produced each model call (`primer.agents.observability`), so 274a behaviour change can be tied to the exact prompt edit that caused it. 275 276**In code:** `PromptRegistry.register` hashes the text into a version name 277and makes it active; `PromptRegistry.rollback` steps back one version; 278`PromptRegistry.active_text` returns the prompt currently in use. 279 280## Canary releases: try it on a few users first 281 282**Everyday picture.** Miners once carried a canary into the mine: if the air 283turned bad, the canary showed it before the miners were harmed. A **canary 284release** sends a small share of traffic to the new version, compares its 285metrics with the current version (the *control*), and grows the share only 286while the canary stays healthy. 287 288**Worked example.** Steps 1% → 5% → 25% → 50% → 100%, allowed drop 2 points, 289at least 100 canary tasks before judging: 290 291| Control success | Canary success | Canary tasks | Decision | 292|---|---|---|---| 293| 95% | 95% | 100 | advance from 1% to 5% | 294| 95% | 90% | 100 | 90 < 95 − 2: **roll back to 0%** | 295| 95% | 90% | 10 | too few samples: stay at 1% | 296 297$$ 298\text{roll back if } \hat{p}_{\text{canary}} < \hat{p}_{\text{control}} - \delta 299$$ 300 301**Symbols** 302 303| Symbol | Meaning here | Worked example | 304|---|---|---| 305| $\hat{p}_{\text{canary}}$ | measured success rate on canary traffic (the hat means "measured from samples") | 0.90 | 306| $\hat{p}_{\text{control}}$ | measured success rate on the current version | 0.95 | 307| $\delta$ | the largest drop you'll tolerate (Greek *delta*) | 0.02 | 308 309**In words:** roll back when the new version's success rate falls more than 310the allowed margin below the current version's. 311 312**On the worked example:** 0.90 < 0.95 − 0.02 = 0.93, so roll back. 313 314**In Python:** 315 316```python 317p_canary, p_control, delta = 0.90, 0.95, 0.02 318round(p_control - delta, 2) # → 0.93 319# roll back? 320p_canary < p_control - delta # → True 321``` 322 323Users are assigned to the canary by hashing their id into one of 100 324buckets: the same user always lands in the same bucket, so nobody flips 325between versions mid-conversation, and growing from 5% to 25% only adds 326users. 327 328```mermaid 329sequenceDiagram 330 participant D as Deploy system 331 participant R as Router 332 participant M as Metrics 333 D->>R: send 1% of users to v2 334 R->>M: success rates, v1 vs v2 335 M-->>D: v2 95%, v1 95%: healthy 336 D->>R: grow to 5% 337 R->>M: success rates, v1 vs v2 338 M-->>D: v2 91%, v1 95%: worse by 4 points 339 D->>R: roll back: 0% to v2 340 D-->>D: alert the owner with traces 341``` 342 343**Reading it:** the deploy system only ever acts on measurements. Each 344growth step waits for enough traffic to judge, and a single bad reading 345sends everyone back to the old version automatically. At worst, a few 346percent of users saw the bad version for one step. 347 348 349 350**Reading it:** each panel is one rollout. The x-axis is the rollout step 351(with the share of traffic on the canary); the lines are measured success 352rates for control and canary. The good version (left) tracks the control 353and reaches 100%. The bad version (right), truly 91% successful, gets lucky 354on the 100 tasks of the 1% step and advances, then falls below the tolerance 355band at 5% and is rolled back to 0%, marked by the red cross. Small steps 356are cheap to get wrong; that's why there are several. 357 358**In code:** `in_canary` hashes a user id into one of 100 buckets. 359`CanaryController.observe` accumulates success counts, and 360`CanaryController.step` applies the rollback rule above: advance, wait for 361more samples, or roll back. `simulate_rollout` drives a whole rollout for the 362figure. 363 364A **feature flag** is the same idea as an on/off switch in configuration: 365it turns a capability on for chosen tenants or users without redeploying, 366and off again just as fast. 367 368## Kill switches and rate limits: bound the blast radius 369 370**Everyday picture.** The red emergency-stop button on a treadmill, and the 371turnstile at a stadium that lets people through only so fast however hard 372the crowd pushes. 373 374A **kill switch** stops the agent instantly, for one **tenant** (one 375customer organisation) or for everyone. A **rate limit** caps actions per 376unit time, so even a malfunctioning agent can only send so many emails or 377change so many records per minute. The **blast radius** is how much damage 378a failure can do before someone notices; these two mechanisms bound it. 379 380**Worked example: the token bucket.** Picture a jar that holds at most 5 381tokens and gains 1 token per second. Each action takes a token; with no 382token, the action is refused. Starting full: 5 actions in a burst succeed, 383the 6th is refused. After 2 seconds the jar has 2 tokens again: 2 more 384succeed, the next is refused. 385 386$$ 387b(t) = \min\bigl(C,\; b(t_0) + r\,(t - t_0)\bigr) 388$$ 389 390**Symbols** 391 392| Symbol | Meaning here | Worked example | 393|---|---|---| 394| $b(t)$ | tokens in the bucket at time $t$ | 2 at t = 2 s | 395| $t_0$ | when the bucket was last updated | 0 s | 396| $b(t_0)$ | tokens left then | 0 after the burst | 397| $r$ | refill rate, tokens per second | 1 | 398| $C$ | capacity: the largest burst allowed | 5 | 399| $\min$ | the smaller of the two: the jar can't overflow | | 400 401**In words:** the tokens now are what was left plus what has dripped in 402since, but never more than the jar holds. 403 404**On the worked example:** min(5, 0 + 1 × (2 − 0)) = 2 tokens, so two 405more actions are allowed. 406 407**In Python:** 408 409```python 410# capacity, tokens per second 411C, r = 5, 1 412def b(t, t_0, b_t0): 413 # what was left plus the drip, capped at C 414 return min(C, b_t0 + r * (t - t_0)) 415# 2 s after the burst emptied the jar 416b(2, 0, 0) # → 2 417# a long wait refills only to capacity 418b(60, 0, 0) # → 5 419``` 420 421 422 423**Reading it:** the blue line is the number of tokens in the jar over time; 424green dots are allowed actions and red crosses refused ones. The opening 425burst drains the jar, refusals follow, and then actions trickle through at 426the refill rate: bursts are allowed, sustained floods are not. 427 428**In code:** `KillSwitch` holds the global and per-tenant off switches. 429`TokenBucket.level` is the formula for $b(t)$, and `TokenBucket.allow` 430spends a token or refuses. 431 432## A tamper-evident audit log 433 434**Everyday picture.** A ledger where every page ends with a wax seal 435pressed from the previous page's seal plus this page's contents. Change one 436number on page 12 and its seal no longer matches, and neither does any seal 437after it. 438 439An **audit log** records which agent took which action, for which user, 440with what inputs. For regulated work, it must be **tamper-evident**: an 441alteration can't go unnoticed. A **hash chain** does this: each entry 442stores the hash of the previous entry, and its own hash covers its contents 443and that previous hash. 444 445$$ 446h_i = \operatorname{SHA256}(h_{i-1} \,\Vert\, \text{entry}_i), \qquad h_0 = 00\ldots0 447$$ 448 449(Entries are numbered from 1 here; `verify()` reports Python list 450positions, which start at 0, so "entry 2" is position 1.) 451 452**Symbols** 453 454| Symbol | Meaning here | 455|---|---| 456| $h_i$ | the fingerprint (hash) stored with entry $i$ | 457| $h_{i-1}$ | the previous entry's fingerprint | 458| $\operatorname{SHA256}$ | the hash function: any input to a 64-hex-digit fingerprint | 459| $\Vert$ | "joined together with" (concatenation) | 460| $h_0$ | a fixed starting value for the first entry | 461 462**In words:** each entry's fingerprint is the fingerprint of the previous 463fingerprint joined to this entry's contents. 464 465**On the worked example:** three entries: refund A200 for 40, refund A201 466for 15, escalate A300. Change entry 2's amount from 15 to 1,500 and 467recomputing its fingerprint no longer gives the stored $h_2$, so `verify()` 468reports position 1. Delete entry 2 instead, and entry 3 moves up to position 4691; its stored previous fingerprint is $h_2$, the deleted entry's, which 470doesn't match $h_1$, so the break is again reported at position 1. 471 472**In Python:** 473 474```python 475import hashlib 476def seal(prev, entry): 477 # SHA256(h_{i-1} ‖ entry_i) 478 return hashlib.sha256((prev + entry).encode()).hexdigest() 479# h_0 480log, prev = [], "0" * 64 481for entry in ["refund A200 40", "refund A201 15", "escalate A300"]: 482 prev = seal(prev, entry) 483 # each entry is stored with its h_i 484 log.append((entry, prev)) 485def verify(log): 486 prev = "0" * 64 487 for position, (entry, h_i) in enumerate(log): 488 if seal(prev, entry) != h_i: 489 # the first broken link 490 return position 491 prev = h_i 492verify(log) is None # → True 493# entry 2's amount changed 494verify([log[0], ("refund A201 1500", log[1][1]), log[2]]) # → 1 495# entry 2 deleted 496verify([log[0], log[2]]) # → 1 497``` 498 499```mermaid 500flowchart LR 501 G["h0 = 000…"] --> E1["entry 1: refund A200, 40<br/>prev = h0<br/>h1 = SHA256(h0 ‖ entry 1)"] 502 E1 --> E2["entry 2: refund A201, 15<br/>prev = h1<br/>h2 = SHA256(h1 ‖ entry 2)"] 503 E2 --> E3["entry 3: escalate A300<br/>prev = h2<br/>h3 = SHA256(h2 ‖ entry 3)"] 504``` 505 506**Reading it:** each box carries the fingerprint of the box before it, so 507the boxes are chained. Editing any box changes its fingerprint, which breaks 508the link to the next box, so verification walks the chain and stops at the 509first broken link. In production, also write the log to storage that can't 510be modified afterwards (write-once storage) and link each entry to its full 511trace. 512 513**In code:** `AuditLog.append` stores the previous entry's hash in the new 514entry and seals it with its own; `AuditLog.verify` walks the chain and 515returns the position of the first broken link, or None. 516 517## Change management: adoption is part of the job 518 519A technically working agent still fails if people don't trust or use it. 520Involve the people whose work it touches from shadow mode onwards (their 521corrections are your best eval data), design around their actual workflow, 522show sources and confidence so they can check answers, and make handing a 523case to a person one click. 524 525## In 20 seconds 526- Earn autonomy in steps: shadow mode, then human approval of each action, 527 then autonomy for low-risk actions only; demote on any incident or drift. 528- Treat prompts as code: versioned (content hashes), reviewed, evaluated, 529 and one step from rollback. 530- Release to a small canary share first and grow it only while its metrics 531 match the current version; roll back automatically otherwise. 532- Bound the blast radius: kill switches per tenant and globally, and rate 533 limits on actions. 534- Keep a tamper-evident audit log (a hash chain) linking every action to its 535 user, authority and trace. 536 537## Self-test questions 538 539**How would you roll out an agent that takes real actions in a company's 540ERP system (its finance and operations system of record)?** 541Start read-only in shadow mode on real transactions, comparing proposals 542with what staff actually did, broken down by action type. Give the agent a 543dedicated service identity with the narrowest permissions, a sandbox or 544dry-run mode for writes, and idempotency keys so retries can't double-post. 545Move to proposing actions that a person approves in their existing 546workflow, starting with the low-value, reversible action types that scored 547best in shadow. Grant autonomy only for those, with value thresholds that 548still require approval, per-tenant kill switches, rate limits on writes, a 549canary rollout for every prompt or model change, eval gates in CI, and a 550tamper-evident audit log tied to traces. Demote automatically on incidents 551or drift, and keep the finance team involved throughout. 552 553**Why start in shadow mode instead of with human approval?** 554Shadow mode costs the humans nothing extra and produces a clean comparison 555against what they actually did, at full volume, before the agent can 556influence anything. Approval mode adds review load and can bias reviewers 557toward accepting the agent's suggestion. 558 559**What does a canary protect against that offline evals don't?** 560Real traffic: inputs, integrations and user behaviour the golden set didn't 561anticipate. Evals gate the release; the canary limits the damage from 562whatever the evals missed. 563 564**Why a hash chain rather than an ordinary log table?** 565Anyone with write access can quietly edit an ordinary row. With a hash 566chain, any edit or deletion breaks every later link, so tampering is 567detectable, and anchoring the latest hash somewhere separate (or using 568write-once storage) makes it provable. 569 570## The papers behind this lesson 571 572- Haber & Stornetta, *How to Time-Stamp a Digital Document* (Journal of 573 Cryptology, 1991), https://doi.org/10.1007/BF00196791. Introduced chaining 574 each record to the hash of the one before, the idea behind tamper-evident 575 logs (and, later, blockchains). 576 577## Further reading 578- Google SRE workbook, canarying releases: https://sre.google/workbook/canarying-releases/ 579- Martin Fowler, feature toggles: https://martinfowler.com/articles/feature-toggles.html 580- Token bucket algorithm: https://en.wikipedia.org/wiki/Token_bucket 581- NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework 582""" 583 584from __future__ import annotations 585 586import hashlib 587import json 588import time 589from collections import defaultdict, deque 590from dataclasses import dataclass, field 591from typing import Any, Callable 592 593import numpy as np 594 595from primer._show import banner, say, table, takeaway 596 597# --------------------------------------------------------------------------- 598# 1. Shadow mode and graduated autonomy 599# --------------------------------------------------------------------------- 600 601 602def shadow_agreement(agent: list[str], human: list[str]) -> dict[str, Any]: 603 """Compare what the agent would have done with what people actually did.""" 604 by_action: dict[str, list[bool]] = defaultdict(list) 605 for a, h in zip(agent, human): 606 by_action[a].append(a == h) 607 overall = sum(a == h for a, h in zip(agent, human)) / len(agent) 608 return {"overall": overall, "by_agent_action": {k: sum(v) / len(v) for k, v in by_action.items()}} 609 610 611LEVELS = ("shadow", "approve_each", "auto_low_risk") 612 613 614@dataclass 615class AutonomyController: 616 """Tracks evidence and moves an agent between autonomy levels. 617 618 Promotions need a full window of recent decisions above a bar; 619 demotions happen immediately. 620 """ 621 622 level: str = "shadow" 623 min_decisions: int = 50 624 min_agreement: float = 0.95 625 min_approval_rate: float = 0.98 626 window: deque = field(default_factory=deque) 627 incidents: list[str] = field(default_factory=list) 628 629 def _rate(self) -> float | None: 630 if len(self.window) < self.min_decisions: 631 return None # not enough evidence yet 632 return sum(self.window) / len(self.window) 633 634 def _push(self, ok: bool) -> float | None: 635 self.window.append(ok) 636 while len(self.window) > self.min_decisions: 637 self.window.popleft() 638 return self._rate() 639 640 def _move(self, level: str) -> None: 641 self.level = level 642 self.window.clear() # evidence at one level doesn't carry over to the next 643 644 def record_shadow(self, agreed: bool) -> None: 645 rate = self._push(agreed) 646 if self.level == "shadow" and rate is not None and rate >= self.min_agreement: 647 self._move("approve_each") 648 649 def record_approval(self, approved: bool) -> None: 650 rate = self._push(approved) 651 if self.level == "approve_each" and rate is not None and rate >= self.min_approval_rate: 652 self._move("auto_low_risk") 653 654 def record_incident(self, description: str) -> None: 655 self.incidents.append(description) 656 self._move(LEVELS[max(0, LEVELS.index(self.level) - 1)]) 657 658 def record_quality(self, success_rate: float, baseline: float, max_drop: float) -> None: 659 """Demote on drift: success falling more than `max_drop` below the baseline.""" 660 if self.level == "auto_low_risk" and success_rate < baseline - max_drop: 661 self._move("approve_each") 662 663 def may_act_alone(self, risk: str) -> bool: 664 return self.level == "auto_low_risk" and risk == "low" 665 666 667# --------------------------------------------------------------------------- 668# 2. Prompts as code 669# --------------------------------------------------------------------------- 670 671 672class PromptRegistry: 673 """Content-addressed prompt versions with one-step rollback.""" 674 675 def __init__(self) -> None: 676 self.history: dict[str, list[tuple[str, str]]] = defaultdict(list) # name -> [(version, text)] 677 self.active: dict[str, int] = {} 678 679 def register(self, name: str, text: str) -> str: 680 version = f"{name}@{hashlib.sha256(text.encode()).hexdigest()[:8]}" 681 versions = [v for v, _ in self.history[name]] 682 if version in versions: 683 self.active[name] = versions.index(version) 684 else: 685 self.history[name].append((version, text)) 686 self.active[name] = len(self.history[name]) - 1 687 return version 688 689 def rollback(self, name: str) -> str: 690 self.active[name] = max(0, self.active[name] - 1) 691 return self.history[name][self.active[name]][0] 692 693 def active_text(self, name: str) -> str: 694 return self.history[name][self.active[name]][1] 695 696 697# --------------------------------------------------------------------------- 698# 3. Canary releases 699# --------------------------------------------------------------------------- 700 701 702def in_canary(user_id: str, percent: float) -> bool: 703 """Deterministic assignment: hash the user into one of 100 buckets.""" 704 bucket = int(hashlib.sha256(user_id.encode()).hexdigest(), 16) % 100 705 return bucket < percent 706 707 708@dataclass 709class CanaryController: 710 steps: list[float] 711 max_drop: float = 0.02 712 min_samples: int = 100 713 index: int = 0 714 rolled_back: bool = False 715 counts: dict[str, int] = field(default_factory=lambda: {"cs": 0, "ct": 0, "ks": 0, "kt": 0}) 716 717 @property 718 def percent(self) -> float: 719 return 0 if self.rolled_back else self.steps[self.index] 720 721 def observe(self, control_successes: int, control_total: int, canary_successes: int, canary_total: int) -> None: 722 self.counts["cs"] += control_successes 723 self.counts["ct"] += control_total 724 self.counts["ks"] += canary_successes 725 self.counts["kt"] += canary_total 726 727 def step(self) -> float: 728 """Decide: advance, wait, or roll back. Returns the new canary percent.""" 729 c = self.counts 730 if self.rolled_back or c["kt"] < self.min_samples or c["ct"] == 0: 731 return self.percent 732 if c["ks"] / c["kt"] < c["cs"] / c["ct"] - self.max_drop: 733 self.rolled_back = True 734 elif self.index < len(self.steps) - 1: 735 self.index += 1 736 self.counts = {k: 0 for k in c} 737 return self.percent 738 739 740def simulate_rollout(canary_rate: float, control_rate: float = 0.95, users_per_step: int = 10_000, seed: int = 0) -> list[dict[str, Any]]: 741 """Walk a rollout step by step with simulated task outcomes.""" 742 rng = np.random.default_rng(seed) 743 ctl = CanaryController(steps=[1, 5, 25, 50, 100], max_drop=0.02, min_samples=100) 744 history = [] 745 for _ in range(50): # safety cap; a real rollout also has a wall-clock limit 746 pct = ctl.percent 747 n_canary = max(1, int(users_per_step * pct / 100)) 748 n_control = users_per_step - n_canary if pct < 100 else users_per_step 749 ks, cs = int(rng.binomial(n_canary, canary_rate)), int(rng.binomial(n_control, control_rate)) 750 ctl.observe(cs, n_control, ks, n_canary) 751 new = ctl.step() 752 history.append({"percent": pct, "canary_rate": ks / n_canary, "control_rate": cs / n_control, "next": new, "rolled_back": ctl.rolled_back}) 753 if ctl.rolled_back or (pct == 100): 754 break 755 return history 756 757 758# --------------------------------------------------------------------------- 759# 4. Kill switch and rate limiting 760# --------------------------------------------------------------------------- 761 762 763class KillSwitch: 764 def __init__(self) -> None: 765 self.global_off = False 766 self.tenants_off: set[str] = set() 767 768 def disable_all(self) -> None: 769 self.global_off = True 770 771 def disable_tenant(self, tenant: str) -> None: 772 self.tenants_off.add(tenant) 773 774 def enable_tenant(self, tenant: str) -> None: 775 self.tenants_off.discard(tenant) 776 777 def allowed(self, tenant: str) -> bool: 778 return not self.global_off and tenant not in self.tenants_off 779 780 781class TokenBucket: 782 """Allow bursts up to `capacity`, then `rate_per_second` actions per second.""" 783 784 def __init__(self, rate_per_second: float, capacity: float, clock: Callable[[], float] = time.monotonic): 785 self.rate, self.capacity, self.clock = rate_per_second, capacity, clock 786 self.tokens = float(capacity) 787 self.last = clock() 788 789 def level(self) -> float: 790 now = self.clock() 791 self.tokens = min(self.capacity, self.tokens + self.rate * (now - self.last)) 792 self.last = now 793 return self.tokens 794 795 def allow(self) -> bool: 796 if self.level() >= 1: 797 self.tokens -= 1 798 return True 799 return False 800 801 802# --------------------------------------------------------------------------- 803# 5. Tamper-evident audit log 804# --------------------------------------------------------------------------- 805 806GENESIS = "0" * 64 807 808 809def _entry_hash(entry: dict[str, Any]) -> str: 810 body = {k: v for k, v in entry.items() if k != "hash"} 811 # sort_keys makes the serialization canonical: same content, same bytes, same hash. 812 return hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest() 813 814 815class AuditLog: 816 def __init__(self, clock: Callable[[], float] = time.time): 817 self.entries: list[dict[str, Any]] = [] 818 self.clock = clock 819 820 def append(self, actor: str, action: str, on_behalf_of: str, inputs: dict[str, Any], trace_id: str | None = None) -> dict[str, Any]: 821 prev = self.entries[-1]["hash"] if self.entries else GENESIS 822 entry = {"at": self.clock(), "actor": actor, "action": action, "on_behalf_of": on_behalf_of, 823 "inputs": inputs, "trace_id": trace_id, "prev_hash": prev} 824 entry["hash"] = _entry_hash(entry) 825 self.entries.append(entry) 826 return entry 827 828 def verify(self) -> int | None: 829 """Index of the first entry whose link is broken, or None if intact.""" 830 prev = GENESIS 831 for i, e in enumerate(self.entries): 832 if e["prev_hash"] != prev or _entry_hash(e) != e["hash"]: 833 return i 834 prev = e["hash"] 835 return None 836 837 838# --------------------------------------------------------------------------- 839# Figures and demo 840# --------------------------------------------------------------------------- 841 842 843def shadow_run(n: int = 200, accuracy: float = 0.95, seed: int = 4) -> dict[str, Any]: 844 """Simulated shadow decisions; returns rolling agreement and the promotion point.""" 845 rng = np.random.default_rng(seed) 846 agreed = rng.random(n) < accuracy 847 c = AutonomyController(min_decisions=50, min_agreement=0.95) 848 rolling, promoted_at = [], None 849 for i, ok in enumerate(agreed): 850 window = agreed[max(0, i - 49) : i + 1] 851 rolling.append(window.mean()) 852 if promoted_at is None: 853 c.record_shadow(bool(ok)) 854 if c.level == "approve_each": 855 promoted_at = i + 1 856 return {"rolling": rolling, "promoted_at": promoted_at} 857 858 859def bucket_trace() -> dict[str, list]: 860 """A burst of 8 requests at t=0, then one request every 0.5 s.""" 861 now = [0.0] 862 b = TokenBucket(1.0, 5, clock=lambda: now[0]) 863 times = [0.0] * 8 + [0.5 * k for k in range(1, 13)] 864 out = {"t": [], "allowed": [], "level_t": [], "level": []} 865 for t in times: 866 now[0] = t 867 out["level_t"].append(t) 868 out["level"].append(b.level()) 869 out["t"].append(t) 870 out["allowed"].append(b.allow()) 871 out["level_t"].append(t) 872 out["level"].append(b.tokens) 873 return out 874 875 876def figures() -> dict[str, Any]: 877 import matplotlib 878 879 matplotlib.use("Agg") 880 import matplotlib.pyplot as plt 881 882 figs: dict[str, Any] = {} 883 884 s = shadow_run() 885 fig, ax = plt.subplots(figsize=(6.5, 3.4)) 886 ax.axvspan(0, 50, color="#dddddd", label="window not yet full") 887 ax.plot(range(1, len(s["rolling"]) + 1), s["rolling"], color="#4c72b0", label="agreement, last 50 decisions") 888 ax.axhline(0.95, ls="--", color="k", lw=1, label="promotion bar (95%)") 889 if s["promoted_at"]: 890 ax.plot(s["promoted_at"], s["rolling"][s["promoted_at"] - 1], "o", color="#55a868", ms=9, label="promoted to approval mode") 891 ax.set_ylim(0.6, 1.01) # low enough for the early, short-window dips (down to 2 of 3) 892 ax.set_xlabel("shadow decisions so far") 893 ax.set_ylabel("agreement with humans") 894 ax.set_title("Shadow mode: earning the first promotion") 895 ax.legend(fontsize=7, loc="lower right") 896 fig.tight_layout() 897 figs["shadow"] = fig 898 899 fig, axes = plt.subplots(1, 2, figsize=(9, 3.4), sharey=True) 900 for ax, (title, rate, seed) in zip(axes, [("good version (95%)", 0.95, 0), ("bad version (91%)", 0.91, 3)]): 901 h = simulate_rollout(rate, seed=seed) 902 xs = range(len(h)) 903 ax.plot(xs, [r["control_rate"] for r in h], "o-", color="#8c8c8c", label="control") 904 ax.plot(xs, [r["canary_rate"] for r in h], "o-", color="#4c72b0", label="canary") 905 ax.fill_between(xs, [r["control_rate"] - 0.02 for r in h], [r["control_rate"] for r in h], color="#8c8c8c", alpha=0.2, label="tolerance (2 pts)") 906 for i, r in enumerate(h): 907 if r["rolled_back"]: 908 ax.plot(i, r["canary_rate"], "x", color="#c44e52", ms=14, mew=3, label="rolled back to 0%") 909 ax.set_xticks(list(xs), [f"{r['percent']}%" for r in h]) 910 ax.set_xlabel("share of traffic on the canary") 911 ax.set_title(title) 912 axes[0].set_ylabel("task success rate") 913 axes[1].legend(fontsize=7, loc="lower left") 914 fig.tight_layout() 915 figs["canary"] = fig 916 917 bt = bucket_trace() 918 fig, ax = plt.subplots(figsize=(6.5, 3.2)) 919 ax.step(bt["level_t"], bt["level"], where="post", color="#4c72b0", label="tokens in the bucket") 920 ok = [(t, 5.3) for t, a in zip(bt["t"], bt["allowed"]) if a] 921 no = [(t, 5.6) for t, a in zip(bt["t"], bt["allowed"]) if not a] 922 ax.scatter([t for t, _ in ok], [y for _, y in ok], color="#55a868", marker="o", label="allowed") 923 ax.scatter([t for t, _ in no], [y for _, y in no], color="#c44e52", marker="x", label="refused") 924 ax.set_xlabel("seconds") 925 ax.set_ylabel("tokens") 926 ax.set_title("Token bucket: capacity 5, refill 1 per second") 927 ax.legend(fontsize=7, loc="center right") 928 fig.tight_layout() 929 figs["token_bucket"] = fig 930 return figs 931 932 933def demo() -> None: 934 banner("1. Shadow mode: compare, don't act") 935 agent = ["refund", "refund", "reply", "reply", "reply", "refund", "reply", "refund", "reply", "refund"] 936 human = ["refund", "escalate", "reply", "reply", "reply", "refund", "reply", "escalate", "reply", "refund"] 937 print(shadow_agreement(agent, human)) 938 print() 939 say("Replies agree every time; refunds only 3 in 5. Let the agent draft replies first.") 940 941 banner("2. Graduated autonomy") 942 c = AutonomyController() 943 s = shadow_run() 944 print(f"simulated shadow run promoted after {s['promoted_at']} decisions") 945 c.level = "auto_low_risk" 946 print(f"autonomous: may act alone on low risk? {c.may_act_alone('low')}; on high risk? {c.may_act_alone('high')}") 947 c.record_incident("refunded the wrong customer") 948 print(f"after an incident: level = {c.level}") 949 print() 950 951 banner("3. Prompts as code") 952 reg = PromptRegistry() 953 v1 = reg.register("support", "Be concise.") 954 v2 = reg.register("support", "Always be maximally helpful.") 955 print(f"registered {v1}, then {v2}; rolled back to {reg.rollback('support')}") 956 print() 957 958 banner("4. Canary rollouts") 959 for label, rate, seed in [("good", 0.95, 0), ("bad", 0.91, 3)]: 960 h = simulate_rollout(rate, seed=seed) 961 table(["canary %", "control", "canary", "next %"], [(r["percent"], r["control_rate"], r["canary_rate"], r["next"]) for r in h], floatfmt=".3f") 962 print(f"{label} version: {'rolled back' if h[-1]['rolled_back'] else 'fully rolled out'}") 963 print() 964 965 banner("5. Kill switch and rate limit") 966 ks = KillSwitch() 967 ks.disable_tenant("acme") 968 print(f"acme allowed? {ks.allowed('acme')}; globex allowed? {ks.allowed('globex')}") 969 bt = bucket_trace() 970 print("burst of 8 at t=0:", ["ok" if a else "refused" for a in bt["allowed"][:8]]) 971 print() 972 973 banner("6. Tamper-evident audit log") 974 log = AuditLog(clock=lambda: 0.0) 975 log.append("refund-agent", "refund", "dana", {"order": "A200", "amount": 40}) 976 log.append("refund-agent", "refund", "lee", {"order": "A201", "amount": 15}) 977 log.append("refund-agent", "escalate", "sam", {"order": "A300"}) 978 print("intact log verifies:", log.verify() is None) 979 log.entries[1]["inputs"]["amount"] = 1500 980 print("after editing entry 1, first broken link at:", log.verify()) 981 print() 982 takeaway("Earn autonomy with evidence, release to a canary first, bound the blast radius, and make every action auditable.") 983 984 985if __name__ == "__main__": 986 demo()
603def shadow_agreement(agent: list[str], human: list[str]) -> dict[str, Any]: 604 """Compare what the agent would have done with what people actually did.""" 605 by_action: dict[str, list[bool]] = defaultdict(list) 606 for a, h in zip(agent, human): 607 by_action[a].append(a == h) 608 overall = sum(a == h for a, h in zip(agent, human)) / len(agent) 609 return {"overall": overall, "by_agent_action": {k: sum(v) / len(v) for k, v in by_action.items()}}
Compare what the agent would have done with what people actually did.
615@dataclass 616class AutonomyController: 617 """Tracks evidence and moves an agent between autonomy levels. 618 619 Promotions need a full window of recent decisions above a bar; 620 demotions happen immediately. 621 """ 622 623 level: str = "shadow" 624 min_decisions: int = 50 625 min_agreement: float = 0.95 626 min_approval_rate: float = 0.98 627 window: deque = field(default_factory=deque) 628 incidents: list[str] = field(default_factory=list) 629 630 def _rate(self) -> float | None: 631 if len(self.window) < self.min_decisions: 632 return None # not enough evidence yet 633 return sum(self.window) / len(self.window) 634 635 def _push(self, ok: bool) -> float | None: 636 self.window.append(ok) 637 while len(self.window) > self.min_decisions: 638 self.window.popleft() 639 return self._rate() 640 641 def _move(self, level: str) -> None: 642 self.level = level 643 self.window.clear() # evidence at one level doesn't carry over to the next 644 645 def record_shadow(self, agreed: bool) -> None: 646 rate = self._push(agreed) 647 if self.level == "shadow" and rate is not None and rate >= self.min_agreement: 648 self._move("approve_each") 649 650 def record_approval(self, approved: bool) -> None: 651 rate = self._push(approved) 652 if self.level == "approve_each" and rate is not None and rate >= self.min_approval_rate: 653 self._move("auto_low_risk") 654 655 def record_incident(self, description: str) -> None: 656 self.incidents.append(description) 657 self._move(LEVELS[max(0, LEVELS.index(self.level) - 1)]) 658 659 def record_quality(self, success_rate: float, baseline: float, max_drop: float) -> None: 660 """Demote on drift: success falling more than `max_drop` below the baseline.""" 661 if self.level == "auto_low_risk" and success_rate < baseline - max_drop: 662 self._move("approve_each") 663 664 def may_act_alone(self, risk: str) -> bool: 665 return self.level == "auto_low_risk" and risk == "low"
Tracks evidence and moves an agent between autonomy levels.
Promotions need a full window of recent decisions above a bar; demotions happen immediately.
659 def record_quality(self, success_rate: float, baseline: float, max_drop: float) -> None: 660 """Demote on drift: success falling more than `max_drop` below the baseline.""" 661 if self.level == "auto_low_risk" and success_rate < baseline - max_drop: 662 self._move("approve_each")
Demote on drift: success falling more than max_drop below the baseline.
673class PromptRegistry: 674 """Content-addressed prompt versions with one-step rollback.""" 675 676 def __init__(self) -> None: 677 self.history: dict[str, list[tuple[str, str]]] = defaultdict(list) # name -> [(version, text)] 678 self.active: dict[str, int] = {} 679 680 def register(self, name: str, text: str) -> str: 681 version = f"{name}@{hashlib.sha256(text.encode()).hexdigest()[:8]}" 682 versions = [v for v, _ in self.history[name]] 683 if version in versions: 684 self.active[name] = versions.index(version) 685 else: 686 self.history[name].append((version, text)) 687 self.active[name] = len(self.history[name]) - 1 688 return version 689 690 def rollback(self, name: str) -> str: 691 self.active[name] = max(0, self.active[name] - 1) 692 return self.history[name][self.active[name]][0] 693 694 def active_text(self, name: str) -> str: 695 return self.history[name][self.active[name]][1]
Content-addressed prompt versions with one-step rollback.
680 def register(self, name: str, text: str) -> str: 681 version = f"{name}@{hashlib.sha256(text.encode()).hexdigest()[:8]}" 682 versions = [v for v, _ in self.history[name]] 683 if version in versions: 684 self.active[name] = versions.index(version) 685 else: 686 self.history[name].append((version, text)) 687 self.active[name] = len(self.history[name]) - 1 688 return version
703def in_canary(user_id: str, percent: float) -> bool: 704 """Deterministic assignment: hash the user into one of 100 buckets.""" 705 bucket = int(hashlib.sha256(user_id.encode()).hexdigest(), 16) % 100 706 return bucket < percent
Deterministic assignment: hash the user into one of 100 buckets.
709@dataclass 710class CanaryController: 711 steps: list[float] 712 max_drop: float = 0.02 713 min_samples: int = 100 714 index: int = 0 715 rolled_back: bool = False 716 counts: dict[str, int] = field(default_factory=lambda: {"cs": 0, "ct": 0, "ks": 0, "kt": 0}) 717 718 @property 719 def percent(self) -> float: 720 return 0 if self.rolled_back else self.steps[self.index] 721 722 def observe(self, control_successes: int, control_total: int, canary_successes: int, canary_total: int) -> None: 723 self.counts["cs"] += control_successes 724 self.counts["ct"] += control_total 725 self.counts["ks"] += canary_successes 726 self.counts["kt"] += canary_total 727 728 def step(self) -> float: 729 """Decide: advance, wait, or roll back. Returns the new canary percent.""" 730 c = self.counts 731 if self.rolled_back or c["kt"] < self.min_samples or c["ct"] == 0: 732 return self.percent 733 if c["ks"] / c["kt"] < c["cs"] / c["ct"] - self.max_drop: 734 self.rolled_back = True 735 elif self.index < len(self.steps) - 1: 736 self.index += 1 737 self.counts = {k: 0 for k in c} 738 return self.percent
728 def step(self) -> float: 729 """Decide: advance, wait, or roll back. Returns the new canary percent.""" 730 c = self.counts 731 if self.rolled_back or c["kt"] < self.min_samples or c["ct"] == 0: 732 return self.percent 733 if c["ks"] / c["kt"] < c["cs"] / c["ct"] - self.max_drop: 734 self.rolled_back = True 735 elif self.index < len(self.steps) - 1: 736 self.index += 1 737 self.counts = {k: 0 for k in c} 738 return self.percent
Decide: advance, wait, or roll back. Returns the new canary percent.
741def simulate_rollout(canary_rate: float, control_rate: float = 0.95, users_per_step: int = 10_000, seed: int = 0) -> list[dict[str, Any]]: 742 """Walk a rollout step by step with simulated task outcomes.""" 743 rng = np.random.default_rng(seed) 744 ctl = CanaryController(steps=[1, 5, 25, 50, 100], max_drop=0.02, min_samples=100) 745 history = [] 746 for _ in range(50): # safety cap; a real rollout also has a wall-clock limit 747 pct = ctl.percent 748 n_canary = max(1, int(users_per_step * pct / 100)) 749 n_control = users_per_step - n_canary if pct < 100 else users_per_step 750 ks, cs = int(rng.binomial(n_canary, canary_rate)), int(rng.binomial(n_control, control_rate)) 751 ctl.observe(cs, n_control, ks, n_canary) 752 new = ctl.step() 753 history.append({"percent": pct, "canary_rate": ks / n_canary, "control_rate": cs / n_control, "next": new, "rolled_back": ctl.rolled_back}) 754 if ctl.rolled_back or (pct == 100): 755 break 756 return history
Walk a rollout step by step with simulated task outcomes.
764class KillSwitch: 765 def __init__(self) -> None: 766 self.global_off = False 767 self.tenants_off: set[str] = set() 768 769 def disable_all(self) -> None: 770 self.global_off = True 771 772 def disable_tenant(self, tenant: str) -> None: 773 self.tenants_off.add(tenant) 774 775 def enable_tenant(self, tenant: str) -> None: 776 self.tenants_off.discard(tenant) 777 778 def allowed(self, tenant: str) -> bool: 779 return not self.global_off and tenant not in self.tenants_off
782class TokenBucket: 783 """Allow bursts up to `capacity`, then `rate_per_second` actions per second.""" 784 785 def __init__(self, rate_per_second: float, capacity: float, clock: Callable[[], float] = time.monotonic): 786 self.rate, self.capacity, self.clock = rate_per_second, capacity, clock 787 self.tokens = float(capacity) 788 self.last = clock() 789 790 def level(self) -> float: 791 now = self.clock() 792 self.tokens = min(self.capacity, self.tokens + self.rate * (now - self.last)) 793 self.last = now 794 return self.tokens 795 796 def allow(self) -> bool: 797 if self.level() >= 1: 798 self.tokens -= 1 799 return True 800 return False
Allow bursts up to capacity, then rate_per_second actions per second.
816class AuditLog: 817 def __init__(self, clock: Callable[[], float] = time.time): 818 self.entries: list[dict[str, Any]] = [] 819 self.clock = clock 820 821 def append(self, actor: str, action: str, on_behalf_of: str, inputs: dict[str, Any], trace_id: str | None = None) -> dict[str, Any]: 822 prev = self.entries[-1]["hash"] if self.entries else GENESIS 823 entry = {"at": self.clock(), "actor": actor, "action": action, "on_behalf_of": on_behalf_of, 824 "inputs": inputs, "trace_id": trace_id, "prev_hash": prev} 825 entry["hash"] = _entry_hash(entry) 826 self.entries.append(entry) 827 return entry 828 829 def verify(self) -> int | None: 830 """Index of the first entry whose link is broken, or None if intact.""" 831 prev = GENESIS 832 for i, e in enumerate(self.entries): 833 if e["prev_hash"] != prev or _entry_hash(e) != e["hash"]: 834 return i 835 prev = e["hash"] 836 return None
821 def append(self, actor: str, action: str, on_behalf_of: str, inputs: dict[str, Any], trace_id: str | None = None) -> dict[str, Any]: 822 prev = self.entries[-1]["hash"] if self.entries else GENESIS 823 entry = {"at": self.clock(), "actor": actor, "action": action, "on_behalf_of": on_behalf_of, 824 "inputs": inputs, "trace_id": trace_id, "prev_hash": prev} 825 entry["hash"] = _entry_hash(entry) 826 self.entries.append(entry) 827 return entry
829 def verify(self) -> int | None: 830 """Index of the first entry whose link is broken, or None if intact.""" 831 prev = GENESIS 832 for i, e in enumerate(self.entries): 833 if e["prev_hash"] != prev or _entry_hash(e) != e["hash"]: 834 return i 835 prev = e["hash"] 836 return None
Index of the first entry whose link is broken, or None if intact.
844def shadow_run(n: int = 200, accuracy: float = 0.95, seed: int = 4) -> dict[str, Any]: 845 """Simulated shadow decisions; returns rolling agreement and the promotion point.""" 846 rng = np.random.default_rng(seed) 847 agreed = rng.random(n) < accuracy 848 c = AutonomyController(min_decisions=50, min_agreement=0.95) 849 rolling, promoted_at = [], None 850 for i, ok in enumerate(agreed): 851 window = agreed[max(0, i - 49) : i + 1] 852 rolling.append(window.mean()) 853 if promoted_at is None: 854 c.record_shadow(bool(ok)) 855 if c.level == "approve_each": 856 promoted_at = i + 1 857 return {"rolling": rolling, "promoted_at": promoted_at}
Simulated shadow decisions; returns rolling agreement and the promotion point.
860def bucket_trace() -> dict[str, list]: 861 """A burst of 8 requests at t=0, then one request every 0.5 s.""" 862 now = [0.0] 863 b = TokenBucket(1.0, 5, clock=lambda: now[0]) 864 times = [0.0] * 8 + [0.5 * k for k in range(1, 13)] 865 out = {"t": [], "allowed": [], "level_t": [], "level": []} 866 for t in times: 867 now[0] = t 868 out["level_t"].append(t) 869 out["level"].append(b.level()) 870 out["t"].append(t) 871 out["allowed"].append(b.allow()) 872 out["level_t"].append(t) 873 out["level"].append(b.tokens) 874 return out
A burst of 8 requests at t=0, then one request every 0.5 s.
877def figures() -> dict[str, Any]: 878 import matplotlib 879 880 matplotlib.use("Agg") 881 import matplotlib.pyplot as plt 882 883 figs: dict[str, Any] = {} 884 885 s = shadow_run() 886 fig, ax = plt.subplots(figsize=(6.5, 3.4)) 887 ax.axvspan(0, 50, color="#dddddd", label="window not yet full") 888 ax.plot(range(1, len(s["rolling"]) + 1), s["rolling"], color="#4c72b0", label="agreement, last 50 decisions") 889 ax.axhline(0.95, ls="--", color="k", lw=1, label="promotion bar (95%)") 890 if s["promoted_at"]: 891 ax.plot(s["promoted_at"], s["rolling"][s["promoted_at"] - 1], "o", color="#55a868", ms=9, label="promoted to approval mode") 892 ax.set_ylim(0.6, 1.01) # low enough for the early, short-window dips (down to 2 of 3) 893 ax.set_xlabel("shadow decisions so far") 894 ax.set_ylabel("agreement with humans") 895 ax.set_title("Shadow mode: earning the first promotion") 896 ax.legend(fontsize=7, loc="lower right") 897 fig.tight_layout() 898 figs["shadow"] = fig 899 900 fig, axes = plt.subplots(1, 2, figsize=(9, 3.4), sharey=True) 901 for ax, (title, rate, seed) in zip(axes, [("good version (95%)", 0.95, 0), ("bad version (91%)", 0.91, 3)]): 902 h = simulate_rollout(rate, seed=seed) 903 xs = range(len(h)) 904 ax.plot(xs, [r["control_rate"] for r in h], "o-", color="#8c8c8c", label="control") 905 ax.plot(xs, [r["canary_rate"] for r in h], "o-", color="#4c72b0", label="canary") 906 ax.fill_between(xs, [r["control_rate"] - 0.02 for r in h], [r["control_rate"] for r in h], color="#8c8c8c", alpha=0.2, label="tolerance (2 pts)") 907 for i, r in enumerate(h): 908 if r["rolled_back"]: 909 ax.plot(i, r["canary_rate"], "x", color="#c44e52", ms=14, mew=3, label="rolled back to 0%") 910 ax.set_xticks(list(xs), [f"{r['percent']}%" for r in h]) 911 ax.set_xlabel("share of traffic on the canary") 912 ax.set_title(title) 913 axes[0].set_ylabel("task success rate") 914 axes[1].legend(fontsize=7, loc="lower left") 915 fig.tight_layout() 916 figs["canary"] = fig 917 918 bt = bucket_trace() 919 fig, ax = plt.subplots(figsize=(6.5, 3.2)) 920 ax.step(bt["level_t"], bt["level"], where="post", color="#4c72b0", label="tokens in the bucket") 921 ok = [(t, 5.3) for t, a in zip(bt["t"], bt["allowed"]) if a] 922 no = [(t, 5.6) for t, a in zip(bt["t"], bt["allowed"]) if not a] 923 ax.scatter([t for t, _ in ok], [y for _, y in ok], color="#55a868", marker="o", label="allowed") 924 ax.scatter([t for t, _ in no], [y for _, y in no], color="#c44e52", marker="x", label="refused") 925 ax.set_xlabel("seconds") 926 ax.set_ylabel("tokens") 927 ax.set_title("Token bucket: capacity 5, refill 1 per second") 928 ax.legend(fontsize=7, loc="center right") 929 fig.tight_layout() 930 figs["token_bucket"] = fig 931 return figs
934def demo() -> None: 935 banner("1. Shadow mode: compare, don't act") 936 agent = ["refund", "refund", "reply", "reply", "reply", "refund", "reply", "refund", "reply", "refund"] 937 human = ["refund", "escalate", "reply", "reply", "reply", "refund", "reply", "escalate", "reply", "refund"] 938 print(shadow_agreement(agent, human)) 939 print() 940 say("Replies agree every time; refunds only 3 in 5. Let the agent draft replies first.") 941 942 banner("2. Graduated autonomy") 943 c = AutonomyController() 944 s = shadow_run() 945 print(f"simulated shadow run promoted after {s['promoted_at']} decisions") 946 c.level = "auto_low_risk" 947 print(f"autonomous: may act alone on low risk? {c.may_act_alone('low')}; on high risk? {c.may_act_alone('high')}") 948 c.record_incident("refunded the wrong customer") 949 print(f"after an incident: level = {c.level}") 950 print() 951 952 banner("3. Prompts as code") 953 reg = PromptRegistry() 954 v1 = reg.register("support", "Be concise.") 955 v2 = reg.register("support", "Always be maximally helpful.") 956 print(f"registered {v1}, then {v2}; rolled back to {reg.rollback('support')}") 957 print() 958 959 banner("4. Canary rollouts") 960 for label, rate, seed in [("good", 0.95, 0), ("bad", 0.91, 3)]: 961 h = simulate_rollout(rate, seed=seed) 962 table(["canary %", "control", "canary", "next %"], [(r["percent"], r["control_rate"], r["canary_rate"], r["next"]) for r in h], floatfmt=".3f") 963 print(f"{label} version: {'rolled back' if h[-1]['rolled_back'] else 'fully rolled out'}") 964 print() 965 966 banner("5. Kill switch and rate limit") 967 ks = KillSwitch() 968 ks.disable_tenant("acme") 969 print(f"acme allowed? {ks.allowed('acme')}; globex allowed? {ks.allowed('globex')}") 970 bt = bucket_trace() 971 print("burst of 8 at t=0:", ["ok" if a else "refused" for a in bt["allowed"][:8]]) 972 print() 973 974 banner("6. Tamper-evident audit log") 975 log = AuditLog(clock=lambda: 0.0) 976 log.append("refund-agent", "refund", "dana", {"order": "A200", "amount": 40}) 977 log.append("refund-agent", "refund", "lee", {"order": "A201", "amount": 15}) 978 log.append("refund-agent", "escalate", "sam", {"order": "A300"}) 979 print("intact log verifies:", log.verify() is None) 980 log.entries[1]["inputs"]["amount"] = 1500 981 print("after editing entry 1, first broken link at:", log.verify()) 982 print() 983 takeaway("Earn autonomy with evidence, release to a canary first, bound the blast radius, and make every action auditable.")