Checks and guardrails
Quality control
The part I care about most: making sure the work is real and the mistakes stay small.
A guardrail is a limit that holds even when the model, the instructions or I get something wrong. The difference between a rule and a guardrail is simple: a rule asks, a guardrail prevents.
I measured this. A written rule about how my assistant should speak held five times out of six. That is fine for style. It is not fine for anything that reaches a real person, so the important rules became mechanisms.
5 in 6
times a written rule held, which is why it became a mechanism
The checks in place
The pre-flight check
34 known failuresBefore a risky change, I go through a list of 34 ways things have gone wrong before, name the ones that apply, and record proof for each.
The last look before sending
Every outgoing message passes a final filter. Internal notes, error text and the model’s own thinking are stopped there.
Silence over a bad answer
If the good model is unavailable, the assistant says nothing. A weak reply to a real person costs more than a late one.
Nobody marks their own homework
Work counts as finished when someone else has checked it and there is evidence. A worker’s own “done” is not accepted.
Yes or no on my phone
Sensitive actions wait for my approval. A spoken command is read back and has to be confirmed before anything starts.
The hard stops
Passwords, payments, my personal documents, and anything a backup could not undo. No worker crosses these on its own.
Tests that can fail
A check that always passes proves nothing. Each safety test is first shown a known bad example, to prove it would catch one.
What it taught me
- 1
Put the control where the worker cannot walk around it.
- 2
Never test on a real person who does not know they are part of a test.
- 3
Measure the control itself. I built a supervision layer, ran it for 17 days, found only two real jobs had completed, and switched it off to fix it.