Backups and recovery
The insurance
What makes a mistake survivable.
Things break. A disk fails, an update goes wrong, a worker deletes the wrong file. The question is never whether, only how much it costs when it happens.
So nearly everything in my system is built to be undone, and I test the undoing instead of assuming it works.
64,977
database rows read back in a restore test
The safety nets
A locked copy elsewhere
The assistant’s data is backed up, encrypted, on a different machine. Nobody can read it without the key.
A restore I actually tried
I brought the backup back on that second machine and checked it: 2,606 files and 64,977 database rows, all readable, without disturbing the live system.
A copy before anything risky
Before a change that could damage something, a dated copy is made first. A safety net is faster than a discussion.
A record of every change
Code, rules and skills are all tracked, so any change can be reversed and I can see who made it.
Retiring things properly
When I remove a tool, I write down how to bring it back and keep what is needed to do it. Then I really remove it.
Alarms I have heard ring
A watchdog warns me before a login expires. I did not trust it until I had made it go off on purpose.
Leaving work finishable
At the end of each task there is a note saying what is left. If I stop, someone else, or another AI, can continue.
What it taught me
- 1
A backup you have never restored is a hope.
- 2
Copying a database while it is in use can produce a broken copy that looks perfect. I caught that in a test, not in an emergency.
- 3
An alarm nobody has ever heard is not an alarm.