11 / 13
Restore & disaster recovery
Recovery kits, standalone restores without the metadata database, and the three key incidents
Routine restore: the recovery kit
AGE_IDENTITY_FILE=/path/identity.txt \
PGPASSWORD='<target db password>' \
sh restore-job7.sh 'postgresql://user@host:5432/restored?sslmode=require' backup-job7.dump.ageOrder of operations: verify the ciphertext SHA-256 (mismatch exits 1 before any
decryption) → refuse a non-empty target database → age-decrypt to a temporary
plaintext → pg_restore → when the manifest carries a table count, assert
EQUALITY and print it. A failed restore may leave partial writes: clean the
target before retrying. Create the target database first (createdb). The
manifest's backupId is the instance-local id (job-N), not the bucket object
key (objects are keyed by the backup UUID).
Supabase backups
Their kit restores with a supabase profile: it needs a server that has
Supabase's extensions, a new empty database and a superuser connection, and
it checks all three before writing anything. SUPABACKUP_PROFILE=generic
or supabase overrides a kit's default. Walkthrough:
Back up a Supabase database.
Standalone restore: without this instance
If the instance is gone entirely, recovery needs exactly:
- the
*.dump.ageand its*.manifest.jsonfrom the bucket (or staging); - your offline age identity;
- the recovery kit script, or manual age-decrypt +
pg_restoreguided by the manifest hash and metadata (the full commands live in the repository playbook, scenario 3).
The restore host needs its own tools: the kit checks for age, pg_restore, psql and sha256sum (or shasum) and exits 3 when missing — the service image embeds the age Go library, not a CLI for your restore box.
The manifest is self-describing: backup id (job-N, the instance-local id),
server version, table count, SHA-256, recipient fingerprint; the bucket
object key's backup UUID lives in the instance metadata database. The SQLite metadata database is NOT a
prerequisite for recovery.
Key incidents (four cases)
Master secret file missing, age identity intact
Ciphertext stays decryptable. The instance generates a fresh master secret and keeps running — stored connection strings and destination credentials become unreadable and must be re-registered. Try to recover the original key from backups before rebuilding anything.
Master secret corrupted / bad permissions / replaced
Corruption and permission problems refuse startup (fail closed): repair the file first. A key that was merely replaced (valid format, different value) lets the instance start, but every old credential fails to decrypt — equivalent to losing it; rebuilding then requires re-entering database connection strings and destination credentials and re-binding them.
Age identity lost, master secret intact
Existing ciphertext is unrecoverable. The correct path for future
protection: preserve the old data directory and recipient record as
audit evidence (do not reuse them), then on a NEW instance (empty
recipient config) run age init, store and age verify the new identity
offline, re-register databases and destinations, restore schedules, and
prove the chain with a fresh backup + restore. age init is not a
rotation interface — an instance with a configured recipient refuses it.
Writing the OLD recipient into the new instance (the DR playbook's
scenario-1 recipe, valid only while the old identity still works) would
keep producing backups nobody can decrypt.
Both lost
Existing backups are unrecoverable. Immediately: pause the old instance,
preserve the data directory and bucket objects (for audit), then complete
age init + destination credentials + database re-registration on a new
instance so the protection chain restarts.
Mapping to the repository playbook
The command-level playbook lives in
docs/disaster-recovery.md
(developer documentation; until the repo has a remote, open the local path
docs/disaster-recovery.md). Scenario index:
- Scenario 1 (SQLite corruption / data-volume loss) → start a fresh instance with an empty data dir, bootstrap, write the original recipient back (only while the old age identity still works), re-register databases/destinations and schedules; historical data recovery follows scenario 3.
- Scenario 2 (master secret unavailable: missing / corrupted / replaced) → the first two key-incident steps above.
- Scenario 3 (whole instance gone, only bucket ciphertext + offline age
identity left) → "Standalone restore"; object keys are
<prefix>backups/<backup-uuid>.dump.age. - Scenario 4 (failed upgrade / interrupted migration) → pre-migrate snapshot rollback, isolate the old WAL/SHM, boot the previous binary.
Last updated