Most self-hosted Supabase teams have backups. Very few have a disaster recovery plan. The difference sounds academic until 2 AM on a Saturday when your VPS provider emails you that your server's disk has failed — and you discover that "we have backups" doesn't answer any of the questions that actually matter: How much data did we just lose? How long until we're back up? Who does what, in what order?
A disaster recovery (DR) plan answers those questions before the disaster. It's not a bigger backup script — it's a document that defines what you can afford to lose, how fast you must recover, and the exact steps to get there. If you're running automated backups already, you've done maybe 40% of the work. This guide covers the other 60%: defining RPO and RTO, mapping your realistic failure scenarios, choosing a DR tier that matches your budget, and writing a recovery runbook someone other than you can execute.
Backups Are Not a DR Plan
A backup is an artifact. A DR plan is a capability. The gap between them is where self-hosters get hurt, and community threads are full of the same painful patterns:
- The backup exists but was never restored. A
pg_dumpthat has never been through a test restore is a hypothesis, not a backup. Corrupt dumps, missing roles, and version mismatches all surface at restore time — the worst possible moment to discover them. - The database was backed up but Storage wasn't. Supabase Storage files live outside Postgres. Restore the database alone and every avatar, upload, and document your users trusted you with is gone. This is common enough that we wrote a whole post on why storage backup is the forgotten piece.
- The data survived but the configuration didn't. Your
docker-compose.yml,.envsecrets, JWT keys, and gateway config are part of your system state. Losing yourJWT_SECRETinvalidates every API key and user session — a self-inflicted disaster on top of the original one. And with upstream changes like the Kong-to-Envoy gateway migration, a rebuilt-from-memory stack may not even match what you were running. - Everything was recoverable but nobody knew how. The one person who set up the server is on a plane, and the restore procedure exists only in their head.
A DR plan closes all four gaps. Start with the two numbers everything else depends on.
Define Your RPO and RTO
These two metrics turn "we should be fine" into an engineering requirement:
- RPO (Recovery Point Objective): the maximum data loss you can tolerate, measured in time. An RPO of 24 hours means losing a full day of writes is acceptable. An RPO of 5 minutes means it isn't.
- RTO (Recovery Time Objective): the maximum downtime you can tolerate. How long from "the server is gone" to "users can log in again"?
Be honest, and let the business — not the tooling — set the numbers. A side project with 50 users can live with RPO 24h / RTO 24h. A SaaS processing customer payments probably can't survive losing more than a few minutes of orders. Ask: if we lost the last N hours of data, what would it actually cost us? The answer sets your RPO. Then ask the same question about downtime for your RTO.
Your answers map directly to architecture and cost:
| RPO / RTO target | What it requires | Rough added cost |
|---|---|---|
| 24h / 24h | Nightly pg_dump + storage sync to S3 | ~$1–5/mo (object storage) |
| 1h / 4h | Scheduled base backups + hourly increments, tested restore runbook | ~$5–15/mo |
| 5min / 1h | WAL archiving with point-in-time recovery, warm standby scripts | ~$20–50/mo + setup effort |
| ~0 / minutes | Streaming replica in a second region, automated failover | 2× infrastructure + real engineering time |
Notice the curve: each tier costs meaningfully more than the last. Most self-hosted projects belong in the second or third row. Jumping to the last row "to be safe" is how teams end up maintaining multi-region complexity they never needed — if you genuinely need it, our geographic redundancy guide covers that architecture honestly, including when it's overkill.
Map Your Failure Scenarios
Generic plans fail because they plan for "disaster" in the abstract. Instead, write down the five or six failures that are actually plausible for your setup, and what each one demands:
1. Accidental data deletion. Someone runs DELETE without a WHERE, or a migration drops the wrong column. This is the most common disaster by far — and full-server redundancy does nothing for it, because the replica happily replicates the mistake. You need point-in-time recovery or frequent backups, and ideally the ability to do a partial restore of a single table rather than rolling the whole database back.
2. Disk failure or corruption. The server is alive but the data is gone. Recovery: provision fresh volume, restore latest backup, replay WAL if you have it. Your RPO is determined entirely by backup frequency here.
3. Total server loss. Provider hardware failure, account suspension, accidental termination, or the billing card silently expiring. Recovery: new server, redeploy the stack, restore database and storage and configuration, repoint DNS. This scenario is why config and secrets must be in your backup set — test whether you could rebuild from nothing but your backup storage bucket and a fresh VPS.
4. Provider or region outage. Rare, but it happens, and it's the one scenario where your backups being in the same provider becomes a problem. The fix is cheap: keep backups in a different provider's object storage (server on Hetzner, backups on Cloudflare R2 or Backblaze B2). Cross-provider separation costs a few dollars a month and removes an entire failure class.
5. Security compromise. Ransomware or an attacker with server access. The critical property here is that backups must be unreachable from the compromised server — a cron job that can write to your backup bucket with delete permissions means an attacker can too. Use object-lock/immutability or write-only credentials, and fold this scenario into your incident response runbook, because recovery here involves forensics and credential rotation, not just a restore.
For each scenario, your plan should state: detection (how do we find out?), decision (who declares a disaster?), procedure (numbered steps), and expected RTO.
Write the Runbook — Then Test It
The runbook is the part of the plan you execute under stress, so it must assume stress: exact commands, real hostnames, no "configure as appropriate." A minimal self-hosted Supabase runbook covers:
- Access: where credentials live (password manager, not the dead server), which S3 bucket holds backups, SSH keys for the standby provider.
- Provision: the server spec and provider, or better, an infrastructure-as-code definition that recreates it.
- Restore order: stack config first, then Postgres, then Storage objects, then verify Auth (can a test user log in? do JWTs validate?), then DNS cutover.
- Verification checklist: row counts on critical tables, storage object spot-checks, a real end-to-end signup.
Then — and this is the step almost everyone skips — run it. Schedule a restore drill quarterly: spin up a throwaway VPS, execute the runbook exactly as written, time it, and fix everything that didn't work. The first drill is always humbling; that's the point. Our guide on testing backup and restore procedures walks through drill formats from a 15-minute smoke test to a full failover exercise. An untested RTO is a guess. A tested one is a number you can put in a customer contract.
Where Supascale Fits
Most of a DR plan is decisions and documentation — no tool writes it for you. But the mechanical layer underneath it is exactly what Supascale automates: scheduled database and storage backups shipped to any S3-compatible bucket (including a different provider than your server, which you now know matters), retention policies, and one-click restores that make quarterly drills a ten-minute task instead of an afternoon of remembering flags. Because restores are cheap and repeatable, the biggest DR failure mode — never testing — stops being the path of least resistance.
The one-time license model also fits DR economics: your recovery capability isn't a subscription that lapses when a card expires. It's worth noting what Supascale doesn't do: it won't set your RPO for you, run your drills, or replace a hot standby if you genuinely need near-zero RTO. It handles the backup/restore machinery so your plan's steps 1–4 are push-button; the plan itself is still yours to write.
Key Takeaways
- Backups are artifacts; DR is a capability. The gap is RPO/RTO definitions, scenario mapping, a runbook, and drills.
- Let the business set RPO and RTO, then buy exactly that tier — most self-hosted projects need PITR or hourly backups, not multi-region failover.
- Back up all three layers: database, Storage files, and configuration/secrets. Missing any one turns a recovery into a rebuild.
- Keep at least one backup copy outside your server's provider, unreachable with the server's own credentials.
- A restore you haven't drilled is a hypothesis. Test quarterly and time it.
