triplicate

Blog

Test your restores: an untested backup is a hope

A checklist with ticks and a download arrow, for a test restore

There is an old saying among people who run servers: nobody wants backups, everybody wants restores. A backup job that reports success every night tells you that something was written somewhere. It does not tell you that you can get your data back, in a usable form, in the time you need.

The only way to know is to try.

Why backups fail without anyone noticing

Some of the ways a backup can look healthy while being useless:

  • The wrong folders. A new project went into a folder that was never added to the backup.
  • Expired credentials. A password changed, an app was logged out after an update, or an access key was rotated, and the job has been failing silently.
  • Full storage. The destination filled up months ago and new data is not being written.
  • Corrupt data. A disk error, a bug or an interrupted upload left files damaged.
  • Lost encryption keys. The backup is fine, but the key or password needed to decrypt it was only stored on the machine that died.
  • Database backups taken wrongly. Copying a database’s files while it is running can produce a backup that will not start.
  • Missing pieces. You restored the data but not the configuration, licences or the software needed to open it.
  • Too slow. The restore works, but takes three days when you needed it in four hours.

Each of these is discovered only by restoring.

A backup you have never restored from is not a backup. It is a hope with a storage bill.

What a restore test should check

Is the data there?

Pick specific files you know should exist, including recent ones, and confirm they are in the backup.

Is it readable?

Open the restored files. Photos should display, documents should open, spreadsheets should have their data. For databases, start the database and run a query.

Is it complete?

Compare file counts or sizes with the original. For a database, check row counts on a few important tables.

Can you do it without the original system?

Restore to a different machine or a clean location, using only what you would have in a real disaster: the backup, the credentials and your notes.

How long did it take?

Time it. Compare with your recovery time objective (RTO). If it is too slow, you have found a problem worth fixing.

A routine for households

Twice a year is enough for most families. Put it in the calendar, for example around Diwali and again in spring.

  1. Pick a random photo album from a year or two ago and restore it to a new folder on a computer. Open a few photos.
  2. Restore one important document (a property paper, an insurance policy) and open it.
  3. On a spare or family member’s phone, sign in and check that the backup is visible.
  4. Confirm you know where the passwords and recovery codes are, and that someone else in the family could find them if you could not.
  5. Check the date of the most recent backup for each phone and laptop.

A routine for businesses

How often What to test
Every backup Job succeeded, and failures send an alert someone reads
Weekly Backup tool’s own integrity check (for example restic check, or pgBackRest’s verify)
Monthly Restore a sample of files and one database to a scratch location; open and query them
Quarterly Full restore of one critical system to a clean machine, timed, following the written runbook
Yearly Restore from the far (off-site) copy, not just the near one

Rotate which system you test so that, over a year, everything important has been restored at least once.

Database-specific checks

  • Restore to a separate server, never over production.
  • Start the database and check it reaches a consistent state.
  • If you use point-in-time recovery (WAL archiving with pgBackRest or WAL-G, for example), restore to a specific time and confirm the data matches what you expect for that moment.
  • Run the application against the restored database if you can; it catches problems that queries alone miss.

Who does the test?

Have someone other than the person who set up the backups run the restore occasionally, following only the written instructions. It is the best way to find steps that live only in one person’s head.

Write down what you found

Keep a simple log:

  • date of test
  • what was restored and from which copy
  • how long it took
  • what went wrong or was unclear
  • what you changed as a result

Over time, this log becomes evidence for auditors, insurers and customers, and a record of steadily improving recovery.

When a test fails

That is a good outcome. A failed test on an ordinary Tuesday is far cheaper than a failed restore during an incident. Fix the cause, update the runbook, and test again soon rather than waiting for the next scheduled round.

A short checklist

  • ☐ Restore tests are in the calendar, with an owner
  • ☐ Encryption keys and credentials are stored somewhere other than the systems being backed up
  • ☐ Restores are done to a separate location, not over the original
  • ☐ Restore times are measured and compared with your targets
  • ☐ The far copy is tested, not only the near one
  • ☐ Results are written down

Closing note

triplicate is being built so that restores from either copy, home region or other continent, are something you can run whenever you like, not only in an emergency. We are still launching; join the waitlist or, for business backups, talk to us.

Keep reading