logoalt Hacker News

TacticalCoder • today at 11:34 AM • 1 reply • view on HN

I spent lots of time thinking about backups. I've got files dating back to 1991 and even some from the 80s (but I don't care about those so didn't bother backing them up: when the last 5"1/4 shall stop reading, those will be gone and it's fine).

I've even got stuff like copies of early websites made by friends, websites long gone and I'm pretty sure I'm the only one to still have backups of those.

To me verifying the backup should be part of the backup procedure itself and so I did just that: my backup procedure does verify that the backups can be decrypted/unarchived and it then verifies the files inside the backup.

Now the thing is: you don't need to verify 100% of your backups all the time (as in I don't decrypt and verify the checksums of 100% of my backups all the time). You can do random sampling: over n backups, you can be reasonably sure at least x% of the files are correct.

So random sampling is part of my backup procedure.

As those are encrypted backups, the verification also ensures that the backups can be decrypted (would be too bad otherwise, wouldn't it?).

I also believe the backuping procedure should not be able to go wild and destroy data. So my backup procedure is done by a podman container that access the volume with all my data (the one that needs to be backed up) read-only.

> To believe in one's backups is one thing. To have to use them is another."

Yes, which is why a proper backup procedure verifies the backup it just created. The greenlight is only given to the backup that's just been made once it's been verified by my verification script.

Some stuff can be verified automatically: for example say you backup a Git repo: a strict, full, git fsck on the backup (that is: on the Git repo once pulled out of the backup, before greenlighting it) works fine.

For other things I've got my own verifications.

> These hundreds of corrupted files have been flowing through my backup system. Now I do not know which files are clean and which are not.

Which is why many of my files, which I know aren't supposed to ever change anymore, have a partial checksum added as part of the filename.

I don't have:

    DSC0983747.JPG
but:

    DSC0983747-b3-7e228491a0.JPG
where 7e228491a0 are the first 40 bits of the Blake3 checksum of the file.

Then it become very complicated for "something" (bit rot or malicious) to silently corrupt my stuff for the backuping procedure (which only has read-only access to data, remember) goes crazy bonkers and warns me as soon as two identically named files (one on the last backup and the current version) have an identical name but different content and one doesn't match the checksum (this is cause for a "stop the world" and immediate enquiry: and, yup, it already allowed me to catch a corrupted file).

I've moreover got a sheet of paper, laminated, that explains how to access the backups and that explanation is saved in many places (safe at the bank, safe at my brother's house in another country, etc.): should say my place burn with me inside, but my wife be safe, she'd be able to access the backups too.

As for my main data, it's RAID on ZFS on a machine with ECC and that's where the backups are made (and automated) from.

Proper backup procedure is not hard: it just takes some time to set up properly, once. Then it's done for a lifetime.


Replies

bigbuppo • today at 5:55 PM

Yeah... I'm going to have to incorporate the checksum-in-the-name idea into my photography workflow.