logoalt Hacker News

pitched • yesterday at 12:51 PM • 3 replies • view on HN

FTA, this is a fault-tolerant server where all hardware is redundant and can be hot swapped while running. The servers you’re thinking about are redundant at the software level so rebooting one server won’t cause the service to go down.

What is remarkable is that apparently that VOS thing has never crashed. I wonder if they also do software redundancy under the hood to keep uptime going during reboots. If the CPU is hotswappable, it must have something.


Replies

wildzzz • yesterday at 1:52 PM

Probably uses formally verified code along with plenty of housekeeping processes to keep any failures from shitting the whole bed. In critical system design, you build in redundancies that work in parallel such that any one failure will not interrupt the system.

The main computer system in the Space Shuttle is an excellent example of this. It had 5 identical IBM System/4 Pi machines. Three of them ran identical code and handled the same work. The fourth ran a completely different codebase to handle the same work, preventing a bug in the main code from killing the whole system. A fifth computer handled other tasks but could be swapped over to the critical role if needed. You could lose 2/5 computers and still have insurance against a cosmic ray flipping a bit.

theamk • yesterday at 3:51 PM

It's really not that hard to keep computers non-crashing, as long as you have good hardware and run a limited subset of software.

Many servers I've owned had multiple years of uptime, and the only reason they'd go down is because they will get decommissioned or because of power outage.

The article says:

> disk drives, power supplies and some other components have been replaced but Hogan estimates that close to 80% of the system is original.

so I am guessing there was no reboots, nor CPU replacements.

serf • yesterday at 1:57 PM

VOS is a parallel lockstep OS. You drop nodes and replace them to keep the whole operational.