If you can turn the problem into a small kernel operating on a heap of data (or a hierarchy), the GPU almost always wins for culling, especially if you pipe the cull into the draw with GPU-driven rendering.
If the author used GPU culling it would likely be faster on modern hardware, they just said they can't because of platform restrictions. But that's what AAA games do.
The modern non-nanite techniques here are basically to regularize to grids, cull the grids on the GPU, and then cull more with a HZB. That's your "mipped occlusion boxes", except it's actually very cheap to do this because you're reusing depth you already had from previous frames, the test for each object just a few texture samples, and it all stays on the GPU. As a bonus, you can do your LODs on the GPU at the same time, saving even more CPU work and memory bandwidth.
Also, depth rejection is not going to help with the problems that culling solves; draw, vertex processing, raster, and then pixel tests is much more expensive than a cull test before doing any of this. And if you're forward rendering w/expensive fragment shader, the overdraw of relying on the Z buffer to do your "culling" can kill you.
I’ve been having a lot of fun implementing most of these on the GPU, and being okay with saying “sorry” to integrated graphics.
The pure bandwidth on a GPU is insane. 1000fps with infinite render distance.
>that's what AAA games do
yeah where the pixel work and the vert count is magnitudes higher. I forgot to mention in my comment that I was talking about low-poly / pixel stuff like this, with a very simple PS and being pretty much API bound or memory-bound in perf