logoalt Hacker News

saithoundlast Saturday at 5:04 PM4 repliesview on HN

> I’m not sure what point you’re trying to make,

Have you skimmed the linked thread?

> especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second

Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

But even if you insist that you alone are doing some very weird stochastic simulation which is somehow bottlenecked on sourcing random numbers fast enough, the falling in planes phenomenon linked above would make xorshift-type generators a poor choice for most sorts of simulations. It introduces spatial correlations into any sort of lattice dynamics simulation (Ising model, percolation) and every high dimensional Monte Carlo integration. Beyond falling in the planes, since xorshift is linear over GF(2), it is also a particularly bad choice for nondeterministic cellular automata and Boolean dynamical systems which use parity, bit masks, or xors.

AES-CTR throughput on a modern CPU is higher than that of xoshiro256++, and much higher quality. No advances in non-CS PRNGs can beat that while maintaining the same quality. If your stochastic simulation is bottlenecked on random bits, CSPRNGs are still the way to go, and they don't interact in nasty ways with any dynamical system you can actually sinulate quickly.


Replies

moregristlast Saturday at 6:25 PM

> Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.

Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.

Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.

Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.

kbolinolast Sunday at 2:20 PM

Whether your recommendation is valid seems to be quite CPU-dependent. Cf:

  cpu: AMD Ryzen 5 5600X 6-Core Processor             
  BenchmarkAES_CBC-12             100000000               10.96 ns/op
  BenchmarkAES_CTR-12             83161234                14.36 ns/op
  BenchmarkPCG-12                 345336063                3.463 ns/op
  BenchmarkChaCha8-12             174143492                6.894 ns/op
  BenchmarkXoshiro256p-12         254343658                4.717 ns/op
  BenchmarkXoshiro256pp-12        266837442                4.496 ns/op
vs.

  cpu: Apple M4 Pro
  BenchmarkAES_CBC-14             162698020                7.370 ns/op
  BenchmarkAES_CTR-14             242501074                4.954 ns/op
  BenchmarkPCG-14                 197000988                6.083 ns/op
  BenchmarkChaCha8-14             237430095                5.050 ns/op
  BenchmarkXoshiro256p-14         252911710                4.738 ns/op
  BenchmarkXoshiro256pp-14        252656401                4.745 ns/op
Code: https://gist.github.com/kbolino/afbb86f3c9b2bd2f87272801d156...
dgacmulast Saturday at 5:16 PM

This is a very weird hill to die on.

I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.

show 1 reply
SideQuarktoday at 7:19 AM

As others here have pointed out, this is nonsense. The vast majority of PRNG calls on the planet are extremely high perf simulations, where crypto secure versions are a ludicrous cost in speed, energy, and sheer stupidity. That you and others do not understand is simply because you don’t see the places it’s required.

I’ve a PhD, have written papers on PRNGs, have worked in both cs prng and high perf prngs, have done decades of HPC projects, scientific sims. I get called in to develop precisely these high performance systems, and when you want to replace trillions to quadrillions of PRNG calls with one costing 10-1000x more, you’d get deservedly fired immediately.

You keep arguing about AES style code on a CPU. That’s not where people do high performance code. Try implementing AES and a fast prng on a GPU. You’ll soon find out how absolutely terrible cs-prngs are at performance. The measuremt isn’t how many ns per prng. It becomes how many thousands of prng generated per ns.

It’s bafflingly shortsighted for people with zero work in this area to continue to argue this. Choose the right tool for the job. Don’t project ignorance as knowledge. Both are useful advice.