No, a cluster can server multiple users at the same time, providers cap the tok/s so that one cluster can run inference on multiple inputs at the same time. OpenAI with their new ultrafast mode is probably reserving the whole cluster or prioritizing requests of ultrafast users above others with a higher tok/s hence the high price and high speed. There's many other knobs providers tweak that they don't show the users, for example I doubt many providers are hosting the full FP16 version.
It's not based on rate limiting at all.
The "expensive part" of generating the next token is streaming in the model weights from memory. The computations are relatively simple, which is called a "low arithmetic intensity" in industry jargon.
So what they do is batch multiple chats together and compute the neuron activations for all of them together.
This is vaguely similar to how some database engines work, where if multiple users need to run a "whole table scan" query, the additional users "join" the streaming workload of the first query mid-way, then loop back around to complete the first part that they missed. The AI accelerators don't do this looping, but the concept is the same: amortize the expensive I/O over multiple computations running in parallel.
The "turbo mode" token rate thing is almost certainly your query getting sent to slower or faster hardware, like B200 vs newer B300 kit.