Pre-training data is pre-tokenized ahead of time before being used to not waste any GPU compute.
A massive speedup like this is a nice efficiency savings on some of these data pipelines for sure.