These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.
I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?
I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?