I'm more interested in the converse question: how well does an LLM perform as a compressor, compared to gzip (ignoring its insanely lower speed)?
hallucinations are lossy compression artefacts
Much better
can they be considered to have compressed the entirety of their training data into their weights?
Top contestant in the Hutter Prize uses a neural network for compression. So fair to say, LLMs would perform pretty well compared to gzip.