I think language itself is compression, so the arxiv paper tracks for me.
Viz. if Language is compression (of thought / culture / the tacit je ne sait quois of being-to-being communication etc.), then definitionally, Language Modelling must also be Compression.
Except, language is an arbitrarily lossy compressor, who's "compression-prediction equivalence" is indeterminate and unstable, because Language co-evolves constantly; both as a function of or response to culture, as well as an influencer of culture.
So, the subjective-objective goodness of Language Models (of any kind of language) would be, at best, upper-bounded by the compression-prediction equivalence of the Languages corpus itself. And that is assuming the language corpus is perfect in every way---it captures all knowledge expressible by language and it is always in-sync with live evolution of all language expression and evolution (i.e. LLM training is not a batch job, but a real-time present continuous process).
For example, to my layperson eyes, the mathematical language of proofs actively weeds out ambiguity of subjective interpretation. Ideally, a proof ought to lead to the exact same conclusion on every single reading by any reader who can follow the steps. A proof also holds only if the rest of the formal, explicit, inviolable, internally-consistent set of axioms and results holds.
So it stands to reason that mathematical prose of proofs, being optimised as mechanical procedure of taking an open question to a deterministically closed solution, has better odds of approximating the tacit aspects of mathematical derivation.
Which makes an LLM able to construct a mathematical proof, which is mind-melting to say the least.
However, I wonder, can LLMs dream of mathematical sheep?
Language is compression of a sort. A dictionary is a decompressor. You look up one word, and you may get a paragraph about its meaning. An encyclopedia can be thought of as roughly the same with more detail for some nouns.
It makes a lot of sense why dictionary-based compression is named the way it is. A shorter symbol is used to store information that would take more symbols in the uncompressed corpus, if the shorter symbol hadn't been assigned to represent it. That's in a way just what an actual dictionary on your English professor's shelf does. The big difference is your compressor is coining new short symbols all the time.