It doesn't explain why it doesn't make "I'm not paid enough for this shit" more statistically likely.
LLMs' processing that reproduces statistical patterns of the training data is modified by post-training. That's why we have LLMisms, for example.
LLMs aren't simple patter-matchers/pattern-predictors. They are incredibly complex systems that capture some aspects of the systems that produce the training data.