logoalt Hacker News

pfdietzyesterday at 5:52 PM1 replyview on HN

I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?


Replies

dparkyesterday at 6:14 PM

They aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.

show 1 reply