logoalt Hacker News

ericd • yesterday at 11:03 PM • 0 replies • view on HN

This seems to imply that training will ever be done? But yeah, I think the idea is that the appetite for thinking-on-tap will be enormous.

Even for smaller models, I think they’ve found that training an enormous, inefficient model and then distilling it internally to something much more efficient to serve is the way to go.