logoalt Hacker News

CuriouslyCtoday at 2:21 PM0 repliesview on HN

Small models can be super smart. Big models mostly give you baked in world knowledge, domain flexibility and long context stability/coherence. I wouldn't be surprised if we see Fable level smarts in a coding model that fits in 24GB by next year, but it'll be a savant style coder that needs in context learning, and it'll get very wonky after >100-200k tokens consumed.