logoalt Hacker News

ericdyesterday at 10:56 PM3 repliesview on HN

Gpt-oss—120b is like 1000 years old in AI years, whereas Qwen 3.8 27b is pretty young. What you’re seeing is that parameters aren’t apples to apples, and at a given parameter level, the new models are much, much better than the ones from a year or two ago. Like, to a comical degree.


Replies

MichaelZuoyesterday at 11:24 PM

Wasnt this known by everyone who cared to pay attention?

It practically became a joke about how a huge amount of the training data for GPT-4 was bottom of the barrel reddit vomit and obvious bot spam. Leading to many bizarre edge cases.

kennywinkertoday at 12:18 AM

Does that not prove my point? Bigger doesn’t automatically mean better. Quality of training data, and model structure, matters as much or more than size

show 1 reply