logoalt Hacker News

Aurornistoday at 6:45 PM6 repliesview on HN

> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.


Replies

ruszkitoday at 8:35 PM

Overfitting to benchmarks. And puff, you have the exact same effect.

talon8635today at 7:36 PM

This is a good point I hadn’t considered, thank you.

Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?

Sorry if it’s a dumb question, I don’t really know much about the topic.

pixl97today at 6:50 PM

Far more likely it's about reducing costs.

mobelkhtoday at 7:05 PM

but there is a gap between benchmarks and user feel.

Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.

AnimalMuppettoday at 8:09 PM

Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.

well_ackshuallytoday at 6:58 PM

* Release new model that scores an arbitrary 100 on a benchmark

* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.

* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.

* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.

* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.

Bis repetitae.

show 2 replies