logoalt Hacker News

epolanskiyesterday at 10:20 PM2 repliesview on HN

I don't get the point of these benchmarks, what are they supposed to represent practically?


Replies

runarbergyesterday at 10:25 PM

For me this looks ideological (or even political), not practical. The theory is that LLMs are approaching general intelligence (whatever that means) and that the more generic of a task they can perform—no matter how badly—the closer we are to AGI.

Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.

show 1 reply
viccisyesterday at 10:24 PM

The ability of an LLM to produce something not in its training data set.

show 1 reply