logoalt Hacker News

spongebobstoestoday at 12:49 PM1 replyview on HN

this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt

test Codex, not Sol. test Claude code, not Opus


Replies

ChrisLTDtoday at 1:31 PM

There are other benchmarks for that