logoalt Hacker News

2001zhaozhaoyesterday at 6:26 PM2 repliesview on HN

It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination


Replies

artemisartyesterday at 7:39 PM

No it's a mean of 5 runs.

> We report FrontierCode’s overall score, a composite measure that grades each patch on blocking functional criteria (held-out unit tests) together with weighted code-quality rubric criteria, as mean@5.

They don't explain more in the system card, I guess higher effort levels could loose points on the code quality / scope / style / maintainability stuff?

dannywtoday at 6:21 AM

If you're doing passes@1, especially for long-horizon agentic benchmarks, you might as well as not do the benchmark at all.