logoalt Hacker News

vessenes • today at 2:01 PM • 1 reply • view on HN

>"CLM-8B also sets a new SOTA on challenging agentic coding benchmarks, including DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).

I didn't see any details on this on the announce page. And I don't believe it. Astra x-high pass@1 on DeepSWE is 74% +/- 3%. (https://deepswe.datacurve.ai).

That said, love seeing some of these new architectures get people exploring. But, surely somebody is incorrect here inre: those numbers.


Replies

kuukyo • today at 2:22 PM

They don't report the pass@1 success rate. They sample multiple solutions from Opus/Fable and CLM decides which one to submit, that's why they get >80%.

➕ show 2 replies