logoalt Hacker News

kuukyo • today at 2:22 PM • 2 replies • view on HN

They don't report the pass@1 success rate. They sample multiple solutions from Opus/Fable and CLM decides which one to submit, that's why they get >80%.


Replies

abeppu • today at 2:37 PM

Maybe I'm misunderstanding this but when would you ever use it this way? If you're already willing to call Opus/Fable, then isn't the obvious comparison whether Opus/Fable can choose among sampled solutions better or worse than their fast model? If you're willing to pay many seconds for many code samples from a slow model, it's contrived to imagine you care about picking between them in ms.

vessenes • today at 2:28 PM

Ah-ha. Interesting! Thanks.