They don't report the pass@1 success rate. They sample multiple solutions from Opus/Fable and CLM decides which one to submit, that's why they get >80%.
Maybe I'm misunderstanding this but when would you ever use it this way? If you're already willing to call Opus/Fable, then isn't the obvious comparison whether Opus/Fable can choose among sampled solutions better or worse than their fast model? If you're willing to pay many seconds for many code samples from a slow model, it's contrived to imagine you care about picking between them in ms.
Maybe I'm misunderstanding this but when would you ever use it this way? If you're already willing to call Opus/Fable, then isn't the obvious comparison whether Opus/Fable can choose among sampled solutions better or worse than their fast model? If you're willing to pay many seconds for many code samples from a slow model, it's contrived to imagine you care about picking between them in ms.