> They do not beat opus on real-world usage
We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios (mostly Rust, some C for microcontroller stuff). Qwen3.6-27B scores only 4% lower for pass@1, n=250 compared to Opus-4.8.
For the labeled dataset, the average PR size they're being measured against is around 1.5k SLOC.
This is very much "real-world usage" for us. The sort of change sets that come in daily/weekly and are solving non-trivial issues in the respective codebases.
As is usually the case, the most broad claims from both the labs and from the consequent pushback are talking past each other.
[dead]
>> We have an internal eval that measures performance on tasks for a handful of embedded systems repos for our mmWave radios
Okay but the parent said real-world usage, presumably meaning coding tasks.
We have a whole bunch of complex evals that Haiku 4.5 passes. That doesn't mean it is a good model for coding.
How much does it score though? 0% would be 4% less if Opus was at 4%. Unless you mean relative fraction not percentage points - but people usually mean percentage points in such situations.
Let us know when you have Qwen vs Qwen comparison stats. As long as there's not a regression, that'd be awesome.