logoalt Hacker News

maxignolyesterday at 3:43 PM1 replyview on HN

> my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.

What do you mean by that ? If the model is higher than 50% on swebench pro then it tends to drift from what you like it to do, like DeepSWE benchmarks ?


Replies

lumostyesterday at 3:57 PM

yes, above 50% I don't observe a significant improvement, the models scoring above 65% like opus4.8 seem worse than gpt5.5/opus4.6. DeepSWE appears to align more closely to my personal experience.