logoalt Hacker News

lumostyesterday at 3:57 PM0 repliesview on HN

yes, above 50% I don't observe a significant improvement, the models scoring above 65% like opus4.8 seem worse than gpt5.5/opus4.6. DeepSWE appears to align more closely to my personal experience.