logoalt Hacker News

lumostyesterday at 12:14 PM2 repliesview on HN

Honestly my anecdotal experience is that fable is benchmaxxed. I have not observed useful gains for Claude since 4.6, with each model iteration making progressively poorer decisions in pursuit of its goal.

The 5.5/5.6 series has performed quite well however. My guess is that my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.


Replies

maxignolyesterday at 3:43 PM

> my use cases stop aligning to swebench pro around 50% accuracy, and more closely align with DeepSWE.

What do you mean by that ? If the model is higher than 50% on swebench pro then it tends to drift from what you like it to do, like DeepSWE benchmarks ?

show 1 reply
carterschonwaldyesterday at 1:30 PM

this mirrors my approx experience.