logoalt Hacker News

recklesslast Sunday at 9:30 AM2 repliesview on HN

5.6-sol would be a better comparison given it's general availability and usage allowances


Replies

hodgehog11last Sunday at 11:11 PM

They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.

ferrouswheellast Sunday at 9:34 AM

And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.