logoalt Hacker News

JCharantetoday at 7:21 PM2 repliesview on HN

I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead.


Replies

baraketoday at 7:29 PM

Anecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents.

andrenotgianttoday at 7:28 PM

Any data or public links you can share? That surprises me