logoalt Hacker News

rootusrootustoday at 4:38 PM5 repliesview on HN

It feels like 5.6-Sol is already fairly close to Fable, and in some ways exceeds it. Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look ... it found an oversight and told me about it, and when I then fed that observation back into Claude it acknowledged the miss.

I've noticed also that 5.6-Sol is more concise with output than Fable (and let's not talk about Opus, which is even more wordy).


Replies

rsyringtoday at 5:04 PM

It's common for different models to find holes in another's work. There are various good reasons for that.

FWIW, we use ChatGPT for our primary model and use Claude to do the reviews. This works better than ChatGPT doing it's own review even with a clean session/context.

show 3 replies
Taronartoday at 6:09 PM

Did you try to say "think more deeply about this problem" to fable after getting your solution, having one model focused on creation then blaming it for not doing proper review when the other model was told to focus sol-ely (pun intended) on review is not a fair apples to apples compaision

show 1 reply
ceejayoztoday at 8:47 PM

> Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look…

You should be doing this for every solution.

Even Fable reviewing itself will find issues, unproven assertions, etc. Same for Codex models. A review loop is critical.

janalsncmtoday at 8:23 PM

The fair comparison would be to also do the reverse: start with Sol then have Fable clean up. Then compare the Fable-Sol and Sol-Fable outputs side by side.

importtoday at 9:10 PM

I used to review each others work, Sol is amazing at review and finding what’s missing.