logoalt Hacker News

jstummbilligyesterday at 10:12 PM6 repliesview on HN

I am trying to estimate if my reaction to seeing GPT-5.6 Sol last on that list is reasonable or or mostly emotional and find that I have no way of telling.


Replies

CuriouslyCtoday at 12:46 AM

Any bench that puts GLM 5.3 ahead of 5.6 Sol is highly sus. They've been my two daily drivers since release, and I like GLM 5.3, but it's definitely not better than Sol, it's more ~Terra, while being significantly slower.

show 1 reply
WD-42yesterday at 10:57 PM

Why would you get emotional over a model? They got you that good?

beefsacktoday at 1:19 AM

There's an issue with GPT-5.6 Sol where it sometimes starts mixing thinking with output and stops working[1]. Once it starts doing that, the session is essentially cooked and you need to do a bit of gymnastics if you want to recover it.

This happens to me more commonly in large projects (>100k LOC) and in those projects it seems to happen every few sessions. I feel this specific benchmark would be impacted by this more than the smaller contrived benchmarks.

[1]: https://github.com/openai/codex/issues/37524

show 1 reply
CompoundEyesyesterday at 10:31 PM

I do think it’s the wizard not the wand at this point given a decent model. These benchmarks don’t have the wizard.

Otherwise I wouldn’t see others in the exact same codebase struggle and underutilize agents while others thrive using the exact same ones.

show 1 reply
didgeoridooyesterday at 10:37 PM

Sol failing mostly on “unverified assumptions” and rarely hitting “integration errors” seems about right to me. I think Sol is second only to Astra (and miles ahead of even Fable) in architecting & engineering the right implementation — but only if you are extremely specific and provide tight guidelines and guardrails. If you give it a one-liner… you’re going to have a bad (SHA-256-hash-verified) time.

show 2 replies
dimgltoday at 1:52 AM

I found 5.6 Sol to be extremely underwhelming.