logoalt Hacker News

the_duke • today at 5:30 PM • 19 replies • view on HN

The GPT 6 release was ... not great.

Sol 6 was so bad that I switched over to Opus 5.5 exclusively.

Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.

Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.

I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.

(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)


Replies

wkcheng • today at 6:23 PM

I agree, and I haven't seen other people mention this! The benchmarks for GPT 6 Sol are great, but realistically it does not seem better than 5.6 Sol. 6-Sol is noticeably worse for code reviews (worse than Deepseek 4.1 flash), has implementation issues (requires more rounds of code reviews and fixes to get to a serviceable state). Opus 5.5 is much much better.

I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.

If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.

stldev • today at 7:39 PM

My experience as well.

For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.

For modeling and artwork, Astra has been great routinely outperforming Kimi.

This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.

I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.

koyote • today at 10:39 PM

I think the fact that Sol 6 appeared higher than Sonnet 4 on benchmarks shows that benchmarks are completely rubbish and useless.

I've never seen such a large degradation in intelligence in a model until I tried out Sol 6 after having used 5.6 almost exclusively for several weeks.

moshegramovsky • today at 7:47 PM

100% hard agree.

I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)

It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.

I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.

OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.

Just because you can, doesn't mean you should.

pampas • today at 8:56 PM

That's my experience too. GPT-6 Sol tends to rabbit hole and over engineer things.

ozgung • today at 6:38 PM

Maybe OpenAI was the only one pacing the frontier.

jrflo • today at 8:08 PM

I'm in the same boat, I'll give 6.1 a shot but I'll probably hop over to Anthropic now that the $200 tier has equivalent weekly usage between the two of them.

jsw97 • today at 8:34 PM

After seeing a number of hit or miss releases from both OpenAI and Anthropic my default is to stay put on what I’m using and then free ride on discerning eager adopters by reading their reviews. (Thanks!) Still on sol 5.6 with an occasional advice from Astra. Also I feel like I kind of get used to the models but maybe that’s just my imagination.

trentnix • today at 6:34 PM

That's not been my experience. My experience with Astra (I use it at home writing Go and C) for coding has been fantastic. Opus 5.5 (I use it for work writing C#) seems faster than Opus 5, but it doesn't seem demonstrably better to my eyes and is still prone to word vomit.

➕ show 2 replies
bitexploder • today at 7:17 PM

I have likewise not been impressed with Astra 6 for most things. It is good, but Opus 5.5 seems just as good or better and I have had Opus 5.5 workers just... hammering since release and cannot spend all of my quota yet.

beebmam • today at 8:30 PM

gpt-5.6-sol is significantly better than gpt-6-sol. Not impressed with this new line.

➕ show 1 reply
nxc18 • today at 5:35 PM

How does this jive with the exponential growth claims? Theoretically sol models are better than the 4 series models I was using at the beginning of the year, but in practice the results don’t seem to be much better. They always nerf the models over the course of the release so it _looks_ like the next version is better but I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.

➕ show 2 replies
NorthSouthNorth • today at 7:18 PM

I shilled so hard to a friend that he actually swapped decided to swap over to Codex. I feel a bit guilty now lol (tbh Astra is a great model, but 5.5 is just brilliant).

setnone • today at 6:36 PM

yeah i can relate, sol 6 is definitely dumber than 5.6, lazier too, i hope it's just roll out pains

sunaookami • today at 7:26 PM

gpt-6-luna is terrible. It leaks tool calls and markers in the output like crazy, there is definitely something wrong here. gpt-5.6-terra works fine. Also, gpt-6-luna was sneakily added to the 1 mio free tokens group instead of 10 mio. like gpt-5.6-luna: https://help.openai.com/en/articles/10306912-sharing-feedbac...

soulofmischief • today at 8:09 PM

I have had the same exact experience. I feel like I'm working with 5.3 again. It is alarming how degraded the experience has become over the last month.

What was a pleasant and productive experience is becoming increasingly frustrating and draining.

jstummbillig • today at 5:39 PM

Eh. What? Is this common sentiment?

I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?

(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)

➕ show 5 replies
jeffybefffy519 • today at 9:04 PM

Its almost like the "frontier" is a load of marketing bullshit and we should ignore it....

btbuildem • today at 6:12 PM

That mirrors how disappointing Opus 5 and Fable were, for anything beyond one-shotted tasks or shiny demos. Maybe OAI is just a step behind Anthropic? Opus 5.5 seems like the real deal again, consistent good results on large, complex codebases.