logoalt Hacker News

dalemhurleyyesterday at 8:40 PM12 repliesview on HN

OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.

Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).

Codex is slightly better than Claude Code.

Good on Sam Altman getting back to basics and turning OpenAI around.


Replies

kroatonyesterday at 8:44 PM

I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.

show 5 replies
andxoryesterday at 10:04 PM

> Sol is so much better than Fable 5

I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?

Sol is a much smaller models and it shows. It often misses the forest for the trees.

zachthewfyesterday at 9:29 PM

I’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.

show 2 replies
ChadMoranyesterday at 10:23 PM

Sol better than Fable? What? I've found it to basically be on part with Opus and I max out 2 accounts on both providers every week.

upupupandawayyesterday at 8:59 PM

Their ads business is also doing well. Not "will recover all compute costs" well, but crossed $1b in a few months.

jeffybefffy519yesterday at 9:03 PM

Its funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to...

I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.

show 2 replies
fastballyesterday at 10:16 PM

Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.

show 2 replies
fnordpigletyesterday at 11:03 PM

Codex is missing a few things that Claude code has had for some time like defined plugin subagents and a few other things. But overall it’s fairly capable. The biggest gripe I have is that codex really restricts context window sizes and compaction leads to a lot of grounding work, and overall codex GPT is too literal in many situations - it’s follows direction slavishly, and when subagent reviewers are used, they tend to find increasingly obscure “flaws” on the instruction following impetus, and the harness agent takes them literally as issues to fix even when it leads to bizarre outcomes. For instance I’ve had several runs where it tries to end up building a hermetic system with sha hashing of everything (including operating system binaries and kernels, tool chains, etc) to certify test results are valid, etc. I have to sort of watch it carefully to be sure it’s not drifting into some insane yak shaving corner, which it will happily do for weeks on end.

Claude has the exact opposite problem, especially opus-5, where I literally can’t trust it to print hello world without taking a shortcut, or just simply lying and saying it printed it when it didn’t, behind a giant wall of inscrutable text. I find it very ironic that Anthropic is the vendor of the lazy lying cheating model that does almost everything you tell it to it do.

I’d really kill for something that balances instruction following and loop escaping behavior better. Fable 5.1 does seem a lot better, feeling more like 4.6 behavior, and honestly Sol has improved as well. I’m pretty psyched for the next generation, as I think the competition has heated up so much that things will improve really fast to the point of marginal utility opportunity being increasingly close to epsilon.

show 2 replies
John7878781yesterday at 9:04 PM

This is what Google needs to do and is probably why Demis has stepped back a bit

openaiscookedtoday at 3:09 AM

Killing Sora was one of the worst mistakes they ever made

show 2 replies
Implicatedyesterday at 9:58 PM

> Sol is so much better than Fable 5.

... looks around ...

bsndjdjdjdjyesterday at 10:05 PM

[dead]