logoalt Hacker News

mattlondonyesterday at 3:42 PM13 repliesview on HN

Currently top at https://deepswe.datacurve.ai - beating Opus 5!

https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!

Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.


Replies

theHocineSaadyesterday at 4:08 PM

As of writing this comment, Claude Opus 5 has an intelligence score of 63, not 59 (it's not the same as Gemini 3.8 Flash).

With a score of 59, Gemini 3.8 Flash is in eighth place, falling behind even Grok 4.6, Kimi k3, and GLM 5.3.

https://imgur.com/a/BMOJBED

show 3 replies
markasoftwareyesterday at 3:54 PM

On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63.

Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference.

show 1 reply
WarmWashyesterday at 3:47 PM

The benchmark also doesn't include speed. You almost think something has gone wrong when using it because it returns full responses so incredibly fast.

show 2 replies
onlyrealcuzzoyesterday at 3:53 PM

The rumor is that 3.9 is an equal improvement in all directions, and that it should be another fast follow on like 3.7 and 3.8 were.

It's almost across the board better than Terra at less than half the price. 3.9 is likely to approach Sol at the 1/10th the price.

Hopefully OpenAI releases Astra first, and it's not only better than Sol but significantly cheaper, too.

show 1 reply
bertiliyesterday at 4:09 PM

A fifth of the cost of Opus 5! Google is certainly pushing the completion with this.

show 1 reply
ttulyesterday at 3:49 PM

Crushing it on DeepSWE is a very big deal. Excited to give this a try.

show 2 replies
Gecko4072yesterday at 3:43 PM

Google - we're so back

show 1 reply
kimjune01yesterday at 4:36 PM

deepswe is public and can be considered contaminated.

sunaookamiyesterday at 3:50 PM

>shows an intelligence score of 59, the same as Opus 5!

...on Medium reasoning. Claude Opus 5 (high) is the default in e.g. Claude Code and scores 61. Still very impressive.

satvikpendemyesterday at 3:50 PM

We'll see about that. I suspect benchmaxxing as all the labs do as I haven't found Gemini models to be nearly as good in agentic engineering compared to Claude or GPT models.

show 2 replies
WhitneyLandyesterday at 4:41 PM

There are important gaps in that hot take.

For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.

jrfloyesterday at 4:24 PM

sidenote, but wow sonnet 5 is shockingly bad on this benchmark.

show 1 reply
pkos98yesterday at 4:11 PM

Wait a week with your judgement - most likely, Google is just bench-maxing very hard. If you look at the previous Flash models and the announcement on Google I/O, it was an absolute disaster. Reality diverged very much from the marketing (supposedly great benchmarks).