logoalt Hacker News

WhitneyLandtoday at 12:48 PM4 repliesview on HN

It’s exciting that a model scoring this high is dirt cheap.

It’s also so inefficient, when they release the full performance numbers it’s not going to be good.

One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.


Replies

Bnjorogetoday at 1:14 PM

Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal

show 2 replies
mdp2021today at 5:14 PM

> inefficient

That depends. Is it also more reliable?

If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.

onlyrealcuzzotoday at 2:20 PM

Roughly equivalent to Gemini 3.6 Flash in capabilities at 1/20th the price...

Mind you, until the recent price cuts to Luna - Gemini 3.6 Flash wasn't even egregiously priced (but oh how things change in just 1 week).

ptole_mytoday at 5:25 PM

you are correct. but in my experience -- not benchmarks -- g flash 3.6 is SO BAD for coding. I'm using all vendors all day and gemini is the worst by far. I built my own semi-deterministic orchestrator for coding agents.