logoalt Hacker News

dumberquestionstoday at 3:27 PM1 replyview on HN

"..and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token."

"3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, as seen in DeepSWE (49% vs. 37%)"

So which one is it? 65% or 49%?


Replies

petutoday at 3:29 PM

First sentence is about token efficiency.

show 1 reply