DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
But is the score really reflective of the quality or are both models benchmaxxing?
how much of it is from reallocation of staff to ai training and labeling
when are we going to stop pretending these benchmarks have any meaning?
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.