logoalt Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

418 pointsby seelosyesterday at 3:29 PM174 commentsview on HN

Comments

postalcoderyesterday at 4:27 PM

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

show 8 replies
gruezyesterday at 5:05 PM

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

https://www.youtube.com/watch?v=tNmgmwEtoWE

As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

show 5 replies
nullbioyesterday at 4:31 PM

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

show 3 replies
thimble_iotoday at 9:53 AM

92.8% on TB2.1 dropping to 27.3% on TB4 is the only number that matters. The rest is marketing.

pkilgoreyesterday at 4:56 PM

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

show 2 replies
TheJCDentonyesterday at 4:28 PM

> SWE-2 is post-trained from Kimi K3

On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

show 1 reply
Ozzie_osmantoday at 3:54 AM

Seeing a lot of Devin-skepticism in the comments, so a data point: here at Monarch we've built a pretty awesome cloud-based coding agent powered by Devin. My team did a bakeoff of all the cloud-based coding agents, and when they suggested Devin, I was really skeptical, but after looking through the data and the tests: Devin worked the best autonomously, it had the most robust security functionality, and it was built to handle multi-repo changes. So we moved forward.

It handles about a third of our pull requests at the moment (mostly smaller changes / bug fixes). Nothing gets merged to prod unless it's reviewed by a human (usually two, the person who prompted and then another reviewer), but Devin has been really powerful for us in terms of productivity on those drudge-y tasks. For more complicated tasks, folks may still have Devin start the implementation and put out a draft before they pull it into Claude Code or Cursor.

(We have invested in a bunch of stuff that makes it easy to use/share across agents, including an MCP gateway, shared repo with skills and context, tools for functional/browser testing and image/video capture for any change, etc).

mydreamofyesterday at 4:10 PM

Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?

show 1 reply
bobtheborgyesterday at 4:40 PM

SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.

Looking forward to 2 -- maybe it'll be usable

captainregexyesterday at 9:26 PM

I am skeptical. Lived experience is what matters and I don’t have anyone in my life (Devin shop) saying good things about SWE other than it’s free. Hope I’m wrong and it’s not so bad this time

CyLithyesterday at 6:10 PM

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.

show 1 reply
Take8435yesterday at 6:17 PM

Post made by account 2 days ago.

show 1 reply
bluelightning2kyesterday at 4:49 PM

I like Cognition as a company and hope they succeed. Seemingly excellent engineering org.

I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.

andaiyesterday at 7:24 PM

Their benchmark used to show other metrics, like output tokens and time, but now only shows cost:

https://cognition.com/frontiercode

Which is too bad, since all of the gains here appear to be from massively reduced output tokens?

The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens.

Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true.

eyerisyesterday at 4:38 PM

Wonder if this was the model that drove factoring the rsa-260

The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.

peloratyesterday at 7:15 PM

Unless it can do CAD via computer-use how can you say it rivals GPT-Astra?

sbseitzyesterday at 6:58 PM

Why doesn't clickbait trash like this get moderated ?

show 1 reply
alansaberyesterday at 6:29 PM

Fair enough that they did a "propoganda and censorship" eval but not sure why i'd care about that in my highly juiced SWE kimi FT.

gexlatoday at 5:26 AM

If everything basically rivals Fable, then why is everything still using it for comparison?

show 1 reply
scronkfinkleyesterday at 4:13 PM

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

show 2 replies
yipinwongyesterday at 10:49 PM

Incest in human biology causes mutations that's bad in the long term.

Same for AI models trained on Kimi-3 or other models like Chinese models do. They suffer from the same issue.

monkeydustyesterday at 4:18 PM

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.

show 2 replies
Tsarpyesterday at 4:11 PM

"SWE-2 is post-trained from Kimi K3"

show 2 replies
llmslaveyesterday at 4:35 PM

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

show 2 replies
gigatexaltoday at 3:53 AM

Not avail on open router?

microdrumyesterday at 11:27 PM

Do they have anything as good as Amp (which is able to use free models)? Amp has been in the lead for almost a year now and doesn't seem to be relinquishing it.

ltsSmittyyesterday at 4:35 PM

Well written and good diagrams. No idea the verity of the TMBB (trust me bro benchmarks) but it was pleasing to look at

m3kw9yesterday at 5:32 PM

I'm using Codex, Gemini etc, they all have desktop apps and have a plan, how do i use SWE-2? Thats is a problem they have. I'm not about to switch out my workflow and plans with a shiny LLM that looks benchmaxxed and graph maxxed.

_doctor_loveyesterday at 4:10 PM

SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.

wqash71yesterday at 6:16 PM

The horrible website is made by Claude or Cognition is distilled. I'm so tired of it all.

logicalleeyesterday at 8:03 PM

all this while Cognition doesn't even make the top page of search results, that's some impressive stealth.[1]

(just to be clear, I am a far-left activist who spends most of my time working on funding Social Security Trust Funds (OASI & DI Solvency), which could impact my search results - this was while I was logged in.)

[1]https://www.google.com/search?q=what%27s+cognition+in+ai or https://imgur.com/a/UdxtnGg

nujan_devtoday at 2:52 AM

[dead]