logoalt Hacker News

mordaeyesterday at 11:07 PM1 replyview on HN

This is a terrible benchmark. It literally tests the models on their ability to track shifting line numbers. If they cannot keep up, no amount of abstract reasoning can redeem them.


Replies

lordmauvetoday at 6:54 AM

Where did you get that idea? It uses mini-swe-agent, same as SWE-Bench.

https://github.com/datacurve-ai/deep-swe

show 1 reply