logoalt Hacker News

cbg0yesterday at 8:06 PM2 repliesview on HN

But is the score really reflective of the quality or are both models benchmaxxing?


Replies

gpt5today at 12:42 AM

Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.

bermudiyesterday at 8:44 PM

Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.