I always find it confusing that a meaningful volume of the comments are saying "this reached parity with SOTA models. Best $/task."
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
I guess both are true and for everyone at some point. All models, even SOTA, fail. When they fail, it is quite frustrating. Additionally, some models are very cheap to run and use. When Deepseek fails the cost was minutes and pennies.
I was in the first group up until last week.. now, the second one.
For anything even moderately complex.. like, even low end of complexity, this model behaves maximum like gpt-5.6-luna-high .. nothing more.
Yesterday itself I gave it a coding task in some existing moderately complex small project, and i was using xhigh thinking effort, it was unable to cover all edge cases... and i had already got it to review, and then fix, 3 more times, after the first initial one.
Still it left 2 edge cases.
Then, reverted full code, gave sol-high the same task, it took well over 20 minutes, and completed it in one go with zero edge cases remaining.
I am not using it for anything serious anymore.