'better results' in terms of what though? A benchmark, or code that I would actually click "approve" on in a pull request scenario?
apologies, I should have clarified the 'better' claim
- same task result (passed) - finished faster - fewer tokens, less cost - fewer requests for inference - fewer tool calls - less peak RAM
apologies, I should have clarified the 'better' claim