If you don’t know the language you can’t evaluate if the models really found bugs.
That’s like translating a text to another language without knowing the language
This is wildly overblown. I've been working with agents for a good while, read tens of thousands of generated Python and the language factor is actually the part they get right that humans don't.
This is wildly overblown. I've been working with agents for a good while, read tens of thousands of generated Python and the language factor is actually the part they get right that humans don't.