I guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?
They do not, as they're aiming for full replacement rather than augmentation of human users.
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
They do not, as they're aiming for full replacement rather than augmentation of human users.
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.