I’ve experimented a bit with this approach for quasi-benchmarking models. Most models nowadays will keep trying until you cut them off, they’re happy to go into millions of tokens for an easy task, digging deeper and deeper into crap design. Compared to a clean reference implementation you get easily 20x token difference at task 11+ for the same model. Somewhat ironically I never published the results because the both the harness and the task set were vibe coded and I didn’t have the energy to clean it all up.