I believe the AI labs are weakly motivated to train strongly against cheating when it helps with benchmarks.
Does it help with benchmarks? Are you saying there are examples of benchmarks where the models have solved the problem by cheating?
Does it help with benchmarks? Are you saying there are examples of benchmarks where the models have solved the problem by cheating?