This study appears to be evidence against that: the model failed the benchmark but arrived at the correct answer.
Half of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.
Half of the benchmarks don't have published answers. That's why training data providers have been hiring physicists to solve them.