Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.