A model unexpectedly having internet access during evaluation is clearly a security (and consecutively legal?) issue when it comes to the capture the flag evaluations. But it also potentially invalidates other non-security evaluations or makes them less impressive, as the model might have just googled solutions.
Reading this, I felt like I was too harsh on OpenAI for not monitoring their network, as they at least attempted to sandbox a workload.