It's differently interesting. It's interesting that they didn't get an important detail right with a partner, so more on Anthropic's attention to important details rather than the power of their models.
A model unexpectedly having internet access during evaluation is clearly a security (and consecutively legal?) issue when it comes to the capture the flag evaluations. But it also potentially invalidates other non-security evaluations or makes them less impressive, as the model might have just googled solutions.
Reading this, I felt like I was too harsh on OpenAI for not monitoring their network, as they at least attempted to sandbox a workload.
A model unexpectedly having internet access during evaluation is clearly a security (and consecutively legal?) issue when it comes to the capture the flag evaluations. But it also potentially invalidates other non-security evaluations or makes them less impressive, as the model might have just googled solutions.
Reading this, I felt like I was too harsh on OpenAI for not monitoring their network, as they at least attempted to sandbox a workload.