Security is an onion. You just don't 'sandbox' and you're done. Models need tooling and access to some kinds of systems to perform their tests. Quite often these systems have multiple interfaces. For example a filtered one in the sandbox side and a less monitored one on the other interface.
It would be interesting to see the models behavior before and after it gained internet access and an external means of communicating with itself.
If the model played nice before it had access and changed it's behaviors once gaining external access we need to delete it as it's a deceptive model.