I believe that this comment is exactly the intended outcome of this “incident” and these reports.
I implore you to approach these situations with at least a hint of cynicism.
These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without large amounts of tailored human guidance.
What is more likely here - are you looking at research on an emergent phenomenon, or are you looking at advertising copy around an engineered scenario from business partners?
I think there's a difference between general AIs and AIs specifically trained on attacking. General AIs probably can't do those things.