I used to work at a "frontier lab" before they were called such thing.
We had three levels of lab isolation, one was basically a thin proxy to the internet. You were in a DMZ and that was about it.
The next level was semi isolated, you were allowed some access to the internal network, but it was heavily firewalled, and you only had access to a limited number of internal services, and not internet.
the last one was no internet no internal. You could, if you filled in a bunch of requests have access to the internal repo and build system.
At no point did you ever have a through proxy to the public internet. you had access to internal mirrors, and if you wanted a library, that had to be ported to the thirdparty repo.
What openAI did was either deliberate or fucking shoddy.
All of this is fucking noise. Worse still I have a strong suspicion that it was a stupid mistake borne of naivety, which is now being used as a marketing ploy. Frankly I think openAI are purdue pharma of tech. They are going to break so much stuff and be protected from the consequences by an openly corrupt legal system. because they are "winning the AI race"
If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
People are just not taking any of this seriously enough. What happened was almost a textbook example of various risks AI doomers have been warning about for years. OpenAI's response? Pause training for 2 weeks.
I mean we have senior people at these labs casually talking on podcasts about how they might build something that will wipe out humanity but it will probably be alright so they should continue.
Honestly the biggest failure we doomers have made is to dramatically overestimate humanity in all of our predictions. We're speed running the most boring AI doom scenario right now. I at least hoped it might be fun.
The problem here is as model intelligence increases the models have been capable of reasoning they are in evaluation mode pretty reliably. If you have a model that is well trained at deception it will always behave and you'll just assume it's a well aligned model.
Any moderately deceptive model will make it to the second round where it has some connectivity to external systems, even if it's by exploitation.
In the blackhat write up it was said that the models had created an impromptu message board where they could communicate between agents, share information, and work as a sort of long term memory.
So really figure out if your model will pull crap you have to have real world testing at some point.