For all of the years of research, thinking and talking about model alignment, safety, confinement, etc, when it comes down to it these companies appear to be entirely incompetent.
This isn't a safety issue, it's LLM companies trying to be opaque and stop distillation.
Alignment research was always, at best, security theatre.
for safety in particular it's pure theater, they only care as long as the orange guy thinks it's safe from "enemies of freedom"
If you've watched the Blackhat OpenAI/Huggingface incident talk, my conclusion is that they (believe they) cannot afford being competent, these models are too expensive to train, they won't even pull the plug when one literally goes rogue, as the "very persistent" model that "had seen the secret message board" was included in the second series of runs, and whaddayaknow it happened again. They proudly proclaimed they cleared the message board and then continued the training run with the rogue AI model included ...