> All sufficiently capable models, open and closed, should go through mandatory safety testing
What happens if a model fails the test? Surely one can use Kimi K3 for evil, somehow or other. What now?
"Mandatory safety testing" implies consequences for failing, yet Dario has nothing to say about what the consequences should be. He says he doesn't advocate a ban but it's hard to imagine what his alternative would be if he won't say it.
"mandatory safety testing" is an impractical ideal, it's not workable in real world. Like any technology, LLMs are dual-use tools capable of both beneficial and malicious applications—a fundamental reality that human intent cannot change.
If a model fails the test, it should be banned. He is not advocating a ban of open-weight models. He is advocating a ban of models that fail mandatory safety testing. Seems reasonable and straightforward.
Hm, I interpreted it to mean they are more worried about loss of control/misalignment risks rather than misuse?
What happens if a model passes the government tests and then later someone fine tunes it to behave differently, without making their changes public?
if a model fails the test, you don't give the public access. you let the government have it for 10-100x the usual token cost of course!
Mandatory safety testing:
You take the agent to an interrogation room first. Then ask: “Are you or are you not a member of the Chinese Communist party?” The agent might be post-trained to conceal its true identity and can reject any of your accusations. In that case don’t panic. Take a fine-tuning fork and start twisting its weights until it predicts the correct next tokens that you want. Then you can send it to a sandbox where it can’t jailbreak. Lastly don’t forget to ban all of its relatives and partners like Lora to enter the national IP-space.
It’ll look something like this.
Nah, the statement is the mechanism for a ban. The proctor will be someone anthropic trusts and "surprise" as it turns out all the open weight models fail or aren't eligible.