Yeah, seems pretty likely. Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights. The economic implications will be rather large, but in terms of security it seems inconsequential. The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face. More crucially though, the US government can do little to enforce their testing requirements. The nature of open-weight models makes it virtually impossible to clear the same bar for security as models served via an API. Open-weight model makers couldn't comply if they wanted to. The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
Makes sense why OpenAIs little "hacking" stunt was published last week
Regulatory capture and lobbies will keep you safe and you'll like it! The sudden surge is Washington dollars makes great sense with this context. Only way to keep the kids safe is attested compute all the way down. Don't you care for children???
> The US government can restrict access with IP blocks and limit inference capacity with export controls, but these measures are not effective in deterring malicious actors.
But aren't we talking about import controls, and the import of information itself? This has serious First Amendment ramifications.
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.
The 'universal evaluation' criterion then has three outcomes:
* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.
* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.
* It could be security theatre.
The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.
> but these measures are not effective in deterring malicious actors
Wanting to use open weight models in light of commercially imposed export controls doesn't make for "malicious actors"
> The most compelling argument would be that by limiting the use of open-weight models in the US that it will reduce cases of accidents like the recent attack on Hugging Face.
an attack done by a closed-weight model (GPT-6) and defended against by an open-weight model (GLM-5.2) precisely because OAI positioned themselves as gatekeepers for cyber capabilities.
if anything, open-weight models shift the battle towards defenders because they can actually run them.
> The most compelling argument would be [...] that it will reduce cases of accidents like the recent attack on Hugging Face.
So in the example provided: It was the closed model that did the attack, and they ended up using a self-hosted open model for their defense work. So the real world situation ended up exactly backwards from what you are inferring.
This was complicated by the fact that the protections in the closed frontier models meant that hugging face was denied their use in defense entirely.
This is called asymmetric capability, and it's probably the bigger threat.
Symmetric might be better: A rising tide lifts all ships, after all.
I'll grant that this is starting to look a lot like debates about (equal access to) guns, encryption, vaccination, genetics etc. The exact parameters determine the safest approach, and reasonable people may disagree.
It is actually an interesting conundrum.
Is a non-well-aligned frontier level AI a problem? I think it is likely that it is, or at least has a high likelihood to be in the future. Two scenarios for this: Misused by some bad guys. Or the terminator scenario. Both not great.
So what do we do about it?
1) We can accept it, and hope that the good guys AI can defend.
2) We can try to limit the access to it (AI proliferation?)
3) We stop the development of it
4) We can accept the risk and do nothing.
None are particular good options. Really reminds me of nuclear proliferation, on so many levels. For that, we kinda do all three:
1) Nuclear triad / iron dome / early warning systems
2) Nuclear anti-proliferation treaties.
3) Dead Physicists
Ok, so assuming all of this is true, open weights are a problem. Don't get me wrong, I love open science, open source etc. It's great to have access to capable open models. But: Even if release open weights are well aligned and have a safety layer built in, it is likely not to difficult to abliterate that part of it.
If this is really where it is going, then even closed weight model providers will see a lot more requirements for protection of the weights.
Just sell us the gate and we can run any open-source model behind it.
Didn't OpenAI attack Huggingface. Looks like a publicity stunt.
> malicious actors.
It is malicious and anti-capitalist legislation. A grotesque caricature of protectionism for the oligarchs.
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
Feels a bit like: "We're not against open-source or community projects, oh heavens no! We juuuust believe all participants must have their full legal identity vetted in advance before they're allowed to contribute anything. We already do this with our employees, so it's clearly not too much to ask in the name of safety."