> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.
Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.
The 'universal evaluation' criterion then has three outcomes:
* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.
* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.
* It could be security theatre.
The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.