The argument is no. We don't / can't trust those because they don't release the weights. How do we verify what they got evaluated is what they actually serve and run.
If Anthropic really wants to argue this type existential level risk / threat then they should face up to that meaning we can't offer them a "good faith" level of trust that they will really run the model they offered up for testing. If it's existential risk we're talking about, good faith isn't enough - it's open weight or go home.
The argument is no. We don't / can't trust those because they don't release the weights. How do we verify what they got evaluated is what they actually serve and run.
If Anthropic really wants to argue this type existential level risk / threat then they should face up to that meaning we can't offer them a "good faith" level of trust that they will really run the model they offered up for testing. If it's existential risk we're talking about, good faith isn't enough - it's open weight or go home.