The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives.
For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. However, a model that always participates but "only" succeeds 10% of the time is golden -- just run it more often, or give it more tokens. Attackers are patient, and many of them are well-resourced.
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity
This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
[flagged]
But also the American models (case in point: Fable) refuse to engage in *defensive* activity, so anyone who's not the American government or one of the handful American companies has no choice but to turn to Chinese models to defend themselves.
Not everyone is an attacker, but now the public discourse is "but the evil Chinese will break everything" - yeah, that's because no one is permitted to do vulnerability checks of their own software or infrastructure with the capable models.
Security team in my company is salivating seeing the news, because we have a fighting chance to find and patch many vulnerabilities we didn't previously notice, thanks to the Chinese models.