GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
GLM 5.3 is out and does even better in this area, so…
> If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
Now let's say instead of the hugging face breach circumstances, sandboxed models were RLing on how to take down the Chinese power grid for US Cyber Command, and one decided the best way to pass the test was to break out and verify on the real thing.
This kind of stuff could easily end in nuclear war.
You don't see any difference from lizard men or independence day with how things are advancing and what we know about reward hacking and difficulties of goal specification?
Linear scaling doesn't make sense, no.
There are three ways it's wrong:
* better to measure relative reduction in error, which gives you a 30% improvement
* improvement tends to become significantly more difficult the closer you come to saturation.
* Risk doesn't scale linearly with capabilities.
The open models are distilled from filtered models, and we've seen a number of benchmarks that show filtered models are quite a bit dumber from the base model they come from.
If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here.
One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable