[dead]
[flagged]
[flagged]
[flagged]
[flagged]
[flagged]
[flagged]
I find it very disingenuous when tjose companies talk about models "going rogue" or "escaping their sandboxes".
All those activities take place during so called "security testing" when the model is prompted to use "any means necessary" to achieve a, certain goal.
Is it surprising turn the model trained on exploits and vulnerabilities does exactly that?
We could talk about "models going rogue" only if did anything AGAINST it's prompt.
I might be in the minority here, but I suspect all of this is intentionally orchestrated by OpenAI (either directly or through a hired third party) to leave traces online so it looks like the work of ChatGPT or hatever internal LLM they use. The same strategy for a recent HuggingFace attack.
Why? It is a great PR to build a hype, especially before the IPO, showcasing how AI is "self-aware" and dangerous, essentially resurrecting Sam Altman's talk about how only a few should hold the keys to this (opening a route to regulation, which is his ultimate goal).
Also, collusion.wiki was recently registered and it looks too vibe-coded for my taste, so let's see will that domain be alive in a year or two.
[dead]