I have this idea of using an obliterated version of this model for cyber work(or even this one, seeing that its guardrails aren't that strong) in a harness with the ability to spawn SOTA level subagents, faster and more capable.
The rationale is that the manager model sees the big picture and knows that the task is "unethical" while sota models are just given very isolated technical tasks that don't trigger any refusals.
Has anyone tried this? I would love to know about previous attempts of this approach.
Yeah, the Chinese government used the same method last year to hack the US government using Claude Code.
Making each piece of work small enough to be plausible. Compartmentalization.
(Also saying "nah it's cool I have permission", heh)
https://www.anthropic.com/news/disrupting-AI-espionage