Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without producing something unaligned to its own values.
We don’t have the bandwidth to distill ourselves that thousands of agents have.
You are assuming it won't solve alignment for itself.