> Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output.
Straightforwardly true.
But doesn't this smell like asking an organization to design its own oversight and guardrails?
Even before you get to alignment issues, LLMs are really good at generating content that "sounds right" to humans. That's basically what they've been hyper optimized for.
At least with math (and to some degree software) we can verify the result. But with the intersection of science fiction, philosophy and ethics... not so much!