>such an RSI-capable agent must _ALWAYS_ be scheming/plotting/hiding its true strength in _ALL_ of its prompts/tests.
From my POV you're over-focusing on a very specific failure story and neglecting a broader swath of possible failure scenarios.
>Is there a hole in my "alignment problem/solve mechanistic interpretability" argument?
The notion of telling an AI which may not, itself, be aligned to solve the alignment problem seems a little dicey.
1. Fair, the story I'm responding to is the senario in If Anyone Builds It, which I assume is Yudkowsky's best/most persuasive argument (else why make it the ONLY scenario in the book.) I'm willing to entertain other failure senarios/arguments, but honestly I'm tired and would like you to propose them yourself instead of having me dream up your arguments for you.
2. 100%. Again, I'm no accelerationist: I have no faith in alignment/mech-interp ever being solved. Anyone saying they know the probability of alignment is lying. My point is that pdoom after RSI is _high variance_. Pdoom pre-RSI is zilch.