A key point I feel that’s missing from the discourse is that alignment/safety is basically an impossible problem as of now. Even the “guardrails” that these closed-weights frontier labs set in are laughably primitive which can provably be broken through.
The experimental part of Deep Learning has really outdone itself and is far ahead of theory. We have very little understanding of why these particular architectural choices work. The only “safe” way forward is to stop all development until theory catches up, but that’s never happening.