I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it?
It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be.
That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.