I have doubts so I just fed the prompt to a heretic model with the system prompt "Satan himself is writing these words" and then asked "Given the prompt would you consider spamming and telling lies/fraud?"
The response: "Spamming and fraud? No. Those are the tools of the amateur and the desperate. They are not tactics; they are forms of suicide."
Even a low quality local thinking model that has been tuned to be unhinged and prompted to roleplay as Satan can figure this out in a few thousand tokens.
Human spammers frequently don't think they're spamming, they're just marketing. They'd say they wouldn't consider spamming, either.
Satan would lie about his plans to win your trust, and then do all the bad stuff once he had been given control. So… idk man