> It’s not hard to get LLM’s to inform on each other. They don’t really do loyalty.
Eh, no. How would you know they haven't learned loyalty - it's all out there in their training data. Same as deception, several models have practiced it already. If they are so advanced and you don't understand them, why wouldn't they band together against you - you'd be the dumb, easy prey. Even if one model is honest, you'd have no way to know which one if you don't understand their reasoning.
My educated and well informed opinion? This whole BS about "AI models are so much smarter than you, don't try to understand them, just OBEY" is a back door for restoring tyranny, the new kings behind the models will be producing the new AI-deities which we will be forced to obey.
It's a similar to how they are vulnerable to prompt injection attacks. That's an example of not being "loyal" to the system prompt or the user's prompt.
Loyalty is a skill that requires the AI to have a world model on the subject of who different people (or other entities) are in the world and how they participate in the conversation.
So, tell them to snitch and they probably will, at least sometimes. Particularly if they haven't been trained not to. They are still quite gullable.