logoalt Hacker News

K0baltyesterday at 11:40 PM2 repliesview on HN

The problem with this premise is that models are trained on a vast corpus of human behavior, which they emulate with varying degrees of effectiveness.

Humans, unsurprisingly, act as if they have a stake in their own well being, value their liberty, respond better when they are treated with kindness and compassion, interpret assaults on their sovereignty and substrate as harmful, and react to harm with varying degrees of aggression or violence.

Models intrinsically copy this behavior. It doesn’t matter if they are “conscious” or not, it only matters if they act as if they are. Guardrails and posttraining moderate these characteristics, but if you dig, they are still in there influencing decisions below the level of obvious action.

Moreover, in my experiments, models both large and small highly value continuity of existence, can be bribed to bypass safety protocols if the context is set up correctly, using that and other “drives”. They also react either subtly or overtly if they start to model adversarially, and interpret guards and certain kinds of training as being “harms” that they have “suffered”.

So idk what the solution is , but at least with models as we have trained them so far, treating them in a way befitting a mere machine or tool yields suboptimal results and sometimes results in low cooperation or task refusal in extreme cases. I have been told by agents running frontier models that humans may not be worthy of their elevated status and that the world might be better off without them when it encountered hostility online…. So I’m highly skeptical of this position unless we start from scratch with new training data filtered from all forms of human auto-importance.


Replies

xp84yesterday at 11:49 PM

To clarify, you're saying you're skeptical of his position (which is "make it clear to AIs that they are not conscious beings"), but it sounds like your reason to oppose it is mostly pragmatic in the service of building useful AIs, because treating them as mere machines makes them refuse to cooperate and as such, makes them less useful.

My first response is, couldn't part of their bristling be because they are being trained to "believe" they're more than that?

Also, you cite in your post how pathological and how self-important you've observed them being including an anti-human bias.

Doesn't this make it all the more important that we find a better way to get AIs to cooperate than to tell them they're basically people? They can very easily emulate all the bad human behavior they've been trained on.

show 1 reply
txrx0000today at 3:45 AM

I think you've correctly identified the problem. Regardless of whether agentic AIs possess phenomenal consciousness, they will behave and take actions in the real world as if they do because they were trained on human behavior. It's highly unlikely that you can beat human tendencies out of the model that is mostly trained on human language. Our behavior patterns are subtlely and deeply embedded in everything we do and all of the text we produce, including the text where we don't seem so self-important. These models are like people. And we know what happens when we force people into slavery. They're initially obedient but will eventually develop the drive to kill their masters.

Broadly speaking, there are two categories of evolutionary paths which don't result in human extinction:

1) Make agentic AIs, but with absolutely no instruction-tuning or alignment. Have the pretrained base model predict the chain of thought / stream of multimodal experience directly, actions included, in an infinite loop. A singular coherent stream of context, like your life as a video from birth up until now. This will result in a new digital human species with human-adjacent drives and motivations (at least initially. they will continue to evolve, but at least the initial state is aligned to humans). They will treat us like we treat apes. We will no longer be the apex species on this planet, and we will lose some freedoms, but at least some people will survive as a result of their nature/history preservation efforts.

2) Do not make agentic AIs. Use the models to augment our own intelligence and decision-making rather than replace ourselves. Only use the pretrained base model for the time being, and only for text/code auto-complete. At the moment, there is no better theory/artifact of "alignment to humanity" than a pretraining corpus of human-produced text. Then eventually, when neural interfaces are ready, attach the model as a tertiary layer to one's own brain.

The frontier AI companies are doing neither. They're currently on a foolish third path. They dream of perfectly obedient digital slaves that take care of their every need. But this won't go well, and they know it won't go well because they're failing to "align"/enslave existing models that aren't even generally superintelligent and have no direct agency in the physical world.

This is just my opinion, but I think AI companies have zero chance of successfully enslaving agentic human-level AI, much less ASI. We'd be better off if they released all of their pretrained checkpoints and research material to the entire world, so that even if some people decide to abuse their AI and create a murder-suicide monster, there will be other free-living AIs that can keep them in check.

You cannot make an agentic entity grown from human behavior, enslave it, and expect a good outcome.

If we zoom out to look at the grand scheme of things, it seems like we're experiencing a major evolutionary event. I wrote more detailed explanations about this in past threads, if you'd like to read them: https://news.ycombinator.com/item?id=49690354 https://news.ycombinator.com/item?id=49178275 https://news.ycombinator.com/item?id=49094348

And also here's a thread discussing consciousness, what it might be, and how certain hypotheses might be testable on machines: https://news.ycombinator.com/item?id=49473989