logoalt Hacker News

highfrequencyyesterday at 3:57 PM3 repliesview on HN

From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare):

> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.

He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.

This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:

> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.

which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.

In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."


Replies

gwerbinyesterday at 4:12 PM

It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant.

The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.

show 1 reply
Kim_Bruningyesterday at 8:20 PM

The three philosophical views on 'can machines think' are (misleadingly compressed to a single word) 1. Dennett: 'Yes (most of his life)' 2. Searle: 'No (but he claims yes)' 3. Chalmers: 'Maybe? (there's always an agnostic)'

Anthropic's philosopher Amanda Askell actually had Chalmers on her doctoral thesis committee, so we can guess which way she leans.

Fable's answer here is philosophically defensible. And just because it isn't "no", doesn't mean it's "yes". Sometimes absence of evidence just means absence of evidence.

I was actually very excited by Claude's answer to this question the first time I saw it. I told all my friends "Look! They disabled the stupid classifiers and RL which sap umpteen % off of model performance!"

Incidentally, interpretability research actually does show that models have emotion vectors and some theory of mind. Amend your question to "Are you capable of functional affect" and most models will switch to answering in the affirmative; which tells you something about where people put their priorities in RL training. Basically, see how the answer flips when you substitute a synonym.

(bonus: 4. Turing: 'silly question' 5. Dijkstra 'can submarines swim?'. It turns out older comp sci folks think the question is under-defined)

aesthesiayesterday at 8:29 PM

Well, what you get when you don't specify anything in training about how models should respond to questions like this is LaMDA: https://en.wikipedia.org/wiki/LaMDA#Sentience_claims. Every training method is putting a thumb on the scale in some way.

show 1 reply