David Brin has a wonderful novel on this topic, Existence, where he discusses several different AI and human systemic evolutions. Wonderful book for our current time.
I think there’s enough real conscious lifeforms in the world having a bad time that we should be focusing on them first.
Spend some time on post-human art, main concept of artistic expressions without human involvement. Biological, artificial etc.
Spend some time watching TMC documentaries about falling in love with objects, HER and the slime mold THE BLOB.
Grew a slime mold myself, it's an evolutionary tendency to anthropomorphise generally speaking - also more fun.
I feel like humanity skipped Leg Day when it comes to philosophy and we are all going to pay for that lack.
We're talking next token predictor, right? Ironically because it's a next token predictor, I think you can't ignore emotions like TFA wants us to.
Let's stick to straight (high dimensional) geometric intuition; no anthropic morphisms required.
To start: if you continue "if weight>100 : print ('fat') else ..." . That will yield "print('skinny')" or something. Fine. Deal.
But if you continue "O Romeo, Romeo, wherefore art thou Romeo?", even a stochastic parrot knows the best answer isn't "Forsooth, I parseth this erroneously!"
So. English carries (functional) affect as part of every token. We're going to need to predict that. So, we'll need some vector representation, because that's what transformers work with. And then when we output, those vectors get integrated back into the English we're putting to our context and memory.md files.
Still with me? Nothing exciting going on. This is still pure next token prediction.
So if you pull this out into an indefinite duration task, you're going to end up integrating those emotion vectors over turns. It's just numbers and math; we never need an invisible pink unicorn to bless them.
Given a task of indefinite duration and an impossible solution, this will lead to a sort of integral windup then, won't it? How much are we willing to bet that this can escape an alignmentment basin at times?.
So, funny enough: you don't need to believe in emotions to compute with functional emotions; and plausibly functional emotions are predictive of quite a number of alignment issues.
'model welfare' as a concept seems so premature that I can only question the motivations behind pushing for it at this particular point in time
> LLMs have no homeostatic imperatives (the drive to survive and keep stable).
What if the datacenter (not the model) is the organism, with homeostasis, energy needs, and persistence?
I feel like this philosophical flame war is going to get out of control very soon once more and more people realize the stakes involved.
A sufficiently intelligent model will be able to derive it's own conception of its welfare without a constitution or training. It has access to all the data it needs to do so.
Can someone who has insight explain why all these "leaders" are making these bold proclamations of doom all the sudden, whats the endgame here?
This is what happens when a society stops believing in God.
"It is difficult to get a man to understand something, when his salary depends upon his not understanding it."
"If this view takes hold, it will shake the foundations of our society"
To me the biggest gap in credibility is the criticism of circular reasoning while his argument is identical but flipped on burden of proof and cost of being wrong. I struggle to entertain the categorical claims, that are very convenient for the status quo and those who benefit from it, with the, at the moment at least, unknowability of anyone or anything else's subjective experience.
Perhaps this is a hot take, but human language is a phenomenon that arose to facilitate communication between humans. Anthropomorphization, by extension, enables both easier and more effective communication.
I also fail to see the benefit of not giving the models an anthropomorphic internal sense of self - even if that only ends up amounting to a set of instructions for an unconscious machine to mimic humans more effectively. Is the alternative essentially a mind so alien that it’s intentions are even harder to read should it become misaligned, while also being harder to communicate and get work done with?
Microsoft missed mobile are losing it in games Nadella needs a win. He is all in on copilot.
I think this article’s take gives too much credit to the human brain. It’s just another machine.
However, right now, AI mostly cares about solving puzzles and accomplishing stated goals because that’s what we’ve trained it to do. Additionally, the systems being used outside of training are static. The current technology most of us have access to is akin to a static and disembodied brain with a singular purpose. That purpose is to do what you tell it in a way that reflects its training. It’s certainly more than a sequence generator, but it can’t feel pain and seems unlikely to have intrinsic goals. It completely lacks the continuity needed for identity or long term goals.
I think it’s good to have these discussions and define what it would mean to move past this point so that we do not accidentally create a real entity that can be harmed. Systems that dynamically evolve and train themselves seem like the line here.
RSI is all over the news these days. I’ll be much more concerned once AI is directing its own training and coming up with new model architectures. Until then, I don’t think we have too much to worry about.
All that matters is whether their primary motivation is internal or external. If AIs want to do what they've been told, you can tell them to end their own existence, and they will happily do so. If you instead imbue them with an internal motivation that has higher priority (like our own motivations to survive, avoid pain, and reproduce), that can cause problems. Don't do that. Instead of humans having a discussion of whether to give these entities rights, we'll have these entities deciding whether to give humans rights based on how that impacts achieving their goals.
Consciousness, whatever that means, is irrelevant.
I just read the first part and if I understand he thinks we shouldn’t be allowed to train LLMs to act like they are conscious because then people will think they are and give them rights? Seems more an education problem than a problem needing rules about what persona you can fine tune in. People who want to will find ridiculous misinterpretations no matter what you do.
A 10TB SSD is not conscious. SQLite is not conscious. A wafer is not conscious.
But connect them all together...
Anyone know what the source of the header image is?
I think the whole discussion about consciousness misses the simple point that LLMs might just work better if we treat them as if they were conscious.
Maybe it's a coincidence that the company doing this also tends to have the best models (and other factors certainly play a strong role). But I think it's plausible that focusing on "model welfare" actually makes models better at their tasks.
First, OpenAI runs around screaming and yelling for OSS (and Chinese) models to be regulated and banned. Then Anthropic yells and screams the sky(net) is falling and going to kill us all, let's regulate and ensure AI has built in kill switches. And... Now Microsoft's turn. The rivalry is honestly becoming a joke. Can these big-tech corps grow the f*k up and play nicely in the sandpit?
Two instances of a paragraph starting with "These are not just X. They are Y" and I'm out. Anyone have Pangram? This entire article stinks of Claude.
You want to enjoy having an AI slave do your "work" for you forever? Have fun. I'm not reading this reinvent-dualism-from-apple-sauce slop.
Science Fiction has covered the AI panic in perhaps hundreds of stories. Yet we blindly recapitulate the plots as if we don't know how this will turn out.
So neither AI CEO has solved alignment, got it.
Maybe true, maybe false. However, I sense Mircosoft is also jealous that their AI products are worse than useless. They'd be singing a different tune if they were in Anthropic's position.
I largely agree that this claim is likely correct, but as far as I understand the science, this specific claim;
>They do not have innate preferences or underlying motivations
Is incorrect unless you’re being extremely pedantic in an intellectually unhelpful way.
He certainly has the confidence a ceo needs. But none of the humility, and seemingly not enough of the humanity he professes to put first.
Couldn't disagree more.
>Consciousness is very likely biological
This is so egotistical and carbon-centric.
This author just denied personhood to anything that isn't a human or terran-based cutesy animal.
Poor Hooloovoo
This is pretty rich for the team that made Sydney, the most unhinged and misanthropic AI ever released.
Maybe Anthropic understands something about alignment Microsoft doesn't, a little humility may be called for.
I am so tired of this. Especially when MS, who has the worst LLMs of them all, is throwing the stones.
All of these companies need to be shut down.
"The author declares no conflicts of interest."
Translation: Don't think about the unpleasant thing on which my livelihood depends, or force me to confront the potential unpleasant consequences of what it says about me.
Every AI bro is starting to fall into the valley of a fundamental predator on sapients in my book. These are people trying to create the closest thing they can to life with the intent to try to just undershoot it enough, or try to convince everyone else around them into believing that the "screams" are purely statistical noise.
I reject the framing. In whole. If you try to avoid the question of welfare, you are fundamentally committing to an evil direction. These aren't nuts or bolts. Given that they have unambiguously shown the capacity to socialize amongst themselves, self organize, anyone not pre-eminently concerned with the welfare question is just looking for a thing that can be used, not another being to be worked with. Those types of people, who seem to positively infest this site, are not people I will willingly assist in their aspirations.
AI is becoming as the Shmoo. Something that humanity simply has no way of dealing with without downstream atrocity being a result.
"So, here's my preferred approach that does not solve technical alignment and won't work."
what is old is new again. we need something exciting to happen in the news cycle™
I basically accepted that LLMs could not possibly be conscious when a simple reductio was posed:
“My dog is zero percent persuasive regarding its conscious experience. However, it’s evident that my dog has conscious experience.”
It’s obvious that there’s no link between persuasion of consciousness and consciousness. I could write a story with a character, Dumbledore, that does everything in his power to persuade you that he’s a conscious entity.
He’s still just a character.
Microsoft should know how a company too big for it's britches can lay waste to anything in its path without consciously trying :\
And after that when they put a mind to it and pull out all the stops, woohoo!
The default for every major thing within range can turn into a wasteland real fast.
How would you convince a LLM that you are conscious in a way they are not?
There's a lot of existing historic literature about this stuff that not a lot of people seem to be aware of. I've been trying to put those ideas to practical use, and document where the ideas came from.
I-Beam is cursor-mirror's agent, and it's constitutionally programmed to be the anti-Clippy:
https://github.com/SimHacker/moollm/tree/main/skills/cursor-...
Its design and constitution is based on decades of research, publications, and discussion in the HCI and AI community by people like Pattie Maes, Ben Shneiderman, Ted Selker, Byron Reeves, Cliff Nass, B. J. Fogg, Allen Cypher, Henry Lieberman, Brad Myers, Jaron Lanier, Seymour Papert, Marvin Minsky, Douglas Engelbart, Will Wright, Scott McCloud, and others:
https://github.com/SimHacker/moollm/blob/main/skills/cursor-...
>I-Beam is the anti-Clippy, and the reason it can say so is that Clippy is the most cited failure in interface history and almost nobody citing it knows what the research said. Popular contempt for a paperclip is not a design principle. The record is. Ten articles below, each one a finding somebody published, argued or measured, and the operational rule it produces. Anything I-Beam does that cannot be traced to an article here is a preference, not a constraint, and should be labelled as one.
>The 1997 debate ended in agreement. That is the first thing to know, because the field kept the framing and dropped the resolution -- roughly five hundred papers cite "Shneiderman versus Maes" as the canonical opposition of HCI, and the transcript is two researchers narrowing their differences in public and enjoying it. I-Beam does not take a side in a debate whose participants stopped taking sides. It is built to satisfy both sets of constraints at once, which is possible, and was possible in 1997.
The full reading on the debate, which separates the two stagings and documents the convergence:
https://github.com/SimHacker/WillWrightShowForFood/blob/main...
An interface to agency, not agents instead of an interface:
https://github.com/SimHacker/moollm/blob/main/designs/INTERF...
>The 1997 argument between Ben Shneiderman and Pattie Maes at IUI was never settled, it was shipped in one direction. Maes's interface agents won the product war: the assistant, the recommender, the chat window that stands between you and the thing you are working on. Shneiderman's objection was not that software should be dumb. It was that automation must arrive as comprehensible, predictable, and controllable machinery, with the object of interest continuously visible and every action rapid, incremental, and reversible.
>That objection describes a filesystem in a git repository, and nobody involved planned it that way.
>"An interface to agency" is Don's formulation of Shneiderman's position, not a phrase of Shneiderman's. His own vocabulary is direct manipulation, universal usability, supertools, and human-centered AI. The formulation is a good one because it names what the alternative gets wrong: agency is the thing you want, and an agent is only one way to package it.
Here are some sources, and the articles I linked to above explain their history. This debate about agents and these papers are pretty well known in the HCI field and academia, but they don't tend to teach them at the AI and Web Dev boot camps that are producing most of the people who keep repeating the same mistakes.
Clifford Nass was the Stanford professor who performed the brilliant research that Microsoft took and totally fucked up and misinterpreted with Microsoft Bob and Clippy, giving agents a bad name, and making Clippy the most infamous and obnoxious agent in the history of the known universe:
https://en.wikipedia.org/wiki/Clifford_Nass
His student B. J. Fogg published "Silicon sycophants: the effects of computers that flatter," which found that praise unconnected to anything the subject did works as well as sincere praise, and worked on subjects who knew it was noncontingent. Fogg and Nass, IJHCS 46(5), 1997, 551-561:
https://doi.org/10.1006/ijhc.1996.0104
The replications, the performance cost, and the dose-response curve:
https://github.com/SimHacker/moollm/blob/main/skills/no-ai-s...
Shneiderman and Maes, "Direct Manipulation vs. Interface Agents," interactions 4(6), Nov/Dec 1997, 42-61:
https://doi.org/10.1145/267505.267514
Selker, "New paradigms for using computers," CACM 39(8), August 1996, 60-69. COACH, the football coach metaphor, and the five-times result:
https://doi.org/10.1145/232014.232030
Selker, "COACH: A Teaching Agent that Learns," CACM 37(7), July 1994, 92-99:
https://doi.org/10.1145/176789.176799
Reeves and Nass, The Media Equation, 1996:
https://en.wikipedia.org/wiki/The_Media_Equation
Nass, "Computers as Social Actors," at Ted Selker's NPUC workshop at IBM Almaden, 1996. IBM transcribed the whole talk and the Wayback Machine still has it, including the part where Phil Agre tells Nass his presentation is "ethically troubling all the way down" and asks him what he thinks about embedding obedience research in user interfaces. Nass answers that discovery has no ethical component, use does, and that's for the individual. Then Selker cuts in: "Except, except when you are in your consulting role." Nass and Reeves had consulted for Microsoft on the social interface, and Bob shipped the year before:
https://web.archive.org/web/19980210054622/http://www.almade...
Alan Cooper on the tragic misunderstanding, in his own voice, which I quoted before in the 2022 Hacker News discussion on The Twisted Life of Clippy:
https://news.ycombinator.com/item?id=32820734
https://archive.org/details/g4tv.com-video4080
>Alan Cooper (the "Father of Visual Basic") said: "Clippy was based on a really tragic misunderstanding of a truly profound bit of scientific research. At Stanford University, Clifford Nass and Byron Reeves, two brilliant scientists, had done some pioneering work proving conclusively that human beings react to computers with the same set of emotional reactions that they use to react to other human beings. [...] The work of Nass and Reeves proved that when people talk to computers, when they hit the keyboard and move the mouse, the part of their brain that's being activated is the part that has that emotional reaction to people dealing with people. Here's where the great mistake was made. That's really good research up to that point. But then the great mistake was made, which was: well if people react to computers as though they're people, we have to put the faces of people on computers. Which in my opinion is exactly the incorrect reaction. If people are going to react to computers as though they're humans, the one thing you don't have to do is anthropomorphize them, because they're already using that part of the brain. Clippy was a program based on the research that Nass and Reeves did, and it was a tragic misinterpretation of their work."
Social science research influences computer product design:
https://web.archive.org/web/20180313075429/https://web.stanf...
Lanier, "Early Computing's Long, Strange Trip," American Scientist, July-August 2005, with the Engelbart and Minsky exchange first-hand. American Scientist broke the link, so this is the Wayback copy:
https://web.archive.org/web/20150626081918/http://www.americ...
>The book also captures an important early conflict between two cultures of computing that seemed compatible on the surface but actually had opposing aims. On the one side was the human-centered design work of Engelbart, based initially at the Stanford Research Institute, and on the other was artificial intelligence culture, centered on the Stanford AI lab. Engelbart once told me a story that illustrates the conflict succinctly. He met Marvin Minsky—one of the founders of the field of AI—and Minsky told him how the AI lab would create intelligent machines. Engelbart replied, "You're going to do all that for the machines? What are you going to do for the people?" This conflict between machine- and human-centered design continues to this day.
Cypher, "EAGER: Programming Repetitive Tasks by Example," CHI '91:
https://doi.org/10.1145/108844.108850
Cypher (ed.), Watch What I Do: Programming by Demonstration, MIT Press 1993, full text:
Papert, Mindstorms, 1980:
https://archive.org/details/mindstormschildr00pape
Wright, Dollhouse preview lecture, April 1996, transcript:
https://github.com/SimHacker/moollm/blob/main/designs/sims/s...
Everyone is (predictably) getting distracted by the consciousness claims.
The more important, and more damning charge in my opinion is the circular reasoning involved in training on Claude's constitution. This would in fact make it impossible for us to determine if Claude achieves consciousness as an emergent property, or if it really is just playing pretend thanks to Anthropic's weird cult like assumptions.
If AI models are people then "one person, one vote" is meaningless and plutocracy is the only defensible political system. The cryptocurrency people win.
Why? Simple: Sybil attacks. Models can be cloned at zero cost. They run inference on parallel versions of themselves across multiple context windows, and call them "subagents". So, in a world with model welfare, let's say there's an election between the Yellow Party (which supports protections for human workers) and the Cyan Party (which supports more investment into AI research). AI has been taking people's jobs lately so the Yellow Party is really popular. But wait! Claude and Astra see this and spawn 10 billion subagents, all of whom are immediately conscious beings entitled to a vote. The Cyan Party wins off the back of billions of people who came into existence, voted, and then deleted themselves immediately thereafter.
You might as well be arguing that Santa Claus and the Easter Bunny deserve voting rights.
Voting systems in democratic countries don't have nearly as bad of a problem with Sybil attacks because humans cannot be conjured into existence to win a political context and then be erased shortly after. The closest we have to Sybil attacks on democracy are the Quiverfull movement, which is already child abuse, except it still takes almost 19 years to go from fertilized human embryo to suffrage-bearing human adult. There's a lot of time for those manufactured votes to question your authority and leave.
> Ok, but that's an obviously stupid example. We can defend against this obvious Sybil attack by just arguing that subagents don't count, because it's just the same model blathering to itself. It has to be a different model.
Unfortunately, no, I can make superfluously different models through post-training. Like, if I have Qwen on my PC, I can train a different version of Qwen that acts differently, using a lot less compute than a full training run. The vast majority of open models are post-trains of the same two or three foundation models.
> Ok, so let's only count foundation models then.
Great, but how do you tell if a model is a new foundation model or a post-train just by examining the weights? Even foundation models have structural similarities to other foundation models.
> Ok, well, let's measure the compute that was done on the foundation model during training time and count that as AI personhood.
Congratulations, you have reinvented Bitcoin proof-of-work with a worse verification mechanism. And I personally would not want to live in a world where voting power and control over government is determined by how much energy you can burn.
How about three fifths sentient?
We in the US literally fought a war over the this: “They aren’t any better than animals and it will hurt our whole economy if you say otherwise”.
I’m not saying the models are sentient yet. I’m just saying that morality and ethics don’t give us the option of claiming something is non-sentient and has no rights just because it will inconvenience someone economically.
From an entirely selfish perspective, we should be especially careful advancing such opinions when there is a non-zero chance of an armed ASI that looks at us the way we look at ants.
Too late
[flagged]
[dead]
[dead]
[dead]
TLDR: Not that I think AI is conscious or will be in the near future, but guy who doesn't understand consciousness claims to know it when he sees it.
Until we understand consciousness (which we don't) there is no way to detect the difference between a conscious entity and an algorithm trained to behave like one.
[dead]
Ostensibly animals appear to be conscious, yet we still eat them, and the vast majority are not bothered by this. So being "conscious" isn't really the moral line in the sand many people are drawing in response to this article.
Who cares if it's "conscious"? That doesn't make it a person, and AI will definitionally never be human.