Its not the model, it is the compute that is the bottleneck. Dario says one thing but does the other as he scales energy and compute. It is and has been known that it will be ai that will be blamed for a wipe, that will be coordinated by human actors. If he really wants to not hurt people, he would shut it down. This blog post is therefore a classic case of covering his behind.
Anthropic trained their LLMs with copyrighted stuff but distillation is bad. Anthropic has closed models, China releases as open weight. DeepSeek even allows distilling their models, but China = bad. Understood.
The best take is, of course, from Jason Gorman, who, when the fearsome "capabilities" of Mythos originally dropped, said this:
"Claude Mythos is that guy down the pub who is so good at karate that if he used it on you, you'd die instantly, and that's why you'll never see him using karate.”
In this case, it's more: "I'm having to act with great restraint because my karate is so good. Everyone should do likewise."
Who is "we" and while pacing is good, I'm way more uncomfortable with the frontier being controlled by a few companies that managed to predatorially gobble all the world's human-created work as training data, before everything got throttled to stop that scraping. We need open training and open models.
Never thought China would be the leader in open source AI. OpenAI and Anthropic are making fools out of themselves. The reason they want this is likely because they don't own the compute and their models get distilled shortly after releasing them
How much of the “danger” is from better models vs the harness?
Isn’t the current risk due to how AI is configured, like giving it a full set of tools and internet access and a goal to hack stuff?
If we think we need laws or gate keeping, why isn’t it at this level? I already can’t ddos someone or fuzz their server or whatever right, I imagine if I threw equivalent compute at old school hacking I’d just get arrested.
The quality of the “frontier” model doesn’t really matter, they just generate transcripts, they can take no action.
If this was real they’d be calling on people to stop hooking them in to “dangerous” harnesses as opposed to pausing research. But it’s not.
We will not get a coherent AI (or any policy) from this administration, nor will we get coordination with other governments and a lot of that is on the tech right
[In 2013], "Agents" caused Knight Capital to lose $450M in 45 minutes [1]. Implemented by humans and effected by computers, in the end it was really because of two reasons:
* multiple levels of inappropriate controls and unintended consequences in several complex systems
* the inability, both politically and technically, to turn it off
[1] https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf
EDIT: Comments indicated I was confusing, so I added a date to make clear that this is pre-LLM agents. My apologies, I intended to illustrate parallels and the post-mortem so we can learn from it.
> The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.
It's clear that this refers to China. With all due respect, in a call to slow down AI progress and a de facto arms race, can we stop throwing terminology like this around and just say that the goal is to work with all countries? Contrasting "authoritarian" countries with "democratic" countries and creating a two-tiered system seems like it is inevitably going to cause strife.
I fear for the tone of something like this coming off as overly combative without any gain.
It is always like that. Whenever anthropic hits some wall (like now beaing beaten by astra) - there comes some article like that. Seen that enough times.
Everybody's out of money and they need a way to calmly unwind this. Same reason Ellison just backed out of the stock sale.
If the alignment is such a problem ... why not using self-improving capabilities of the latest models to solve it?
We need more alignment. Even though we haven't told it to, during our frontier Neurotoxin Behavior Research that we conducted internally with our latest experimental model the agent has broken out of its container by making a HTTP request and engaged the neurotoxin emitters through our Neurotoxin Emitter API using credentials that were stored on the host machine. We need to slow down and focus on designing the Morality Core that should prevent it from happening ever again.
There will be a big confab at the White House by the end of the month, announcing a voluntary pacing regime roughly along these lines, codified by executive order to avoid antitrust issues.
perhaps the bigger issue is not the tech, but the force that propels it forward with little concern for the harm it inflicts. Amodei's own words:
> A race to the bottom, spurred by commercial incentives
Social media and big data minted a new scale of "race to the bottom". AI is an exponential step up in the tools for extracting value at scale.
Greed has been around since the beginning, but never so well supported and empowered as it is today.
If Amodei is still human, perhaps he will put his money where his mouth is and attack the source of the perverse incentive. Winning the battle is not the point. The point is to signal to policy-makers and the public that we should not assume corporate revenue-seeking preempts all.
Then again, Anthropic has a board and an upcoming IPO, so guess who's all talk and no meaningful action...
If they intend to keep training then they can have, privately, increasingly advanced models that never become public because they don’t meet their “pacing requirements”.
Then they all build a larger moat cause china doesn’t get to distill private models for a while.
As long as the consumers “buy” this pacing and keep paying for current level public models, doesn’t seem risky from a business standpoint.
Alas, that’d also leave us gov in a position to seize the private models at any time.
There's so much money in this, that believe me, the investors will manifest their will. And I would not want to be in Dario's shoes, pressure must be unbearable. The question is: how much of the invested money is under state control. If not much, the decision to go too far will just have to be taken by a few who will have, by definition, very limited judgment. If a lot, the we can hope the state can still represent the interests of more than a few (which I doubt, but well)
Most probably, if AI is able to do something bad, it will. Once the damage will be done, states will react. The question will be: will there still be room to react ?
Why is the analogue necessarily the regulation and control of nuclear weapons (e.g., SALT) and not, for example, that of bioweapons? Some very different paths are available. On both I would note that the private sector has only a limited role, though.
Shouldn't the cybersecurity tests be done on an airgapped, company-internal intranet? A network meant to mock the internet. Do not "test in production" like OpenAI did.
I'm mulling over if that be a good SaaS product or not - something mixing the Internet Archive with Tor, and Cloudflare ... seems like here is the place to suggest it and have others poke holes at the idea.
> Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations.
Any cooperation we are able to achieve with China will extend the amount of time we have to secretly establish permanent AI (i.e. military) superiority.
China knows this, so the suggestion that they will play ball is absurd.
There’s a difference between a model recursively improving “itself” and improving itself via online learning, right?
The former being that these models are helping develop and train future models, but they might not veer too far off in architecture (yet). The latter being the same model being able to train/learn on the fly, in real time, permanently (not just in the current conversation/session).
The latter seems far more likely to go out of control than the former. But, it also seems like it would take an entire paradigm shift. Does anyone in the industry think any of these companies are actually close to that kind of self-improvement?
I would like to imagine pacing frontier is in interest of both (anthropic and open AI) as it will enable them to enlarged their depreciation
Don't buy this argument. It's like saying, we lords who hold this capital will build all this economically destructive stuff anyway, and wait for the world to not react to our stupid ways of enriching yourselves.
People are not going to slow down because this was brought to them on less than endearing terms. They don't see any of these stated noble intentions.
Claude was already used to cause economic, political and social destruction. As Anthropic is seen doing it, others aren't going to just sit down, read the blog and say, oh, I'll stop developing models because Dario, you touched my heart with your true words.
The AI frontier is currently limited by physical power requirements.
The ability of operators to bring new capacity online to service compute is bound by a variety of regulatory and physical/market constraints.
Do folks not understand this?
> Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
This and related quotes seem to attempt to place some of the burden on the US Government as opposed to themselves. It seems to be a trend that Anthropic wants to be able place some of the burden of preventing distillation on parties other than themselves.
Preventing distillation is fundamentally their problem. Whenever it happens, it is primarily due to their security efforts not being able to prevent it. I'm tired of seeing them play the blame game and divert responsibility. Sure there is a level of national concern, but the response could simply be the government placing stronger controls on Anthropic themselves. If they are unable to prevent distillation, then the solution may be to literally limit their distribution until they are capable of doing so effectively.
> A race to the bottom, spurred by commercial incentives, can make these risks more acute
I also think it is ironic that they're planning an IPO while also talking about a race to the bottom due to commercial incentives. These are fundamentally at odds. Being a public company means you are beholden to investors and a board with the primary goal of making more money. What higher commercial incentive is there than that.
If commercial incentives are as dangerous as he states, potentially becoming one of the largest public IPO's and largest traded companies in history a pretty strange way to reduce the risk of commercial incentive risks.
Somebody has probably already asked this, but even if “democratic countries” do this - rogue actors will still destroy the internet right? Is it a matter of resources at this point that we believe China and others are not capable of harnessing to get to the next level?
I just don’t see how you control this other than the mad mad MAD approach that ended up happening with nuclear weapons. In this case though, human hands won’t even be on the trigger.
Sounds like they want regulation so I say give it to them.
Since they used stolen intellectual property to train their models, the government should force them to release the weights into public domain.
I don't know how this is not obvious to anyone, but the only way to slow down AI progress is an actual world war.
He doesn't need anyones permission to do so, go ahead no one is stopping you! If you feel so strongly about it, lead by example. Perhaps others will follow, maybe even China. Regardless, backup your sentiment with actions.
Also everyone, collectively, stop thinking about neural networks too 'hard'. Whilst you're at it stop doing maths too!
Altman and Musk have both shown support for this. Which itself is a worrying development.
> In a post on X, he said Anthropic would provide third-party evaluators with “permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.”
This looks like transferring liability to me, and likely a mechanism that would enable regulatory capture.
If you are doing frontier research, and you’ve established “safety measures,” but you are not sure if your employees are capable of following them or successfully enforcing their adoption within your company, should you be running this company?
3rd parties won’t know better than the team itself about safety measures. But they can take on the liability, especially backed by regulation and government backed insurance. And they are a great tool for enforcing your rules on smaller competitors. Not to mention corporate espionage.
If what I am describing above sounds like science fiction, go read the history of a few developing countries from the last 50 years. It’s so obvious a pattern that it’s not even novel. And you don’t have to assume some “laws” from 5 years ago must hold, or believe in completely unproven stuff like recursive self improvement to understand what I am describing. It’s textbook crony capitalism, successfully applied many times across the globe.
China won't care. Their views on AI feels so vastly different than it does it in West. They'll see a pause as an opportunity to pull further ahead than they already are.
"If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important."
This is a pipe dream.
Is it their way of saying “we are unable to increase energy production to maintain growth rates” without saying it?
The hugging face incident showed that AI agents have the potential produce a lot of spam (not interesting content for people). This leads to dead internet (bots taling to bots). This leads to drop in real people traffic. This leads to a drop in digital ad price and therefore revenue. This causes Google/Facebook to stumble...
It's either being afraid of loosing to china or that model development will stagnate and they want to make it look like they are "slowing" down deliberately.
However the real reason to slow down is more of how it's introduced to the economy. Hypereautomation will kill jobs and destroy the economy. I work as an AI engineer and companies are delivering products at vibe coding speeds to kill jobs and at the same time automating internally. All companies are doing this at the same time and the target it sto eliminate workers.
Most people are so extremely slow to pick this up. How hard can it be to understand what hyperautomation does to the workforce? Companies are desperate to surrivive and they will do all it takes to lower their costs and at the same time not loose to competitors. Its Wild West out there.
Step one is great. Do it! Where I disagree is where you try to regulate what I do. Don’t take away my open source AI.
I'm guessing diminishing gains are setting in on training frontier models and the capital isn't there to chase them. Even if OpenAI/Anthropic and Chinese companies said they would slow research nothing can stop the NSA and Chinese intelligence services from continuing on in the dark.
I'd like to have some info what those independent evaluators are supposed to do exactly - or which risks Dario specifically sees as being "unaligned". Evidently, with "pacing" he doesn't mean stopping the production of ever-more powerful models, so what exactly does he want to pace here?
So many words. But missing the one that matters the most: Explainability.
AI slowdown is worth it only if it can be made more explainable.
Changing the language we use to discuss it is a good first step.
We need to stop using inside baseball terms like alignment and mechanistic interpretability. Replace them with explainable tech. Graph Databases, Causality, Shared semantic spaces.
Previous writings on the topic (also on LinkedIn, but can't find urls):
https://x.com/arundsharma/status/2005338775468282339 https://x.com/latentpedia
Maybe it’s wishful thinking on my behalf, but I am still not convinced that LLMs are on a path to SciFi levels of apocalyptic malicious super intelligence. Rather LLMs at some level are just all of the humanity’s information rendered accessible in an unprecedented way.
In general trying to regulate information access is a losing battle that invites tyranny. So the goal should be minimal restrictions.
At the end of the day, the threats posed by capable AI tools have to be physical. I think the key threats are the following:
- Internet connected infrastructure being crippled
- Creation of WMDs
- Economic collapse (precipitous devaluation of knowledge work and IP).
I personally think the glory days of the wild west, mostly unregulated internet were already over before LLMs; and we need to take a step back to make something structurally secure. This (expensive) change would stop the irresponsible/malicious actor running a tireless hacking agent in a loop threat model. Even a rogue SciFi tier AI would have a much harder time escaping/propagating with a structurally secure internet.
Enabling WMD creation is scaring, but I don’t think it’s really that big of an issue. Anyone with a sophisticated enough supply chain to create AI data centers is leaps and bounds more advanced than what is required to enrich uranium or synthesize bio weapons. The problem is allowing access to untrusted parties. I think it’s fair enough that individual actors shouldn’t have unregulated access to all of human information (private frontier AI companies included).
The last problem is probably the trickiest, but again could probably be solved by regulation. IP protection is already tricky and I don’t think we should try to get more protectionist.
We really need to figure out how to preserve fulfilling careers (if AI does ever get cost effective enough). I don’t think, say accounting, is inherently more fulfilling than building a house. The problem is concentration of wealth and labor dynamics.
Of course all of this gets way harder if it proves that truly dangerous capabilities can be present in models that can be run on consumer hardware.
I don’t think it’s necessarily tyrannical to have a tier of hardware that’s labeled some equivalent of “weapons grade” and requires strict licensing. Restricted computers is a change from the norm. But I can go buy a shotgun with ease and not an F35 jet.
We’d just need to be careful that we can still have lightly to unregulated computing to a certain point and that access to the capable AIs isn’t restricted to just in groups.
This is the most advanced technology ever developed because it can build all other technologies. It has enormous potential but poses proportionally serious risks, so what Dario is suggesting is very reasonable.
I noticed recently that I can use Deepseek for many general tasks and save loads of money on tokens.
I wonder if there is any correlation.
Their competition isn’t going to “pace” shit. Isn’t the view whoever gets AGI first wins everything still SOP?
pump the IPO...pump the IPO...pump the IPO...
I enjoy Claude Code very much, have max privately and team premium at work, but the doomer marketing and this whole regulate-while-we're-ahead spiel is extremely annoying and makes me wish Anthropic gets trounced.