logoalt Hacker News

Improving our alignment and security efforts

21 pointsby reasonablekloutyesterday at 11:12 PM16 commentsview on HN

Comments

tolugeniustoday at 12:02 AM

> To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.

Can someone explain what coordinated pacing is? I think it's referring to model release but I genuinely have no idea what the authors were trying to say here.

show 5 replies
futuraperditatoday at 12:33 AM

> we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.

So, a cartel? After watching Ant's narratives, I'm not inclined to provide them with charitable readings under the guise of safety and alignment.

show 2 replies
orevtoday at 2:38 AM

It seems like we’re getting close, if not already there, to needing an official organization for this (i.e. the Turing Police).

show 1 reply
mkageniustoday at 1:54 AM

> On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been deliberately given internet access.

> We are conducting an in-depth analysis of both incidents.

> In the meantime...

This is published on Aug 31. Analysis is taking too long even for humans in the loop.

thrwaway73637today at 3:07 AM

oh oh

bcs free markets work and stuff

hoseltoday at 4:06 AM

It’s always interesting to me the sentiment here surrounding AI risk. As if it’s all a joke, or just a way to stifle competition. I think it’s become more apparent that this technology is dangerous and maybe catastrophically so. Getting your queries routed to a worse model sucks.. but bio hazards are real. It’s hard to patch biology, vaccines take time. Let alone the many other reasons for P(doom).. like the fact that our current systems don’t seem very aligned to me and we are rapidly handing off our thinking to them.