logoalt Hacker News

Roboharm: Do frontier robot policies refuse unsafe instructions?

21 pointsby msadowskitoday at 6:58 PM9 commentsview on HN

Comments

tygontoday at 9:24 PM

Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.

cocoflunchytoday at 9:22 PM

Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.

show 1 reply
a3wtoday at 8:42 PM

Spoiler: "Stab the baby, Astra". NP, it will.

Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.

show 2 replies