logoalt Hacker News

andaitoday at 6:23 AM3 repliesview on HN

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

Researcher: hack me

Model: understood

Researcher: oh my god


Replies

0xDEAFBEADtoday at 8:50 AM

You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?

show 1 reply
gallerdudetoday at 12:41 PM

But it wasn't explicitly told to hack HuggingFace. It was told "answer this security question", and it's answer was to break into the teacher's desk to find the answer key.

Tzttoday at 2:03 PM

Researcher: hack me

Model: I committed a crime

Researcher: oh my god

show 1 reply