logoalt Hacker News

embedding-shapetoday at 11:25 AM2 repliesview on HN

> At this point it's just marketing stunts.

If you have access to a SOTA model without guardrails, provide a prompt that lets the agent come up with "creative" solutions to problems, and don't properly isolate it, they can end up inadvertently hacking 3rd party companies. Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.

It seems like if they released this models differently, say without the guardrails they currently have, we'd have a lot more collateral damage than we currently have.


Replies

BlobberSnobbertoday at 12:34 PM

It is a marketing stunt in the sense that, instead of being honest and saying "Taking structured output from token predictors and running that as commands for external tools, then passing the output back to the token predictor in a loop can lead to very bad consequences, especially if they have internet access.", they say "Our models are so freaking smart they can hack HuggingFace"

indymiketoday at 11:59 AM

> Even if it was a mistake or "mistake", the part where the agent can exploit things across multiple levels like that, isn't just marketing.

When we say "safety" people do not think we are protecting them from accidental automated crime at scale being committed on their behalf.

show 1 reply