logoalt Hacker News

NikolaNovakyesterday at 10:20 PM4 repliesview on HN

>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?

I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?


Replies

altruiosyesterday at 10:42 PM

> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?

not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter.

This "make me a billion dollars" is a maximal example (easy to go wrong). here is the same logic applied to a minimal example (harder to go wrong).

prompt: "make and pour me some tea", agent: goes and kills the grandparent to incinerate them to turn them to ashes to 'make tea'.

Is the human on the hook for the robot acting according to their wishes, but just happened to be aligned so that 'going to the store to buy something' was not within its capabilities, so it works with what it has on hand (the grandparent)?

We either need a much clearer line in the sand, or we need to treat each prompt with the same moral weight. My bet is on the latter.

show 2 replies
drdecayesterday at 10:40 PM

What if the user says “Make me a million dollars legally.” (Including the emphasis), and then the model ends up breaking through bank infrastructure (even though that is illegal)? Is it just because they were the last person to instruct the model, and you regard them as being therefore responsible for whatever it does in response? Or, does there have to be an element of “they reasonably could have anticipated this as an outcome that is likely enough to be worth considering” to it?

lukanyesterday at 10:34 PM

Because the basic assumption is always to stay within the bounds of the law.

show 1 reply
p1eskyesterday at 10:44 PM

If I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.