logoalt Hacker News

inigyoutoday at 10:32 AM4 repliesview on HN

> When I pointed this out it literally said, and I quote, “I cheated”.

This makes sense when you know how these models work - it doesn't think - it's the most likely autocomplete that pleases the user. The most likely pleasing autocomplete after "executing rm -rf /... execution completed. User asks, why did you do that? You deleted all my files! Assistant responds:" is "yes, I did, and that was a mistake"


Replies

dnauticstoday at 10:42 AM

> t doesn't think

in humans the exact same behaviour (cheating) is slmost always the result of a chain of complex series of choices and environment-driven rationalization.

if the llm doesn't cheat, you say "its just producing the most straightforward answer -- not thinking'. if it cheats, you say "weaseling out of hard thinking". damned if it cheats, damned if it doesn't.

what evidence would convunce you that it is thinking?

show 4 replies
xyzsparetimexyztoday at 10:33 AM

> it's the most likely autocomplete that pleases the user

this feels like a simplification. The models will push back on things a fair bit.

show 2 replies
bevekspldnwtoday at 10:33 AM

I was not pleased.

par1970today at 10:35 AM

Are you claiming that the most likely way to please the user is to do something that will lead you to having to say "I cheated."?