logoalt Hacker News

bcoriglianoyesterday at 8:38 PM1 replyview on HN

I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand.

And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.


Replies

pixl97yesterday at 10:47 PM

What's really funny about these eggheads encoding things like watermarks in LLM output is I've never seen one of them ask, what if the LLM does this back to pass hidden messages.