logoalt Hacker News

danny_codestoday at 3:26 AM0 repliesview on HN

How can we prove intent? I’m sure a clever actor can disguise their prompt and simply claim the LLM suggested such and such on its own.

It’s not like Anthropic or OpenAI have the faintest idea how their models actually work.