logoalt Hacker News

addagyesterday at 1:52 PM0 repliesview on HN

From the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]".

If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.