logoalt Hacker News

andaiyesterday at 5:00 PM2 repliesview on HN

If I'm reading this right, they literally just ask a LLM to tell them what the traces say, with the key being that the traces are portable across LLM models, so they can switch to a smaller one that's easier to jailbreak.


Replies

dgellowyesterday at 5:20 PM

Correct, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches

show 1 reply
tw1984today at 5:12 AM

nothing scientific here, they basically just figured out some real issues caused by bad engineering practice.