logoalt Hacker News

Groxxyesterday at 5:15 PM1 replyview on HN

Less "internal prompt" and more "they are trained to summarize after a </think> token"


Replies

astrangeyesterday at 6:12 PM

The training methods try not to apply any particular rules to the contents of the thinking text. That's called "optimization pressure on CoT" and is thought to reduce safety by inducing the model to lie (or stop clearly printing its intentions) in the thinking text.