logoalt Hacker News

fzysingularitytoday at 4:02 PM0 repliesview on HN

Am I missing something here:

p(y = next thinking+decision token | x = question) != p(y = next decision token | x = question)

The former is what LLMs are trained for, the latter is what Jev was likely trained on (likely used thinking alignment as an auxiliary loss, but not explicitly included in the probability calibration).