logoalt Hacker News

sothatsitlast Tuesday at 11:10 PM1 replyview on HN

The probability values don’t really represent confidence in modern LLMs though, especially after RLHF and RLVR.

System One says they use RLCD, Reinforcement Learning for Calibrated Decisions, which presumably has accurate probabilities as an explicit optimisation goal.


Replies

nkozyrayesterday at 3:05 AM

How is that different from RLVR?

show 2 replies