logoalt Hacker News

frabcus • today at 8:36 AM • 1 reply • view on HN

Presumably this is much less good than Jev, because the normal LLM models have been trained with RLHF and to be agents. Especially on a large model, I'd expect it to decide in an earlier layer.

I'd hope whatever Jev's Reinforcement Learning for Calibrated Decisions (RLCD) does is better at training the models to give accurate probabilities in the weights.


Replies

jampekka • today at 8:54 AM

Empirically this approach is more accurate, faster and about the same price as Jev.

https://github.com/Mushroom-Systems/lichen