logoalt Hacker News

extrtoday at 4:01 PM2 repliesview on HN

It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.


Replies

causaltoday at 4:05 PM

Does not explain timing

show 1 reply
lossolotoday at 5:20 PM

This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.