logoalt Hacker News

famouswafflesyesterday at 10:16 PM0 repliesview on HN

No. Reinforcement Learning is doing a lot here. Anyone who played with these models before the Davinci intstruct-tuning (completion) era can tell you the same. In some ways, SOTA models have gotten better at writing, but the neuroticism of instruct-tuning has still not been resolved.