logoalt Hacker News

rad-byesterday at 5:35 PM1 replyview on HN

Great read and interesting analysis! I’m less charitable toward Anthropic supposedly not distilling ChatGPT for training purposes. Maybe not today, but during the GPT-4 era when Anthropic was the underdog - I can see it happen. Packaged along with some of Amodei’s clever jumping through hoops to prove how that is, in fact, virtuous.


Replies

cmatoday at 5:31 AM

Distilling is only prohibited from chats you prompt. But there are lots in the open like the LMSYS Chatbot Arena data. It's a essentially one of the largest public sources of preference data, maybe still useful for boostraping RLHF and successor techniques (possibly a more highly educated contributor base than the low paid contractors), and was full of GPT-4-era inputs.

show 1 reply