logoalt Hacker News

MisterMunchkinyesterday at 6:09 PM1 replyview on HN

They've corrupted their data set by padding it with generated slop, in the misguided belief that you need 10PB of data to train a brain. Every training round they load more AI slop into it, further amplifying the slop language.

It's fascinating, Sonnet 4 is still available via API and it's so much less moronic than the current model. All of the em-dashes and nonsense are a result of the repeated rounds of reinforcement learning using slop data.


Replies

jp57yesterday at 6:45 PM

But why did that style of writing get rewarded? The people at Anthropic are ultimately responsible for the reward signal and what it produced.

show 2 replies