logoalt Hacker News

CuriouslyCtoday at 2:53 PM0 repliesview on HN

Model output is pretty mid at augmenting, it can lead to distribution collapse. It's useful for smaller models because nobody wants manually to curate a specialized corpus and those models can't represent the diversity anyhow, but if the plan for infinite scaling was just to keep feeding the biggest model more of its predecessor's slop, that's not going to work out so well. It might work as a supplement for "thin" areas that have outsize importance for the amount of training data available for them though.

Big models are going to "tap out" on non verifiable fields within ~2 years, just because the pool of experts able to reinforce the models is going to get very small, and as the nuances get finer, the signal from reinforcement is going to get progressively less aligned with the intent. Math and code will be mostly tapped out in that time frame as well, even though we can technically scale them "infinitely," just because the cost benefit won't line up. At that point, most RL will be "gyms" with games that are designed to model designated valuable economic activity.

In the next few years, we'll get small domain specific distillates that are ridiculously smart in their domain (imagine if Qwen 3.X 27B went super saiyan), and even frontier labs will be routing to experts/orchestrating because the cost to serve/TPS difference is huge. They'll still train the god models for PR/marketing, c-suite use and distillation, but using them for day to day work would be like making houseware out of solid gold.