logoalt Hacker News

_345today at 4:05 AM1 replyview on HN

"87.3%

Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling."

_______

This is hard for me to believe. I have a lot of skepticism that frontier models like GPT 5.5 that are likely 2T+ parameters in size only got about 12% more accurate than an untrained 9b parameter LLM.


Replies

baqtoday at 5:29 AM

Why? This is a very narrow task, it’d be surprising if the results were different actually; more interesting question would be how an even smaller model performs in the same finetune.