logoalt Hacker News

Wowfunhappytoday at 4:55 PM1 replyview on HN

When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)


Replies

chrswtoday at 5:12 PM

I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

show 3 replies