logoalt Hacker News

AmazingTurtletoday at 5:51 PM5 repliesview on HN

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.


Replies

boardwaalktoday at 6:52 PM

I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.

show 4 replies
zeroonetwothreetoday at 6:07 PM

Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.

show 2 replies
AtHeartEngineertoday at 6:48 PM

flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans

redanddeadtoday at 7:13 PM

Serving compute is their main value prop

Yet… even Altman called out Anthropic for serving dumbed down models.

Shits weird man

show 1 reply
ramesh31today at 5:58 PM

>Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.

show 2 replies