logoalt Hacker News

adrianNyesterday at 5:46 PM2 repliesview on HN

When efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.


Replies

bryanlarsenyesterday at 7:31 PM

Very few consumers are going to spend multiple thousands of dollars to save $10 per month. Companies absolutely will to save hundreds per month per employee, but that's not consumer hardware.

show 2 replies
pseudosavantyesterday at 6:31 PM

It could be quite a while before we reach that point though. 5+ years easily.

I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model.

This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed.

Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models.

I really want to buy instead of rent my AI, but the economics are truly terrible.