logoalt Hacker News

pseudosavantyesterday at 6:31 PM0 repliesview on HN

It could be quite a while before we reach that point though. 5+ years easily.

I've been keenly interested in the ability to run local models, but the hardware is just not there. Consumer RAM speeds and capacity will have to significantly increase before local models will be able to perform as well as even the lowest end GPT-5.6 Luna model.

This is on the backdrop of RAM becoming prohibitively expensive. And without the speed and quantity of RAM, it becomes impossible to generate tokens at interactive speeds, regardless of model. There is a fundamental dependency between calculating all of the active params with the given RAM speed.

Even with a model that has been quantized all the way down to Q4, the DGX/RTX Spark chip with 128GB of RAM can only generate ~18 tokens/sec for a MoE model with only 30B active parameters. There haven't been any broadly useful models below 30B active parameters. And that is for a $5000+ piece of hardware that will be one of the best for running on-device models.

I really want to buy instead of rent my AI, but the economics are truly terrible.