logoalt Hacker News

fy20today at 1:35 AM0 repliesview on HN

It doesn't really matter though. Hardware performance is still growing. The new Mac Studio could just about run this model locally (rather slowly) - something that sits on your desk, that you as a consumer can buy.

Imagine prosumer desktop hardware 10 years from now. The 2036 DGX Spark. For a few thousand dollars you will be able to buy something with hundreds of GB (maybe TB if manufacturers step up) of unified RAM, memory bandwidth in the 10-20TB/s range. Overall AI "compute" will increase 10-20x, while at the same time AI model capability per byte will increase 5-10x.

The hardware would fit today's models, something like Kimi K3, quite comfortably and give performance of maybe 100 tokens/second. So what needs data center hardware today will run on your desk.

But if we also assume the models become more efficient, a 2036 Fable-class model (in terms of intelligence/capabilities, not size) will easily run on this thing at hundreds of tokens per second.

Unfortunately it'll still slow to a crawl with 5 Chrome tabs open, and every Electron app will need at least 200GB of RAM.