Prediction - we are going to figure out SOTA AI performance without requiring 1TB of memory within a year or so.
Of course CXMT, Micron, and family will still be profitable, but maybe not 'surge 470% from IPO' profitable.
And by the time the models that require 1TB memory will be better.
It's like people saying mobile chips are going to be better than PC chips... until they realize PC can be made of mobile chips too if that comes true.
It sounds like you have strong faith that if we can do in 100GB what we can do in 1TB today, 10x-ing our model size[1] from today won't be better?
[1]whether it is 10x in size, or maybe running 10x the models to have things like judges etc.
The fundamental issue is that for these systems bigger is essentially always better (if affordable). So if we can squash something like Kimi K3 down to run on a 'normal' system, that just incentivizes devs to increase the model size until once again we're at the limit of what can be run.
I think the chipmakers are pretty securely set up here, the usual cyclical issues aside. Between the stubborn resistance of the data to being compressed without significant loss, and to a lesser extent the ability to do more with more memory (we’re surely at diminishing returns for LLMs already, but diminishing doesn’t mean zero) it seems like a fairly safe guess (I am not an expert on anything) that for the next decade at least you’ll still want at least 128-256GiB of video RAM to have a comfortably “SOTA-enough” experience. Then the other side of the rectangle comes from putting that VRAM in every consumer and office-worker device. Selling everything at a big premium to a few efficiency-obsessed SaaS providers is what you do when you’re temporarily supply-constrained. Sidelining the SaaS middlemen and selling to gloriously inefficient consumers who will leave their laptops closed for most of the day is the dream. This is probably not unrelated to the fact that the country which is currently making a big strategic push into memory manufacturing is also championing open models.
What do you base this prediction on? It seems very unlikely, unless by SOTA you mean the current SOTA.
And besides llm's memory requirements:
-HBM has a lot of room to grow(bandwidth/power/density). And an HBM memory takes 3x the amount of wafers of regular DRAM.
-HBM + better DRAM designs could increase random memory access speed by alot. I've seen research talking about 7x.
At what point it makes sense that half of consumer GPUs and most of the server GPU/CPUs start using HBM?
SoTA AI will move well past 1TB+ memory requirements in 1 year or so.
Also, you will be able to play with much more competent models locally. They will still feel like children compared to the adults living in the SoTA region.
Are you thinking of fundamental architectural changes (like the Transformer)? Or incremental (like MoE or GQA)? Are there specific neolabs or techs you are following that lead you to this prediction?
Indeed it seems quite possible we are one architecture breakthrough away from existing chip stock driving us all the way to ASI.
Nobody has done AFAIK the information theory to prove it’s not possible.
Yeah, you can directly print the SOTA AI model on the chip. I also believe this can be done on older process nodes. The most important factor would be speed. How fast can you go from new model to new chip?
Even if we do, we are going to need a lot of RAM for the billions of agents running everywhere.
I'm looking forward to being able to buy CXMT chips for running SOTA models locally. :)
You can get _better_ that SOTA AI performance by burning the weights directly to silicon, the problem is you're locked into that model
"we" being China?
NOPE. The opposite will happen. CXMT will flood the market thus making 1 TB models affordable.
that would be perfect timing
Hyperscalers are screwed, data center mania couldn’t even be completed during this massive spending spree, all while the people seemingly sitting on the sidelines are working on getting these things baked into the OS and chipsets in a way consumers wouldn’t notice
[dead]
How would we do that? There's no historic precedent for that. Its fundamentally an information theory thing: what's the max amount of intelligence you can get out of 1 KB/MB/GB? There has to be a limit and I'm not convinced that it's far off.