And besides llm's memory requirements:
-HBM has a lot of room to grow(bandwidth/power/density). And an HBM memory takes 3x the amount of wafers of regular DRAM.
-HBM + better DRAM designs could increase random memory access speed by alot. I've seen research talking about 7x.
At what point it makes sense that half of consumer GPUs and most of the server GPU/CPUs start using HBM?