It is entirely possible that 5 years from now, there's >100x LLM inference going on - but demand for AI chips (including memory) is only 2x or less.
I doubt it. If LLM inference efficiency is 50x better than today, then there could be 1000x increase in inference volume due and we'll end up needing even more chips.Jevons paradox should win out for a long time for AI.
When internet connections got faster than 56k modems, we didn't use the same amount of bandwidth but faster. We used more bandwidth doing things like 4k streaming. I see the same in AI inference. If AI inference is that much more efficient, it will just enable more use cases for AI.
See for example, internet traffic over time: https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcS779VS...
Even after so many years, internet traffic continues to grow at an increasing rate.