I think efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute; and we are not going to run out of economically useful things to do with it anytime soon on the demand side.
The harder thing to forecast for me is if we hit a wall on increasing efficiency, either on the model weights side or silicon side, with current approaches. If we have to switch to something like burning the model weights into silicon to continue to make gains, then the current math on general purpose accelerators might be upside down.
> efficiency is unlikely to result in lower demand for compute, instead more useful compute per watt increases the value of that compute
I buy that. Jevon's Paradox, sure.
> and we are not going to run out of economically useful things to do with it anytime soon on the demand side
This I don't buy. Not fully, at least. Whether or not there's demand for LLMs in some particular field is one thing, whether or not there is a sustainable business model to be built out of that demand is another thing entirely.
There is a staggering amount of money pouring into startups looking for novel use cases for LLM-based agents. As usual, 99% of them will fail, but those other 1% are going to have to look harder and harder to find a novel use case that can actually be served profitably.
First of all, there's only so many places where a chatbot is going to sell. But, that also seems to be the only interface anyone can come up with that allows a user to steer an agent through a long-running task reliably. I'd love to be proven wrong here.
Also, if current trends plateau and large datacenters are still needed for complex tasks, that would stimy growth of LLM usage across entire industries.
But, if present trends continue, then local inference will become feasible for most tasks. That would lower the barrier to entry across tons of heavily-regulated and/or cost-sensitive industries. But, widespread local inference will almost certainly come with a painful market correction centered around hyperscalers, which would itself dry up the pool for ventures into new markets.
When efficiency reaches the point where local models on consumer hardware are good enough, demand for cloud tokens could rapidly shrink.
> think efficiency is unlikely to result in lower demand for compute
You can't save yourself rich.
I agree; I don't think there's any reason to assume Jevons paradox won't apply.
> If we have to switch to something like burning the model weights into silicon to continue to make gains
I think that's already being considered semi-seriously [0][1]
[0] https://taalas.com/products/
[1] https://ir.amd.com/news-events/press-releases/detail/1296/am...