There's been a ton of optimizations already, it hasn't remotely reduced demand even temporarily. More efficiency just makes the compute have even higher ROI per $ and watt spent.
With sufficient optimisation, there ought to be a tipping point beyond which local inference is good enough. And, sure, datacentre compute will still be needed for training but one of the biggest current uses will begin to taper off.
The question really is how soon we reach that tipping point, and whether it's before or after the current bubble runs out of steam for some other reason.
With sufficient optimisation, there ought to be a tipping point beyond which local inference is good enough. And, sure, datacentre compute will still be needed for training but one of the biggest current uses will begin to taper off.
The question really is how soon we reach that tipping point, and whether it's before or after the current bubble runs out of steam for some other reason.