logoalt Hacker News

apitoday at 1:34 AM8 repliesview on HN

This is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear.

GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.


Replies

mindwoktoday at 1:40 AM

Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.

If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.

But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.

aurareturntoday at 7:49 AM

By the way, this is the same argument that Michael Burry used to short Nvidia.

He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0]

The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example.

New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot.

New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion.

In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion.

We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices.

[0]https://inferencex.semianalysis.com/inference

show 1 reply
petratoday at 2:43 AM

I wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use?

In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system?

And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy.

So unless there are some other bottlenecks, we'll use a lot of simulation servers.

RachelFtoday at 3:14 AM

True, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO.

A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.

winridtoday at 1:35 AM

On the plus side, lots of cheap servers to swoop up :)

show 1 reply
jeffybefffy519today at 6:41 AM

Exactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.

__turbobrew__today at 3:30 AM

By the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.