logoalt Hacker News

MBCooktoday at 1:59 AM4 repliesview on HN

But that means your different chips all have different sets of weights and are different generations.

If none of that is baked into the chip as now then all the chips are running the latest weights every time.

Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.


Replies

WithinReasontoday at 6:22 AM

Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.

kurthrtoday at 5:51 AM

Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.

If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.

What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.

geysersamtoday at 5:37 AM

Why would it be useless in 3 months?

show 1 reply
Certhastoday at 5:50 AM

Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on).

Would you use it?

show 1 reply