logoalt Hacker News

bjackmantoday at 9:28 AM3 repliesview on HN

I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW.

Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.

So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.


Replies

ttoinoutoday at 10:13 AM

We could make LLM inference 100x cheaper to run at home efficiently, but that solution might need to be updated every 1-2 years, whereas current GPU are useful for various others tasks and last longer

Tepixtoday at 11:42 AM

"Never" is a long time. Just think about how much ram we had 10 or 20 years ago. 1.5TB isn't a lot really.

show 1 reply
taneqtoday at 9:48 AM

‘Never’ is a big word in the computing world. 10 years from now a model this size will probably run on a high-end phone.

Of course, by then we’ll want to run something commensurately larger.