logoalt Hacker News

JamesSwifttoday at 12:18 AM1 replyview on HN

They are both on borrowed time. The only thing keeping them afloat business wise is that current hardware costs are prohibitive which keeps local models artificially constrained. The next era will be open weight / local model for the masses and SOTA for the big players.


Replies

Tareantoday at 10:42 AM

If your prognosis is true I'm not sure it's fully positive. Inference without batching is just massive less efficient, even if local hardware can be more energy efficient per operation and can skip out on cooling.

Open weights, obviously beneficial. Local compute when not necessary, seems like it'd be significantly worse for the environment?