logoalt Hacker News

fnordpigletlast Sunday at 7:31 AM2 repliesview on HN

I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years.

That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding.

Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.


Replies

torginuslast Sunday at 10:16 PM

Afaik, for a MoE model, total size doesn't matter that much for how heavy it is to run, the size and number of active experts at any given time does. Of course, you still have to store the whole model in fast memory, so there's that, but the reason these things are getting this large is because it doesn't really affect runtime that much.

sneaklast Sunday at 8:26 AM

The APIs for the frontier models via the US hosters do the exact same thing wrt saving the requests and responses for data mining. Let’s not pretend that pervasive surveillance is an eastern thing.

show 1 reply