logoalt Hacker News

logicalleetoday at 2:17 AM4 repliesview on HN

(Where did you see that?)

This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use?

I think the frontier providers keep the size of their models carefully hidden.


Replies

redox99today at 5:20 AM

You don't really need to train a 10T model to test cerebras against a 10T model. You can feed it an untrained (randomly initialized) model and benchmark it. Result will be gibberish but performance the same.

nltoday at 4:47 AM

Mythos/Fable are around 10T:

> According to FT, industry estimates say Anthropic's most advanced Mythos 5 has about 8 trillion parameters and Fable 5 about 5 trillion

https://www.reuters.com/technology/bytedance-targets-mega-ai...

I believe this report has confused Opus (which is known to be around 5T) and Fable.

Other reports say 10T. See for example https://eu.36kr.com/en/p/3760679047267075?ref=explainx where Musk talks about the models being trained on Colossus2

show 2 replies
alightsoultoday at 4:44 AM

pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks

ewildtoday at 2:19 AM

It's rumored fable is around that 10T number

show 2 replies