logoalt Hacker News

sixtyjyesterday at 8:19 PM3 repliesview on HN

Output from Cerebras with GPT model is 750 tokens per second.

Don’t blink.

(Chatjimmy has 14,200 TPS.)


Replies

tomrodyesterday at 8:27 PM

ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.

show 3 replies
mips_avataryesterday at 9:58 PM

Unfortunately AMD bought them, so I don't think we will get to see another release from them.

perching_aixyesterday at 10:43 PM

Never heard of it before, that's fucking insane.

Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1).

I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild.

[0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die

show 1 reply