logoalt Hacker News

perching_aixyesterday at 10:43 PM1 replyview on HN

Never heard of it before, that's fucking insane.

Apparently they baked the Llama 3.1 8B model weights [0] into silicon (the actual hardware is called Taalas HC1).

I guess for the trillion parameter models this would not scale due to cost? Imagine buying GPT 6 in the form of a PCI-E card, pulling these speeds, with up to 120 cct agent sessions. It'd be beyond wild.

[0] the weights are also using some cut down small format, but HC2 will have regular FP4 supposedly, and support for 20B params on one die


Replies

TacticalCodertoday at 2:09 AM

> Never heard of it before, that's fucking insane.

They've been acquired by AMD. Those saying the model sucks are completely missing the point: it was a proof-of-concept.

The question is: what happens to a model like Anthropic's Fable 5 that does, what, 70 tokens/s (and requires lots of output tokens) when the latest open-weights model is etched on silicon and does 14 000 tokens/s?

Shall the better model still have the upper hand or will the raw speed compensate?

show 1 reply