logoalt Hacker News

jubilantitoday at 3:17 AM1 replyview on HN

> Taalas needed a giant chip (6nm) for an 8B model.

You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon. Which is also not new, it goes back to the 1980s with fixed function digital signal processors and little linear regressions or hardware classifiers for industrial control systems, all are the same basic principle.

It's just usually not worth it to go super small process node, because most models people thought to turn into silicon were pretty small parameter sizes. We're talking 10-100 weight regression or at most 2-4k weight neural net, used in some instrument or factory equipment. You can do a decent MNIST OCR with a 4k weight neural net. For this, 180/130nm is fine.

Or you might think it's required with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.


Replies

Aurornistoday at 5:56 AM

> You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon.

It is.

As I said, they could have shrunk it with a smaller process node, but that's not at all close to what would be required for a GPT Sol size model.