logoalt Hacker News

reliabilityguyyesterday at 10:36 AM1 replyview on HN

> You could run MACs directly in RAM

Sure, MACs are nice. However, unless there other, PIM-specific/optimal, algorithms, regular matrix multiplication algorithms like tiling-based won’t work here I think — how would the tile be shared? By doing read/write all the time?


Replies

whatshisfaceyesterday at 6:42 PM

Attention calculations aren't shared across more than one vector during next token prediction (thinking and writing) which this sounds almost perfect for. Per attention layer, for deepseek at 1M context, you want to broadcast a single 1KB vector to 4GB of dot products, and map reduce a 1KB vector back.

show 1 reply