logoalt Hacker News

whatshisfaceyesterday at 6:42 PM1 replyview on HN

Attention calculations aren't shared across more than one vector during next token prediction (thinking and writing) which this sounds almost perfect for. Per attention layer, for deepseek at 1M context, you want to broadcast a single 1KB vector to 4GB of dot products, and map reduce a 1KB vector back.


Replies

reliabilityguyyesterday at 7:24 PM

How exactly the map-reduce will happen though? Won’t you need to do it host-side, or make a lot of reads and writes?

Also, doesn’t it mean that you forgo batching?

show 2 replies