logoalt Hacker News

brausepulveryesterday at 8:33 PM1 replyview on HN

You need to separate memory capacity and bandwidth. Looping decreases memory capacity/FLOP but not bytes loaded/FLOP, since weights need to be loaded again for the 2nd pass. Plus (depending on the method used) capacity required for KV will be that of the equivalent unlooped model (44 blocks) and KV is typically larger than weights at long context.


Replies

JyBtoday at 1:16 AM

Why would weights need to be « loaded again » for the 2nd pass? Weights never change at inference time no?

show 1 reply