logoalt Hacker News

lostmsutoday at 2:11 AM1 replyview on HN

KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?


Replies

walrus01today at 2:18 AM

Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.

show 1 reply