logoalt Hacker News

YetAnotherNickyesterday at 8:56 PM1 replyview on HN

LLama 3 405B had the most unoptimized kv cache usage by far. Deepseek v4 pro uses 2.4GB for the same context length[1].

[1]: https://vllm.ai/blog/2026-04-24-deepseek-v4


Replies

philipportnertoday at 8:57 AM

Good point, thanks! I haven't been keeping up with most of the new model internals.