logoalt Hacker News

WhitneyLandtoday at 2:05 PM1 replyview on HN

How do you figure that?

When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB.

And this implementation is already cutting down the 1M token context window you would normally get.


Replies

Tepixtoday at 7:35 PM

For sure if you want to properly utilize the model with several users in parallel and large context you'll want two MI350P.