Not necessarily. Servers serving the model likely has enough traffic that they are batching decodes ...

ydj • today at 4:09 AM • 1 reply • view on HN

Not necessarily. Servers serving the model likely has enough traffic that they are batching decodes already. MTP reduces latency and increase efficiency only when the server can’t batch enough concurrent streams to be compute bound rather than memory bound.

Replies

imrozim • today at 4:16 AM

Fair didn't think about batching makes more sense for self hosted models then.

alt Hacker News

Replies