logoalt Hacker News

kristopoloustoday at 3:12 PM1 replyview on HN

q4km is about 48 tps on a 4090. my llama.cpp params are --flash-attn on --parallel 1 --load-mode mmap


Replies

m_ketoday at 3:14 PM

With spec decode should easily get to >100tps

on my dual 3090s qwen 3.5 27b was running at around 110tps using the config from https://github.com/noonghunna/club-3090

make that 200tps on a single 5090, 4x faster than opus https://x.com/radixark/status/2088285681131110446

devs about to get handed a two 5090 box each and told to max that out