logoalt Hacker News

decide1000yesterday at 9:27 PM1 replyview on HN

On the DGX I get 44.5 tokens per second (NVFP4). With 8 concurrent it's 241 t/s total.

I am using the PrismaAQUA

standard 9.7 t/s

+ Dflash2 30 t/s

+ torch-compile 37 t/s

c8 = 177 t/s


Replies

SwellJoeyesterday at 9:33 PM

What model? Also, I don't know what "the PrismaAQUA" means, ddg thinks it's a CPAP machine, which seems unlikely to help with inference performance.

Also, 4-bit has measurable intelligence loss. Sometimes worth it, but, at this size models are barely smart enough at 8 or 6.

show 1 reply