logoalt Hacker News

apitoday at 4:37 PM0 repliesview on HN

I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.