logoalt Hacker News

jbellistoday at 5:53 AM0 repliesview on HN

The only scenario is if you have enough work to do batch inference. Using a tiny fraction of GPU capacity to decode a single request at a time just doesn't make sense, as you say.