logoalt Hacker News

matteotomtoday at 6:13 PM0 repliesview on HN

I find it difficult to believe the inference only providers (Baseten, Fireworks, Digitalocean, etc) are all selling tokens at a loss.

Asking Claude for a rough estimate based on publicly available throughput and cost data for open weight models on modern GPUs suggests serverless, pay-as-you-go inference is profitable on owned GPUs with reasonable utilization (30-50%).