logoalt Hacker News

sheepscreek • today at 9:10 PM • 2 replies • view on HN

> 450M tokens / $150 / 5kWh

Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.


Replies

RussianCow • today at 9:35 PM

That number sounds about right, if a little low. According to OpenRouter, the weighted average input cost (which includes cache discounts) of GLM 5.3 is $0.2337 and the output cost is $3.291 per million tokens. If we assume 80% of the tokens are inputs, the cost of 450M tokens should be right around $300, which is the correct order of magnitude. And it depends highly on the provider(s) that the author is using, the ratio of inputs to outputs, etc.

It really makes you see how heavily subsidized the subscriptions are.

Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.

ThibWeb • today at 9:32 PM

The high usage was due to omp in vibe mode overnight, probably working way too hard through things. The high cost, yes we’d rather pay extra to work with providers that provide other benefits than just lowest cost possible (open models, no training, EU DC, etc)