logoalt Hacker News

ethanzhang1024today at 2:18 AM8 repliesview on HN

If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?


Replies

walrus01today at 2:22 AM

Without having any inside information, one possible theory:

All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use.

or

The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.

wmftoday at 4:24 AM

Cerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.

show 1 reply
aurareturntoday at 7:24 AM

Probably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.

show 1 reply
smallerizetoday at 2:25 AM

If you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.

gamplemantoday at 8:52 AM

Cerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.

aseipptoday at 2:48 AM

The WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.

doctorpanglosstoday at 2:21 AM

it only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1.

imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)

show 2 replies
petesergeanttoday at 9:39 AM

I use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.