How's that jive with the fact that they're introducing a new model every other week?
Pipeline the burn into silicon, lower the latency as much as you can, for the 10-100x operation cost it's worth it. Imagine if frontier models cost $5/mtok and the 2nd or 3rd tier models cost $5/billion tokens for 3-month-old models.
The new model every week is not necessary at this point really. What if you could run opus 5 for the next couple years at 1/20 the cost?