I don't think the "don't really save you money" hot take holds water in every case.
Coding, maybe.
But for operationalized/repeatable tasks it definitely does.
For example I have a workflow that I was running in April that effectively would cost $30k in token spend for each full run.
However now, with GLM 5.3-flash, we've brought the cost down to $7k-9k with our evals showing we've had no loss in recall, precision etc..