Very good points. Incentives are just terrible for this.
Add in:
- harder to audit
- move cost of failure/reprompts to the user
- kind of noted by you, but all kinds of quantization, model pruning, model routing, A/B testing becomes invisible and without any repercussions. The ways to cost-optimize are just crazy.
IIRC Thinking Machines had a mode with deterministic numerics but it's more expensive to run due to limitations this imposes on cross-batch ops and ordering of floating point reductions, and their model is not great overall.