It's one part technical, one part a product decision. The technical part is that billing is not actually instant. As a most basic example, a VM reports its billing units every X period of time it is active. If there is some network blip but it's still running, then that billing data could be delayed.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
Regarding hobbyists: some clouds, and also the bigger clouds, do target hobbyists quite a lot, presumably with the intention that some of those hobbyists will eventually grow and stick with them. So I don't think "not wanting hobbyists" is entirely true.
But yes, the amount of money they spend is less, so it makes less sense to implement features that only hobbyists want.
Amazon's new feature for this specifically says that it won't delete any of your data for 90 days:
> If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.
From https://docs.aws.amazon.com/accounts/latest/reference/create...
I think there is a middle ground between deleting data and allowing 5000 VMs to be created to mine bitcoin. Obviously there are a lot of different scenarios to consider but the explosive costs seem to be constrained mostly to a couple of features which would be fairly safe to cap.
There is a middle ground here: Make this setting configurable, off by default, and with all the associated warnings of what will happen if you switch it on. Better yet, make it settable on each billable service. Keep Route 53 going, but halt and delete any VMs that exceed some limit.
> The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups
Hit the nail on the head.
Nearly all businesses would prefer a cost overrun than services going offline.
> Are we including deleting RDS data? S3? Glacier storage?
If you're billing per GB of storage, then you can put hard caps on storage capacity, and then hard-reject any operation that would take the total stored size over that capacity.
A lot of billing systems are organized around event delivery. The system does what it does and reports usage. This reporting is asynchronous and can be done E.G. via cron jobs running on a 24 hour cadence in certain cases. There's an internal guarantee that billing records for a given period are delivered by a certain time. Nobody checks whether the user has enough money to do what they're trying to do, just whether they're authorized to access the system in the first place. Shutting down accounts due to non-payment is more of an abuse / fraud concern, and happens long after the bill is delivered.
Storage is almost never the main thing racking up bills so it should be handled in a different way. If you hit a limit you can't store any more data but everything you have stays there and readable.
This is more about compute, VMs, LLM inference and services like hosted database. These are all safe to stop if the system triggers a normal shutdown when costs hit a limit.
A provider could, in theory, calculate the cost of storage for X days and prevent you from uploading an object that would push you over the limit.
For S3 a reasonable option is a data limit as the primary limit. If you set it a TB over your real needs you only waste one dollar per day after it locks. And there's no need for any other paused service to charge more than that for idle data.
You're right that nobody wants deletion. Spending limits do not imply deletion.
Yeah, you can't just implement it as a pure "stop all services immediately once I hit a set amount"
It needs to be more like "don't allow spinning up additional services after you hit this amount", although that still allows you to go over the limit by a lot, since most services are billed hourly.
It really is difficult to implement a spending cap that doesn't risk shutting down important things.