logoalt Hacker News

We're going to need default hard budget caps on pretty much everything

579 points • by elffjs • today at 12:20 AM • 296 comments • view on HN

Comments

motionlessveloc • today at 3:00 AM

I used to work on a support team of a well known backend type service that had hard budget caps.

It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.

Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.

This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.

➕ show 18 replies
kqr • today at 5:46 AM

This goes beyond dollar charges. For production code to be reliable, everything needs to have a hard limit.

Queue lengths, request sizes, response wait duration, message payload size, authentication attempts, allocation rates -- there's always some upper number beyond which the system is so messed up you'd rather it crashes.

> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.

Indeed. If you want a surprise $10,000 bill that's still not an argument against a hard cap -- just set it at $9,999,999 instead, or wherever you don't want the surprise bill. There's always a number that indicates something has gone insane. There's always a sensible upper limit to any operation.

mstaoru • today at 11:29 AM

I find Google AI Studio / Dev Platform / whatchamacallit one of the worst in this.

When video models just came out, I wanted to make a short clip. Gemini didn't have enough control back then, so I started in Studio. I had $10 on my account. Try, try, try, not good, retry, altogether maybe 20-30 retries with 4 choices for a 10 second video.

Wake up next morning to an email from Google - my Studio account is frozen because of negative balance.

I check it, it's at -$160. Not financial ruin, but a painful sum for a 10-second video I didn't even use in the end.

I don't think there was (or is) a setting in there that says "stop everything when I'm in the negative".

Almost a year later, it's still as bad. Trying to make some music via Lyria, I need to wait 10-12 hours for the API charge to show in my costs. Now I had the bitter lesson, I wait after every major novel experiment to see how much it costed me.

Coding is solved, my ass.

➕ show 2 replies
hyperhello • today at 12:36 AM

These shouldn't even exist without a negotiated contract.

I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.

Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.

➕ show 1 reply
brap • today at 8:27 AM

When I first started working with cloud providers it was shocking to discover this feature doesn’t exist, basically anywhere.

It’s such a basic thing, not having it has to be deliberate to make you accidentally spend more than you’d like.

➕ show 2 replies
chrismarlow9 • today at 2:56 AM

Network saturation is difficult. Even if you turn off the endpoint you can still saturate the network in between. And it's still bandwidth.

I actually think network ACL triggers based on billing might be the only way to really enforce this.

I witnessed a DDoS attack once that changed how I think about billing. It was locally provisioned hardware and the attackers had saturated the switches. Naively I said "just block the CIDRs" but the problem was the incoming ram is so saturated that it can't even get to the point of "deny" in the firmware.

So from a technical perspective if there's an internal DDoS at AWS what do you do? Do you turn off the endpoint? Do you drop the sources from hitting it at the router? And even that costs money. Anyway that incident gave me a different level of appreciation for this challenge.

Edit: this is mainly targeted at the people complaining why this took so long. At some point in scaling even telling you "no sorry" in a nice way is expensive. I'm sure recruiters can sympathize with this nowdays.

➕ show 1 reply
akd • today at 1:42 AM

Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry

➕ show 1 reply
threecheese • today at 3:13 PM

The value proposition for software - and the reason why it's highly paid and has eaten the real world - is its "write once run forever (ish)" property. Unlike every other consumable, it doesn't expire.

In this proposed world where this static software is replaced with dynamic software ("AI will build it JIT"), doesn't that destroy the value proposition?

And in the world where normal software becomes AI-integrated ("software has intelligence "), isn't this tantamount to very very high hosting costs - where each execution now costs centi-dollars instead of nano-dollars?

Seems like the only use case where the cost model is equal or better is "AI writes the software".

joshdavham • today at 1:56 AM

It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.

Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.

➕ show 12 replies
modeless • today at 12:32 AM

Wait, Google Cloud finally added hard caps on spending per service? I've been wanting that for so many years! They sure took their sweet time.

https://cloud.google.com/blog/topics/cost-management/new-ear...

Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.

➕ show 2 replies
HamadMalikKhan • today at 9:01 PM

This, plus the implementation on the AI providers' side is still choppy. I remember one of our implementations had a $500 limit on OpenAI with auto-recharge enabled. It burned through $600 despite the lower limit.

You can be down serious money even with a limit set.

gausswho • today at 12:38 AM

I wonder if BigCorp adding spending caps is due to them getting sick of customers solving it for themselves with virtual cards.

➕ show 1 reply
karmelapple • today at 1:00 AM

Having a monthly summary or estimate of how your spending is going would be really useful, too.

Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.

I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.

I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.

sdcfgy • today at 11:53 AM

Let’s just call it what it is. We’ve been conned into accepting microtransactions for everything in computing. At the same time centralising and exposing everything publicly thus creating a huge risk surface area with respect to billing and security.

The whole idea was insane. We shunned paid compute services in favour of personal computing and then sold our souls back again.

fwlr • today at 5:45 AM

This is something payment providers and banks should be offering. Otherwise you’re hoping each of N sellers will spend engineering money to individually implement a feature that realistically reduces their potential profit, on the vague promise that it’s some kind of beneficial feature that will bring them profit, which is never going to work like you hope.

➕ show 2 replies
BodyCulture • today at 12:25 PM

Just imagine we could buy one piece of hardware where we could run all our software on, wouldn’t that be awesome? We could call it a „Personal Computer“ and it could start a revolutionary new way of doing business independent of gatekeepers! Freedom for free people!

➕ show 1 reply
thatsit • today at 9:05 PM

Well, we don’t need to reinvent pre-paid services, just because it’s „AI“

kinnth • today at 3:51 PM

The reason I will never use AWS again was all related to billing and budget caps. We had a small kubernetes cluster that failed to remove all of itself on a shutdown, it hung around like a ghost consuming resources for 3 months and nothing could see or find it, but we got billed!

Was a 6 month nightmare proving we did nothing wrong.

amelius • today at 12:20 PM

This should be a feature of our payment systems.

It should be possible to open my bank app, see all my recurring payments (subscriptions), and be able to cancel them with one button. And this should count as an official termination.

Also, limits, etc.

➕ show 1 reply
WhyNotHugo • today at 2:27 AM

So, prepaid services which you top up?

I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.

Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.

➕ show 1 reply
ozgrakkurt • today at 2:32 AM

Where I live, you can just not pay for something and it is cancelled.

Phone service, bank card, home internet etc.

If you don't pay your bill than they just cancel your membership and it works ok.

People in western countries are just getting shafted by companies for (mostly) no reason because an alternative balance is just inconceivable.

➕ show 2 replies
bl4kers • today at 12:38 AM

It should be illegal to not have them

➕ show 5 replies
biophysboy • today at 1:12 AM

My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”

➕ show 3 replies
kamyarg • today at 12:25 PM

Learnt recently DigitalOcean does not have hard budgets limits, only emails. I am surprised and torn between this being intentional vs. them not caring.

One solution is banking apps that let you create extra cards and assigning hard limits to spending. Wise & Revolut apps can do that.

Have to mention OpenRouter also, I logged into OpenRouter through pi, on the auth screen it optionally let me set a limit which is very smart and user friendly.

➕ show 3 replies
traceroute66 • today at 9:35 AM

Hard budget caps are readily available in Europe.

And everywhere in Europe you have clear price sheets (unlike the deliberately opaque mess of the US price sheets with more small-print than a packet of pills), which means even if you are at an EU provider with no hard caps you can still accurately reason and predict your costs.

Just a few examples....

Cloud providers:

    - Upcloud
    - Exoscale
Inference providers:

    - Verda
    - PrivateMode

I really don't buy the stories the US providers tell you that "its too difficult" or "what if you suddenly go viral".

The "viral" bit is easily solved through basic monitoring of metrics that everybody should be doing. I believe the cool-kids give it the fancy name of Site Reliability Engineering (SRE). All you need to do is top-up your balance / adjust your cap if your metrics are trending upwards for an explainable reason. Its not rocket science.

As for the "too difficult" that's just a lie. It just suits the US cloud providers better to have you spend spend spend on their messy soup of random interdependent microservices.

➕ show 1 reply
bwanab • today at 2:47 PM

Not making this political since this is true across countries and the various political spectrums, but you can see the same thing happening in national budgets in terms of programs set up by legislation that have built in financing formulas. Years later people wake up and realize the program is now spending an order of magnitude more money than anyone would have supported in the beginning. It's the same phenomenon over a much bigger scale and timeframe. If only they'd inserted a hard budget cap.

andai • today at 9:56 AM

I call this "blast radius". Earlier this year during the Claw hype I was reading about all kinds of elaborate schemes to prevent the agent from getting API keys.

I realized, what am I actually afraid of. Well, overspend. So I just set then all to disable auto-reloading. Now if it blows up, I'm down $5.

Same story with containers. Just give it root on a VPS, and if it blows up, I'm down $3.

vincentmathis • today at 7:14 AM

It should definitely be default for all payed APIs. I think every AI provider have this by default.

I've had an AI coding agent enter a doom loop for hours multiple times now. having a budget limit for my openrouter key is helpful, but it's already spent then.

I've been building a thing you can hook into ai workflows that uses multiple detectors if an agent is starting to loop and will then send a kill command to that chain. I don't have any testers for it though.

altcognito • today at 12:57 AM

So weird, cause it seems a lot of services are suddenly adding them. Huh, wonder what changed?

theodpHN • today at 3:27 PM

Nice to see the tech giants finally launch hard spending caps in 2026, matching some but not all of the resource/spend limiting features ($, compute, disk, tape, print, time) that were baked into mainframes and time sharing systems in the early 1970s! Everything old is "modern" again!

steveBK123 • today at 9:12 AM

The challenge with cloud has always been the selling point of “infinite scaling” vs the flip side “infinite billing”.

It’s not in their interest to make cost controls work well.

Ideally you’d be able to set something granular like “allow this service to scale up only 10x, measured at an hourly level, and alert me when it happens. Drop all requests that exceed 10x”.

And then you are mostly in a throttling situation until the burst clears or a human can review & accept increased usage is ok/increase thresholds. I’d rather have services go slow during excess load (like a real server) than go dark for remainder of month.

This seems a lot better than brute force “turn everything off at $X level of monthly billing” or “no limits you can charge me infinity dollars”.

SubiculumCode • today at 8:00 AM

More than just hard budget caps, we need agents to be aware of those hard budget caps, and plan accordingly. Too much agent behavior is driven by the immediate proximal goal instead of long term, strategic considerations. What we have now is agents deciding to on embark on expensive, circuitous routs to a goal, perhaps with potential but not certain in any case, benefits, and intervention depends on either an attentive user or forces stop after the money pot dries up. Prior research (#?) has Dem nsttated agent behavior to adjust behavior in response to budget considerations that appeared to enable more efficient token usage, and if that means it needs to be imposed on a custom harness (and not the self interested ai provider), so be it.

Havoc • today at 1:37 PM

It needs to be selectable.

Hobbyist needs to not get wiped out. Business need stuff to stay online.

In order to learn how to use the platform, employees in business and startups need a relatively safe space to learn too so even they have mixed needs.

Big tech is definitely the worst offender though on handing out footguns and relying on "beg support for mercy" model.

gozucito • today at 9:30 AM

Why don't I have the ability to delete my payment details from Codex or Claude? That way I cannot be billed more until my subscription is over.

Turns out it's impossible to do. You have to delete your entire account.

➕ show 1 reply
iambateman • today at 3:26 AM

I learned this the hard way with OpenAI last week. A key got hacked and a Chinese-language bot used $300 in tokens in an hour.

I had a spending limit on for $30, so why did it keep charging? Because the spending limit is meaningless without a hidden checkbox called “enforce spending limit,” which is (or at least was for me) off by default.

To OpenAI’s credit, they refunded the money.

sailfast • today at 2:54 PM

To create these caps wouldn’t AI companies need to tell us how much we are spending in realtime? Not sure they want to do that. The complete lack of transparency is a feature and not a bug. Just like in the healthcare space.

darkstars • today at 1:09 PM

Surprised and happy to see this. Made an account just to comment.

Just last week I was looking for similar functionality in Cloudflare as I am exploring publicly exposing apps I have built (as opposed to just me and some friends on a home lab).

My main concern on public cloud platforms is costs. I never want to spend more than, say, 10 EUR on a simple service. In my mental risk matrix, the chance of a cost-related issue has been increasing more-and-more. I am creating and hosting more than ever, and the chances of malicious activity/abuse are in my opinion higher than ever before.

It feels like cloud providers either are unaware of this, or riding a wave of increased turnover - which I think a more likely.

baxtr • today at 5:45 AM

I am a little bit surprised by the lack of depth in discussing such a major product change.

I get it, the problem is definitely worth solving for. Waking up with a $100k bill isn’t great.

At the same time, from a product perspective the proposed solution might be a bad idea. Simply having hard caps as default will definitely turn out to be as bad for some people as a $100k bill, see for example (1).

You can’t come up with good product changes if you don’t discuss the potential negative effects of a change.

(1) https://news.ycombinator.com/item?id=49950196

➕ show 2 replies
tkgally • today at 1:42 AM

> In an ideal world, our agents could help with this.

Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].

A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.

[1] https://github.com/tkgally/je-dict-1

[2] https://github.com/tkgally/eex-dict

jcims • today at 4:04 AM

I work at a place that has an eight figure monthly AWS bill. They won’t use this.

I had a personal development account for ~15 years. I tinker with infrastructure stuff and had built some centralized event reporting. One day about two years later I turned on sqs data events into cloudtrail. What I didn’t realize was that this closed a feedback loop and over the next couple of hours my run rate went to about $4k per day in cloudtrail+sqs usage.

I didn’t realize it until I hit the next months billing alarm immediately the next month. I’d racked up $25k in usage fees.

I’ll be using this feature. Nothing I run is worth that risk.

➕ show 1 reply
genxy • today at 1:14 AM

We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.

wiether • today at 10:06 AM

Should be noted that the AWS feature discussed is not pure AWS, it's the "Builder" thing that is built on top of AWS, some kind of PaaS.

A regular AWS account doesn't have the notion of "project" and Cost Management doesn't have those spend limits.

neom • today at 1:18 AM

I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.

mattlondon • today at 8:02 AM

When I've used GCP in the past this has really really annoyed me.

Someone told me some years ago that it was "impossible" for Google to apply an upper usage caps due to technical design reasons which I found absurd. You can build a globally distributed continent-scale SQL database with consistency guarantees, but you can't stop me paying more than X USD when my usage ticks over that? Huh?

It smelt much more to me like Google's business model made it impossible to stop customers spending money, not their engineering.

I voted with my wallet.

gadders • today at 5:06 PM

1000% agree. I'm sure it is technically possible.

dpedu • today at 12:40 AM

I'm surprised this isn't law. Should it be?

➕ show 2 replies
pwndByDeath • today at 2:48 AM

I read the title and thought it was going to be a rant about how nations will get out of crushing debt...

jherdman • today at 1:37 AM

"Oh hai. I'm hooked on the drugs. Please stop me from taking more. kthnxbai."

Seriously. You all asked for this.

➕ show 3 replies
Terretta • today at 1:40 PM

FTA:

> Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.

Monitoring and a circuit breaker. If you are given the tools, make your own heuristic and flip the breaker yourself. Don't let someone else turn it off, and be hostage to their process failures for getting it back on. By self-selection, if you are a hard limit customer, you are not front of their service line.

> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.

Based on a career serving the Fortune 50, CTO for trillion dollar bank, etc.: "No." Assuming actual business is being done by the machinery in question...

The lights must stay on.

Cost management cannot shut off the enterprise. Cascading costs, even before reputation, are incalculable.

➕ show 1 reply

🔗 View 42 more comments