You would not give a new hire a company card with no limit and a vague brief. You have almost certainly given an agent exactly that.
Prompts get all the attention. Budgets get none, right up until an infinite loop produces an invoice that requires an explanation.
The Three Budgets Every Agent Needs

- Token budget. A hard ceiling per task and per day. Not a monitoring alert, an enforced limit that stops execution.
- Action budget. How many tool calls, retries and hops before the agent must stop and escalate. This is your loop protection.
- Money budget. Where the agent can spend actual funds, whether that is cloud resources, API credits or genuine purchases. Separate, tighter and always visible.
Most teams implement the first as an alert, skip the second entirely, and never consider the third until an agent gains a tool that costs money.
Why Alerts Are Not Controls
An alert tells you the money is gone. A limit stops it going. The difference matters most at 2am on a Saturday, which is when runaway loops happen, because nobody is watching to notice the alert.
Enforce at the gateway rather than in the agent. An agent asked to respect its own budget is being asked to police itself using the same reasoning that just went wrong.
What Happens at the Ceiling
| Behaviour | When appropriate |
|---|---|
| Stop and escalate to a human | Default for anything important |
| Degrade to a cheaper model | High-volume, low-stakes work |
| Queue for the next budget period | Batch work with no urgency |
| Hard stop with no fallback | Anything touching money |
Whatever you choose, decide it in advance and write it down. A budget with undefined overflow behaviour becomes an incident conversation at the worst moment.
The Attribution Problem
You cannot enforce budgets you cannot attribute, and MCP has no defined multi-tenancy or cost attribution model. When agents invoke tools autonomously, cost attribution and rate limiting are undefined at the protocol level.
So build it yourself at the gateway. Tag every call with the agent identity, the owning team and the task. Without that, your only lever is a single global limit, which means one team runaway loop throttles everybody.
The Loop That Costs the Most
The classic expensive failure is not a huge single request. It is a retry loop where an agent hits an error, tries again with a slightly different phrasing, fails again, and repeats through the night.
Two controls stop it. A hard retry cap per task, and error messages that tell the agent what to do differently. If your errors are actionable, agents stop guessing. If they are bare status codes, agents guess forever, and each guess costs money.
Conclusion
Give every agent a token budget, an action budget and, where relevant, a money budget, all enforced at the gateway rather than requested in a prompt. Define what happens at the ceiling before you need it, tag every call for attribution, and cap retries hard. An agent with a good prompt and no budget is not a productivity tool, it is an unbounded liability that means well.
Frequently Asked Questions
What is a sensible starting budget?
Take the cost of a typical successful run and set the ceiling at three to five times that. Then watch what actually hits the limit, because those cases are usually bugs rather than genuinely hard work.
Should agents know their budget?
Yes, tell them in context. Agents behave more sensibly with a stated budget, choosing cheaper paths. Just never rely on that instead of enforcement.
How do we budget for agents that make purchases?
Separate, much tighter, and always human-approved above a threshold. Emerging payment protocols use signed mandates precisely because an agent spending money needs cryptographic limits, not polite instructions.