How to set spending limits for AI agents.
A useful spending limit is not one number. It is a small authority envelope that defines who may request, what they may buy, how much they may spend, and when a person must intervene.The short version
Start with a narrow job, a monthly ceiling, a smaller autonomous transaction cap, and explicit review conditions. Then add merchant, category, geography, time, and lifecycle controls. Test representative purchases before activation, and make pause and revoke immediately available.
Begin with the agent’s job, not its budget
A budget without a purpose is only a larger blast radius. Before choosing a dollar amount, write down the job the agent is expected to perform. “Procurement agent” is still broad. “Renew approved software subscriptions and order routine office supplies for the San Francisco team” gives the policy designer something concrete to constrain.
The purpose should guide every later field. A software-renewal agent probably needs recurring SaaS merchants and a narrow software category. A travel-booking agent needs airlines and hotels, different transaction amounts, and more geographic flexibility. Combining those jobs into one identity makes the policy harder to understand and the audit trail harder to defend.
When two jobs need materially different permissions, create two agents. Least privilege is easier to maintain when authority follows a clear operational role.
Separate monthly budget from autonomous authority
The monthly budget answers how much the agent may consume during a budget period. The maximum autonomous transaction answers how large one request may be before a person reviews it. They should not be the same number.
Suppose a procurement agent has a $2,000 monthly budget. A $2,000 autonomous limit would allow the entire budget to leave in one request. A $250 autonomous limit lets the agent handle ordinary renewals while escalating an $899 laptop. The remaining budget still matters: an otherwise valid $96 request should decline if only $40 remains.
A practical starting point is to review recent human-approved purchases for the job. Set the monthly ceiling near the legitimate operating need, not the largest imaginable emergency. Set the autonomous cap around the upper edge of routine purchases. Exceptional spending belongs in a review path, not inside everyday authority.
Use an approval threshold for ambiguity
Amount is not the only reason to ask a person. A request can fit under the transaction cap and still deserve review because the merchant is new, the category is adjacent to the agent’s purpose, or the geography is unusual.
Think of approval conditions as deliberate friction. They preserve speed for known, repeatable work and create a checkpoint when context matters. Common conditions include every new merchant, any transaction above a threshold, every international request, unusually rapid purchase attempts, and any category outside a tight allowlist.
Do not confuse approval with a soft decline. A blocked category such as gambling or cryptocurrency should remain declined. If the business intends to permit an exception, an operator should deliberately change the mandate and submit a new request. That preserves evidence that the old authority did not allow it.
Prefer allowlists for narrow jobs
There are two basic ways to express scope. An allowlist says which merchants, categories, or countries are eligible. A blocklist names known prohibitions. For a narrowly defined agent, allowlists usually provide a clearer boundary because everything outside the expected job starts in a restricted state.
A software-renewal agent might allow software and cloud infrastructure, allow Notion and AWS, require review for other merchants, and limit countries to the United States. A general operations agent may need broader categories but can still block crypto, gambling, cash equivalents, gift cards, and known risky merchants.
Merchant names require normalization in production. “Amazon,” “Amazon Web Services,” and a payment descriptor may represent different business relationships. Category data may also be incomplete or disputed. Treat uncertain classification as a reason for review, not a reason to silently infer permission.
Add time and lifecycle boundaries
Authority should not last forever by default. Mandates can expire at the end of a project, contract, travel window, or fiscal period. Temporary agents should receive temporary authority. Permanent agents still need periodic review.
Status is a separate control. Pausing an agent should stop new requests without erasing its history. Revoking it should end the relationship and invalidate credentials. Requiring approval for all future transactions is a useful intermediate state when behavior looks suspicious but the operator still needs visibility into attempted work.
Budget changes and threshold changes should create new mandate versions. Historical transactions must continue pointing to the version evaluated at request time. Without versioning, today’s settings can misleadingly appear to justify yesterday’s decision.
Model a complete authority envelope
| Control | Question it answers | Example |
|---|---|---|
| Agent identity | Who is requesting? | Procurement agent |
| Monthly budget | How much may it consume this period? | $2,000 |
| Autonomous maximum | How large can one request be without review? | $250 |
| Approval conditions | When must a person decide? | Every new merchant |
| Allowed scope | What is eligible? | Software and office equipment |
| Blocked scope | What is never eligible? | Crypto and gambling |
| Geography | Where may the transaction occur? | United States |
| Expiration | When does authority end? | End of quarter |
The fields should be stored as typed data, not only as prose. Natural-language instructions are helpful for drafting, but the activated version must be inspectable and deterministic.
Test the policy before activation
Create a small test suite from realistic purchases. For the example procurement agent, try Notion at $96, AWS at $212, Apple at $899, Binance at $600, and a second Notion request after most of the budget has been consumed.
For each scenario, predict the result before running it. Notion may be approved because it is known, allowed, and below the autonomous cap. AWS may be approved under the same logic. Apple should require review because the amount is high or the merchant is new. Binance should decline because crypto is blocked. A later Notion renewal should decline if it exceeds the remaining budget.
Inspect the rule evidence, not just the final state. A correct-looking outcome can still be produced for the wrong reason. The trace should identify the agent, mandate version, budget snapshot, merchant status, category result, country result, risk factors, and final decision.
Design the human review moment
An approval queue should show enough evidence to decide without forcing the reviewer to reconstruct the policy. Include the amount, merchant history, remaining budget after approval, exact review-triggering rules, risk factors, and the scope of the action.
“Approve once” should resolve one request. It should not quietly add the merchant to an allowlist or raise the agent’s threshold. If the reviewer wants a durable policy change, that should be a separate, versioned action. Decline should also preserve the original request and the reviewer’s note.
After resolution, the item should leave the pending queue, appear in the transaction history, and append a human-resolution audit event. The interface should provide a receipt rather than leaving a stale error banner or inviting a second resolution attempt.
Monitor limits as the business changes
Limits are not set-and-forget configuration. Review requests that repeatedly require approval, merchants that become routine, categories that are never used, and agents that consistently approach their budgets. The goal is not to maximize autonomous approvals. It is to make the smallest safe authority envelope fit the actual job.
Watch for changes in transaction velocity, amount distribution, geography, and merchant novelty. Risk signals can make an eligible request more restrictive, but they should never turn an explicitly blocked request into an approval. Read why risk scores should not override policy for the reasoning behind that one-way relationship.
A practical activation checklist
- Give the agent one specific operational purpose.
- Choose a monthly budget based on legitimate recurring need.
- Set a much smaller autonomous transaction maximum.
- Allow only the categories, merchants, and countries the job requires.
- Block categories the business will not delegate.
- Require review for new merchants and ambiguous cases.
- Add an expiration or scheduled review date.
- Test ordinary, exceptional, prohibited, and exhausted-budget scenarios.
- Confirm pause, revoke, and approval-for-all controls work.
- Activate one version and preserve every later change as a new version.
Mandate’s simulator models this workflow without moving money. The important production principle is portable: authority should be explicit, narrow, testable, revocable, and separate from the agent asking to use it.