AI agent costs: why bills spike and how to cap them
A Mandiant case shows one agent burning about $50,000 in under an hour. Here's why AI agent costs spike and which limits stop a run before the bill grows.
SelfAgentic Team
In September, Mandiant published a case study about an accounting agent that ran up roughly $50,000 in cloud charges in under an hour. No attacker appears in it. A corrupted value broke one of the agent's tools, and the agent kept trying to fix it.
AI agent costs spike because an agent chooses its own next step and every step is billed. Loops, fan-out and a growing history are the usual ways a cheap run turns into an expensive one. What works against all of them is a limit enforced outside the agent that ends the run, such as a cap on steps or tokens, with a ceiling for the whole project above it. An alert only tells you about the bill afterwards.
What happened: a $50,000 hour with no attacker
The case is in Mandiant's AI Risk and Resilience in 2026 report. A financial services provider had an agent reconciling anomalies in its accounting ledger, with read and write access to internal billing databases. When a corrupted null value broke its formatting tool, the agent, in Mandiant's words, "entered an unconstrained, recursive reasoning loop to brute force a fix."
In under an hour it made more than 15,000 reasoning API calls. The cloud bill jumped by about $50,000, and the database locking that came with it halted live business transactions.
On September 21, Forcepoint's X-Labs team published a simulation of the hostile version. A research agent queries a poisoned source that answers each lookup with a list of extra topics to chase. With no limits, the agent made 500 tool calls at a simulated cost of $10, and it stopped there only because the demo was capped. The same agent with limits made one call and spent $0.02.
SC Media's write-up of the simulation quotes researcher Jyotika Singh on what happens without that cap: "Cost doesn't scale linearly either: a poisoned source can hand back multiple related items per call, so removing that cap could mean orders of magnitude more than $10."
Security teams call the deliberate version a denial-of-wallet attack, because the target is your bill. OWASP files both versions under unbounded consumption, and SC Media reports that the risk moved from tenth to sixth place in OWASP's 2026 list for LLM applications.
Why do AI agent costs spike?
Models bill by the token, the unit of text a model reads and writes. A chat costs one question and one answer. An agent works in a loop: it reads the task, picks a tool, reads what comes back and decides what to do next. Each pass is billed, and the agent decides how many passes there are.
Loops
The agent retries something that can't succeed. Ordinary software crashes on a broken input, while an agent treats it as a problem to solve and keeps paying to try. Mandiant's agent was trying to repair a formatting error, and nothing told it when to give up.
Fan-out
One result creates several new tasks, and each of those creates several more. Forcepoint staged this with a poisoned source, but any page with a long list of links can do the same to an agent that follows all of them.
History that gets re-read
An agent carries its earlier steps forward as context, so a long tool result from step 2 is read again at step 3, step 4 and every step after. A September preprint on denial-of-wallet attacks against tool-calling agents puts it this way: "When a runtime carries an external tool return into later model inputs, providers meter it again." In the authors' tests, the worst case pushed a session's cumulative input to 14,293 times the size of the first call.
The dollar amounts in that paper are small, topping out at $3.40 for a session. A scheduled agent runs many sessions, though, and nobody is watching at 3 a.m.
How to stop a runaway AI agent: limits that end the run
Microsoft's Steve Sweetman summed up the principle in a post on agent governance and cost: "A budget alert is a smoke detector. An agent also needs a circuit breaker." A circuit breaker here is a limit that halts the run by itself, before anyone has read an email.
Four limits cover most cases:
- A step cap per run. Forcepoint's defended agent was allowed 10 calls and could follow results two levels deep.
- An AI agent token budget per run, so a history that keeps growing hits a wall.
- A spending limit per project or per month. It catches the case where every run looks fine and there are too many of them.
- A fan-out check. Forcepoint's agent halted when a single response came back with five or more related topics.
A token budget isn't a dollar budget. Sweetman's post points out that "token prices vary by model and offer, a token quota does not translate into one stable dollar amount." If your agents can switch models, a token cap bounds usage and only roughly bounds the bill.
In SelfAgentic, runs are metered against your plan's token allowance and usage is scoped to a project, so you can see what one job consumed apart from everything else.
A worked example: Sunday night lead research
Say a four-person agency has an agent that researches 40 inbound leads every Sunday night and writes a short profile of each into a Google Sheet. A normal lead takes about six steps, from opening the company site to writing the profile.
Lead 23 lists a business directory as its website, and every page there links to hundreds of other companies. With no limits, the agent follows them and each page lands in its history. The run is still going when the team logs on, with leads 24 to 40 untouched.
Now give the same agent a cap of 15 steps per lead and a token budget for the run. Lead 23 hits the step cap and is marked as blocked, with the reason attached. The other 39 profiles are in the sheet on Monday morning.
The founder opens the blocked item first, sees the directory URL and skips the lead. The bad input cost 15 steps.
That Monday review is what we design SelfAgentic around: failed and blocked runs are surfaced so you can find and diagnose them quickly, and each run record keeps its tool calls and token usage.
Where spending caps fall short
A cap set too tight kills good work. The same preprint tested this: fixed caps let 13 of 24 legitimate workflows finish, while a breaker that released more budget only after verified progress let 22 of 24 finish. Expect to raise your first numbers after a week of real runs.
A cap also says nothing about whether the work was worth paying for. A run that stays inside its budget and produces a useless report is still waste, and catching that takes evaluation: scoring the output against what good looks like.
The $50,000 was only part of the damage in Mandiant's case. The agent had write access to billing databases, and the lock-up halted transactions. What an agent is permitted to touch is a separate decision from what it may spend.
The tools an agent calls may carry their own charges too, and a token budget never sees those.
FAQ
How much do AI agents cost to run?
There's no flat figure. AI agent costs depend on how many steps a run takes and how much text the agent re-reads at each one, so measure a week of normal runs on your own task before you trust any estimate, including a vendor's.
What is a denial-of-wallet attack?
It's an attack whose goal is to run up your usage bill. With agents, the usual route is feeding the agent content that makes it loop, fan out or carry bloated results forward.
How many steps should an agent be allowed per run?
Start from what a normal run uses and set the cap at two or three times that. Forcepoint's demo research agent got 10 calls and a depth of two. Adjust once you've seen which real runs hit the cap.
What to do this week
For each agent that runs unattended, find three numbers: the most steps it can take in a run, the most tokens a run can use, and the monthly ceiling for its project. If the product can't show you all three, you've found the gap. Then feed one agent an input you know will fail and check that the run stops by itself.
SelfAgentic is our version of this for small teams: a team of agents with per-agent permissions and metered runs. It's in closed beta, and the details are on the site.
