Putting a price ceiling on an agent run

  • CLI
  • Chat
  • Portal
  • Budgets

A chat reply is one model call, and you can roughly guess what it costs. An agent run isn't. You give it a task, and it reads files, runs commands, calls the model again with what it found, maybe hands part of the work to a sub-agent, and keeps going until it's done. One prompt can turn into a few calls or a few dozen, each carrying a longer conversation than the one before.

Monthly budgets are a good backstop, but they catch a runaway run only after it has spent the money. What you want is a limit for each run: stop this task once it has cost this much.

A limit for one run

In the CLI, it's a flag:

bash
tempr --max-cost 0.50 "migrate the payment tests to the new fixtures"

The run keeps a total of what its model calls have cost, sub-agents included, from the cost each call reports. Once the total reaches the limit, the run stops before its next model call, the same way Ctrl+C stops a turn. The CLI exits with code 6 and status cost-limit, so a script or CI job can tell this apart from an error. With --json, the final result line also carries spentUsd.

The IDE extensions have the same limit as a setting: "Stop an agent turn after (US dollars)" in VS Code, JetBrains and Visual Studio. It's blank by default.

Enforced on the server too

A limit that only the client checks has gaps. The client doesn't see every call billed as it happens, and a run that's stopped on your machine can still have work going on the server. So Tempr's server enforces the limit as well. Each run is sent what is left of its limit, and the server stops the run before any model call that would start once the limit is reached.

Two things to know about the edges:

  • The limit is checked between calls, so a run can finish a little above it, by the cost of the call that was already running when it crossed the line. Set the limit with that call in mind.
  • A model with no known price can't be counted. Tempr prices a call from what the provider reports or from the model's listed price. When it has neither, as with a custom model you haven't priced, the call counts as $0, and the CLI warns you rather than pretending the limit still holds.

A limit for your whole organization

On a team, the person setting the limit often isn't the one running the agent. An organization's owners and admins can set the most one agent run may cost on the Portal's Budget page, and it applies to every member, in the CLI and in the IDE extensions. Members can set a lower limit for themselves, but not a higher one, and --max-cost can only lower it. Changes to the limit appear in the organization's audit log.

When a run hits the organization's limit, only that turn stops. The session carries on, and the next run starts with a fresh limit, so a developer can look at what happened and decide whether to continue.

Seeing where the money went

A limit stops a run. It doesn't tell you which runs are worth a closer look. The Portal's Usage page now lists this month's most expensive agent runs, with when each started, who ran it, which model it used, how many model calls it made and what it cost, sub-agents included. On a team, the same list appears on the Budget page for owners and admins. Each run links to its trace, so you can see which step took the most.

One more budget tool shipped alongside this. Monthly budgets on virtual keys, on members and on an organization's hard cap now have a Strict option. Without it, several requests running at the same moment can each see room left in a budget and, together, go past it. With it, they can't. Strict mode adds about 5 milliseconds per request, and up to about 30 when many requests share a budget at once, so it's off unless you turn it on.

The details are in the docs: capping what a run spends in the CLI guide, and the CLI reference for every flag.