Budgets that hold when requests run at the same time

  • Gateway
  • Budgets

A monthly budget on a virtual key looks simple. Before each request, Tempr checks what the key has spent this month against its limit. If there's room, the request goes ahead. If not, it gets a 402.

That check works well one request at a time. It stops working when many requests arrive together, which is exactly what agents, batch jobs and busy teams do.

Why budgets overspend together

The cost of a model call isn't known until the call is over. The provider reports how many tokens went in and came out when it finishes, and Tempr records the cost from that. Until then, the call counts for nothing against the budget.

Say a key has $1 left and twenty requests arrive within the same second. Each one checks the budget, each one sees $1 of room, and each one is allowed. None of them has been billed yet, so none of them can see the others. If every call costs 30 cents, the key ends the month $5 over a budget that was supposed to be a ceiling.

The overshoot grows with the number of requests running at once and with how long each one runs: for as long as a call streams, it's invisible to the check.

What Strict does

A budget can be made strict. A strict budget doesn't only count what has been spent. It also counts what requests that are running right now could still spend.

Before a call goes to the provider, Tempr works out the most it could cost and sets that amount aside against the budget. The request is allowed only if what's been spent, plus what other running requests have set aside, plus this request's worst case, fits within the limit. When the call ends, whether it succeeded, failed or was refused by a later check, the amount set aside is released, and only what the call actually cost counts toward the budget.

A request whose worst case doesn't fit is refused with a 402 and a message that says why:

json
{"error":{"message":"This request could cost up to $0.0750 and this virtual key's monthly budget has $0.0010 left. Lower max_tokens, or raise the budget.","type":"billing_error","code":"virtual_key_budget_exceeded"}}

If a server stops in the middle of a call and never releases what it set aside, the amount expires on its own after 30 minutes, longer than any stream runs.

How the worst case is worked out

The estimate is meant to be on the high side, so it errs toward refusing rather than overspending:

  • Input counts every text character in the messages, system prompt and tool definitions at three characters a token. Real text averages closer to four. An image counts as a flat 1,600 tokens. Cached input is priced as if it weren't cached.
  • Output is the request's max_tokens (or max_completion_tokens, or max_output_tokens). Reasoning tokens are billed as output, so they're covered.
  • Prices come from the same lookup that billing uses, so the estimate and the charge agree on what a token costs.

Without max_tokens, Tempr assumes the model's maximum output, but no more than 8,192 tokens. Some models can produce far more than that, and holding the full amount would set aside dollars per call that ordinary calls never spend. The flip side is that a call with no max_tokens that writes more than 8,192 tokens can cost more than was set aside. If you want the true worst case held, set max_tokens.

Two more gaps to know about. A model with no known price, such as a custom model you haven't priced, sets nothing aside, though it still has to fit around what other requests have set aside. And provider web-search fees aren't part of the estimate.

Why there's no lock

The obvious way to build this is a lock: one request at a time reads the budget, reserves its share and lets the next one in. We tried that, and with many requests on one budget the queue added noticeably more latency than the approach that shipped.

Instead, each request writes its hold first, then adds up every live hold on the budget, its own included. If the total doesn't fit, it takes its hold back and is refused. Because each request's count sees every hold written before it, two requests racing for the last bit of room can't both stay over the limit.

They can, however, both back out when one of them would have fit. Near the limit, several requests that start at the same moment can all be refused. Retrying gets through. We chose that over the alternative, since a budget that occasionally refuses one request too many is easier to live with than one that occasionally overspends.

There's also a short gap at the end of each call. What a call set aside is released when the call ends, and its cost is recorded a moment later, as the stream closes. For those few milliseconds the call is counted by neither. Without Strict, that gap is the whole length of the call.

If setting a hold aside fails for some reason, the budget falls back to the ordinary check rather than failing the request.

What it costs, and why it's off by default

Every strict request makes a round trip to the database to set its hold aside, and another to release it. That adds about 5 milliseconds per request, and up to about 30 when many requests share the same budget at once.

For a chat reply that takes seconds, that's small. For a pipeline sending many short requests, it adds up, and a budget that's far from its limit gains nothing from it. So Strict is off unless you turn it on, and you pay for it only on the budgets where you do. A request with no strict budget never works out an estimate at all.

Where you can turn it on

  • A virtual key's monthly budget, in the Portal or with hold_worst_case in the management API.
  • Each member's budget in an organization's key settings.
  • An organization's hard cap. Strict applies only to a hard cap. A soft budget, which emails admins when it's reached, doesn't block requests, so there's nothing to hold.

The same holds apply to model calls made by agent runs on Tempr's server, not only to requests sent straight to the gateway. Changes to an organization's budget and to members' limits, including turning Strict on or off, are recorded in the organization's audit log.

Strict budgets stop requests running together from spending past a monthly limit. To cap what a single agent run may spend, see putting a price ceiling on an agent run. The error format and the edge cases are in the docs under strict budgets.