One plan, many keys: how limits work across a team's Gateway keys

  • Gateway
  • Budgets

A Gateway account starts with one key. It doesn't stay that way for long. The production app gets its own key, the staging app another, the nightly CI job a third, a contractor a fourth, and every developer on the team ends up with one too.

Each key fails in its own way. A CI job stuck in a retry loop, a key leaked in a public repository and a developer trying an expensive model all look different. But the keys also share one plan, one monthly allotment and one bill. Tempr's limits follow that split: some count every key together, others count only one key or one person.

The plan's rate limit counts every key

Every Gateway plan has a requests-per-minute limit: 10 on Free, 60 on Pro and 240 on Scale. That limit is shared by all the keys on the account.

It wasn't always enforced that way. Until recently each key got the whole limit, so an account with five keys could send five times what its plan allows. Now the plan's limit counts every key together, as the plan describes. If you run many keys at once, this is the change you're most likely to notice.

A refused request gets 429 rate_limit_exceeded with a Retry-After header: the seconds until the current minute ends. A client that honors it waits exactly as long as it needs to.

The monthly allotment is shared the same way: 10,000 requests on Free, 250,000 on Pro and 2,000,000 on Scale, drawn on by every key. Past it, Free stops, while Pro and Scale keep serving on metered overage up to a cap you set.

Each key has its own limits too

A shared limit protects the account, but it doesn't protect one key from another. If a CI job uses up the whole minute, your production app waits behind it. So a virtual key can carry its own limits, checked on their own:

  • Requests and tokens per minute. When a key has its own limit, the stricter of the key's limit and the plan's applies. Giving the CI key a low ceiling keeps it from crowding out the app.
  • A monthly budget. The most the key may spend in a month, counted from the provider cost of each request. A key over budget gets 402 virtual_key_budget_exceeded.
  • A model allowlist. The models the key may call. A request for any other model is refused before it reaches a provider, so a key meant for a small model can't be pointed at a large one.
  • An expiry date. After it, the key gets 401 key_expired. This is for keys that are meant to be temporary, such as a contractor's or one CI job's, so they stop working even if nobody remembers to revoke them.
  • An IP allowlist. Up to 50 addresses or CIDR ranges the key may be used from, and 403 ip_not_allowed from anywhere else. A key that only your servers use is worth much less to someone who finds it in a log file.

You set these when you create or edit a key in the Portal. On Pro, Scale and Enterprise you can also set them from your own scripts with the management API, which is how a CI pipeline can create a short-lived key with a small budget for each run and revoke it afterwards.

The idea behind all of these is to give each key only what its job needs, so a bug or a leak is bounded by that key's ceiling rather than the account's.

People, not just apps

In an organization, access also goes to people. When developers work in Tempr Chat, the CLI or the IDE extensions, the organization's pooled provider keys serve them, and an admin sets each member's budget, rate limits and allowed models. Members get the models they need without anyone pasting provider keys around, and the organization pays one provider bill.

Above every key and every member sits the organization's monthly budget. It can be a soft limit, an email alert to the organization's owners and admins, or a hard cap, which refuses requests with 402 organization_budget_exceeded once it's spent. A soft limit tells you without stopping work; a hard cap is for when a missed alert would be expensive. The Budget page shows the organization's spend by day, by member and by model.

These limits stack. An app's request has to fit within its key's limits and the organization's, and a developer's within their member limits and the organization's. Each layer answers a different question. A key's limits answer "what may this app do", a member's "what may this person spend", and the organization's "what may we spend in total".

A refusal that says what's left

When a budget refuses a request, the next question is how close it was. A 402 virtual_key_budget_exceeded or 402 organization_budget_exceeded response now carries x-tempr-budget-remaining: the dollars left of that budget this month, to four decimals. It sits next to x-tempr-budget-limit or x-tempr-org-budget-limit, which give the budget itself.

That lets a client tell two cases apart. With nothing left, retrying is pointless until the budget is raised or the month turns over. With money left, the refusal came from a strict budget, which also refuses a request whose worst case doesn't fit, and a smaller request may still get through. A separate post covers strict budgets.

Hearing about it before the refusal

Every plan is emailed at 80% and 100% of its monthly requests, and at 80% of the overage cap on a metered plan. The same alerts can go to a webhook as gateway.quota_alert events, signed so you can check they came from Tempr. In an organization, the emails go to its owners and admins.

The alert at 100% now says what happens next: on an account with overage switched off or no overage cap, new requests are refused until the quota resets.

Where to start

If you hand out more than a couple of keys, a sensible setup is:

  1. One key for each app and each CI job, with a per-minute limit on the noisy ones.
  2. A monthly budget and a model allowlist on every key, and member limits for each person.
  3. An expiry date on anything temporary, and an IP allowlist on anything that only runs on your own servers.
  4. An organization budget as a hard cap, if a surprise bill would hurt more than a stopped request.

The details are in the docs: what a virtual key scopes, rate limits and quotas and limit headers. Plan limits are on the pricing page.