Errors
Last updated
Errors come back in the format of the API you called: OpenAI's error object from /v1/chat/completions, /v1/responses, /v1/embeddings, and /v1/models, and Anthropic's from /v1/messages. Error-handling code written for either SDK keeps working unmodified.
Error shape
OpenAI format, used by every endpoint except /v1/messages:
{
"error": {
"message": "This virtual key is not permitted to use model 'openai/gpt-4o'.",
"type": "invalid_request_error",
"code": "model_not_allowed"
}
}
Anthropic format, used by /v1/messages. It has no code: error.type follows the HTTP status, as it does on Anthropic's own API, and the message is the same one the OpenAI format carries.
{
"type": "error",
"error": {
"type": "permission_error",
"message": "This virtual key is not permitted to use model 'openai/gpt-4o'."
}
}
| Status | Anthropic error.type |
|---|---|
400 | invalid_request_error |
401 | authentication_error |
402 | billing_error |
403 | permission_error |
404 | not_found_error |
429 | rate_limit_error |
503 | overloaded_error |
Other 5xx | api_error |
Every response, errors included, carries an x-tempr-request-id header. It matches the request in your request logs; include it when contacting support about a specific request.
Error codes
Every error the gateway itself returns, by what it concerns. An error from the provider behind a request is covered under provider errors.
Authentication
| Status | Type | Code | Meaning |
|---|---|---|---|
401 | authentication_error | missing_api_key | No virtual key was sent. Send it as Authorization: Bearer tvk_…, or on /v1/messages in x-api-key. |
401 | authentication_error | invalid_api_key | The key doesn't exist or has been revoked. |
401 | authentication_error | key_expired | The key is past its expiry date. Create a new key, or extend its expiry in the Portal. |
403 | permission_error | ip_not_allowed | The key has an IP allowlist, and the request came from an address that isn't on it. |
402 | billing_error | gateway_inactive | The key's owner has no active Gateway subscription: a personal key's account has none, or an organization's plan is past due or cancelled. |
403 | invalid_request_error | unsupported_key_kind | This kind of key can't be used on the Gateway API. |
The request
| Status | Type | Code | Meaning |
|---|---|---|---|
400 | invalid_request_error | invalid_body | The body couldn't be read, or isn't valid JSON in the endpoint's format. |
400 | invalid_request_error | missing_model | The body has no model. |
400 | invalid_request_error | invalid_reasoning | /v1/chat/completions with a malformed reasoning_effort or reasoning: an effort that isn't one of none, minimal, low, medium, high, xhigh, max; a max_tokens that isn't a positive integer; or effort and max_tokens together. See Reasoning. |
400 | invalid_request_error | guardrail_blocked | The request matched a guardrail set to block, so it wasn't sent on. The message says what matched: a sensitive-content pattern (email, card, jwt, bearer-token, aws-key, secret) or, where it's enabled, the prompt-injection check. |
400 | invalid_request_error | previous_response_id_unsupported | /v1/responses with previous_response_id on a model that isn't served by OpenAI's or Azure OpenAI's own Responses API. Tempr doesn't store responses; send the conversation in input instead. |
Embeddings
From /v1/embeddings, all before any quota is spent.
| Status | Type | Code | Meaning |
|---|---|---|---|
400 | invalid_request_error | missing_input | The body has no input. |
400 | invalid_request_error | invalid_input | input isn't a non-empty string, an array of non-empty strings, an array of token ids, or an array of token-id arrays, or it holds more than 2,048 inputs. |
400 | invalid_request_error | invalid_encoding_format | encoding_format is neither float nor base64. |
400 | invalid_request_error | invalid_dimensions | dimensions isn't a positive integer. |
400 | invalid_request_error | embeddings_unsupported | The model's provider has no embedding models Tempr can call (Anthropic has none at all), or on AWS Bedrock the model isn't an Amazon Titan Text or Cohere Embed model. The message lists the providers that work. |
400 | invalid_request_error | unsupported_input | The input is token ids and the provider (Google, Cohere, AWS Bedrock) takes text only. |
Prompt templates
| Status | Type | Code | Meaning |
|---|---|---|---|
400 | invalid_request_error | prompt_invalid_header | The x-tempr-prompt header isn't a JSON object with a slug. |
400 | invalid_request_error | prompt_not_found | The key's owner has no prompt template with that slug, or no such version of it. |
400 | invalid_request_error | prompt_variable_missing | The template uses variables the header's variables didn't supply. The message lists them. |
Models & providers
| Status | Type | Code | Meaning |
|---|---|---|---|
400 | invalid_request_error | unsupported_provider | The provider/ prefix isn't a provider the gateway supports. |
410 | invalid_request_error | provider_retired | The model's provider has shut down, for example Meta's Llama API (meta-llama/). The message names the providers that still serve the same models. See Supported providers. |
400 | invalid_request_error | model_not_found | None of the providers you have a key for serves this model. GET /v1/models lists the ones that do. |
400 | invalid_request_error | ambiguous_model | The model has no provider/ prefix and more than one of your providers serves it. Add the prefix. |
403 | invalid_request_error | model_not_allowed | The key's model allowlist doesn't include this model. |
400 | invalid_request_error | endpoint_not_configured | The provider has no API endpoint set up on the gateway yet. |
402 | billing_error | provider_key_required | No usable API key is set up for the model's provider, or its Azure OpenAI, Azure AI Foundry, or AWS Bedrock configuration is incomplete. Add it in the Portal. |
402 | billing_error | provider_auth_failed | Tempr couldn't sign in to Azure with the Microsoft Entra ID credentials you configured. |
Limits & billing
| Status | Type | Code | Meaning |
|---|---|---|---|
429 | rate_limit_error | rate_limit_exceeded | The requests-per-minute ceiling was reached: the key's or your plan's, whichever is lower. |
429 | rate_limit_error | tpm_limit_exceeded | The key's tokens-per-minute ceiling was reached. |
429 | rate_limit_error | quota_exceeded | The plan's monthly request allotment is used up and metered overage isn't on. |
402 | billing_error | overage_cap_exceeded | Metered overage has reached your overage cap for the month. |
402 | billing_error | virtual_key_budget_exceeded | The key's monthly budget is spent. If the budget is strict, also when this request's worst case doesn't fit in what's left; the message says what the request could cost. See strict budgets. |
402 | billing_error | organization_budget_exceeded | The organization's monthly budget is spent, and it's set to a hard cap. If the hard cap is strict, also when this request's worst case doesn't fit in what's left. |
403 | permission_error | end_user_blocked | The request's end user is blocked on this key. |
429 | rate_limit_error | end_user_rate_limit_exceeded | The request's end user reached their requests-per-minute limit on this key: their own, or the key's default for each end user. Retry-After says when to try again. |
402 | billing_error | end_user_budget_exceeded | The request's end user has spent their monthly budget on this key: their own, or the key's default for each end user. If the key's budget is strict, also when this request's worst case doesn't fit in what's left. Other end users of the key aren't affected. |
Fallback chains & backup keys
| Status | Type | Code | Meaning |
|---|---|---|---|
503 | upstream_error | no_healthy_candidates | No model in the fallback chain is usable right now: each is cooling down after failures, not on the key's allowlist, or missing a provider key. |
502 | upstream_error | all_candidates_failed | Every model in the fallback chain was tried and failed. The message names the last provider and the status it returned. |
502 | upstream_error | all_keys_failed | Every API key set up for the provider was tried and failed. The message gives the last status. |
Remote MCP proxy
| Status | Type | Code | Meaning |
|---|---|---|---|
403 | invalid_request_error | mcp_server_not_allowed | The key isn't allowed to reach this MCP server. |
404 | invalid_request_error | mcp_server_not_found | The key's owner has no remote MCP server with this slug. |
403 | invalid_request_error | mcp_connection_requires_user | The server signs each member in with their own account, and this key belongs to an organization with no member behind it. Use a member's key, or give the server a static header. |
401 | authentication_error | mcp_sign_in_required | The server signs each member in, this member hasn't, and the request body isn't JSON-RPC, so there's nothing to answer in its place. The message has the Portal link to sign in. |
Some refusals come back as JSON-RPC errors inside a 200, so MCP clients show them as tool errors:
-32602: a tool call's arguments were blocked by the MCP guardrails. The message names what was found, for example "an AWS access key". The request log recordsmcp_guardrail_blocked.-32042(MCP'sURLElicitationRequiredError): the member needs to sign in first.data.elicitations[0].urlis the Portal sign-in link. The log recordsmcp_sign_in_required.-32001: the server refused the call because the member's sign-in lacks a permission (insufficient_scope). The message names it. The log recordsmcp_insufficient_scope.
A withheld tool result is a normal result with isError: true and the reason as its text; the log records mcp_guardrail_withheld.
See remote MCP proxy for what comes back from the server itself.
Members' own requests
An organization's members reach models through Tempr's IDE extensions, the CLI and agent runs, not with a virtual key. Their requests answer these errors, as { "error": "…", "message": "…" }, plus currentSpendUsd and limitUsd for a budget. See Group budgets and models.
| Code | Status | When |
|---|---|---|
model_not_allowed | 403 | The model isn't among the member's own allowed models, or their groups' (the message names the groups). |
member_rate_limit_exceeded | 429 | The member's own requests-per-minute limit is reached. |
member_tpm_limit_exceeded | 429 | The member's own tokens-per-minute limit is reached. |
member_budget_exceeded | 402 | The member's own monthly budget is spent. |
group_budget_exceeded | 402 | One of the member's groups has a hard-capped monthly budget, and its members' spend this month has reached it. The message names the group. |
organization_budget_exceeded | 402 | The organization's hard-capped monthly budget is spent. |
Every budget applies, and the first one used up refuses the request. In an agent run on Tempr's servers, the run ends with the same message.
Limit headers
A request refused by a limit carries headers saying which limit and where you stand:
| Header | Sent with | Value |
|---|---|---|
x-tempr-rpm-limit | rate_limit_exceeded, end_user_rate_limit_exceeded | The requests-per-minute ceiling that applied. |
x-tempr-quota-used | quota_exceeded | Requests counted this month. |
x-tempr-quota-limit | quota_exceeded | The plan's monthly allotment. |
x-tempr-overage-used | overage_cap_exceeded | Overage requests counted this month. |
x-tempr-overage-limit | overage_cap_exceeded | How many overage requests your cap allows. |
x-tempr-quota-resets | quota_exceeded, overage_cap_exceeded | When the monthly counters reset, as an ISO 8601 UTC timestamp: the end of your billing period on a paid plan (a month from the day it started), or the start of next month on Free. |
x-tempr-budget-limit | virtual_key_budget_exceeded, end_user_budget_exceeded | The key's monthly budget, or the end user's, in US dollars. |
x-tempr-org-budget-limit | organization_budget_exceeded | The organization's monthly budget, in US dollars. |
x-tempr-budget-remaining | virtual_key_budget_exceeded, organization_budget_exceeded | What's left of that budget this month, in US dollars, to four decimals. |
x-tempr-end-user-budget-remaining | end_user_budget_exceeded, and every successful request whose end user has a budget | What's left of the end user's monthly budget on this key, in US dollars, to four decimals. |
Strict budgets
A budget checks what's already been spent, so requests that run at the same time can each pass the check and together spend past it. A key's budget, or an organization's hard cap, can be made strict: in the Portal, or with hold_worst_case in the management API. A strict budget sets aside each request's worst case while it runs, and refuses a request whose worst case doesn't fit in what's left:
{"error":{"message":"This request could cost up to $0.0750 and this virtual key's monthly budget has $0.0010 left. Lower max_tokens, or raise the budget.","type":"billing_error","code":"virtual_key_budget_exceeded"}}
- The worst case is the request's input at its model's price, plus
max_tokensof output. Withoutmax_tokens, Tempr assumes at most 8,192 output tokens, so set it when you need the true worst case held. A model with no known price sets nothing aside. - A strict budget adds about 5 milliseconds to each request, up to about 30 when many requests share it at once. Near the limit, requests that start together can all be refused when one of them would have fit; retrying gets through.
- What's set aside is released as soon as the request ends; only what it actually cost counts toward the budget.
Retries & fallback
Tempr already retries transient upstream 429/5xx errors against the same provider before giving up, and — if you've configured a fallback chain on the key — falls back to the next candidate model automatically. When a provider asks for a wait longer than ten seconds, Tempr doesn't sleep through it: the 429 or 503 comes back at once with a Retry-After header carrying the provider's wait. A rate-limited model also sits out fallback chains until then, while that provider's other models stay available. A client-visible error means Tempr's own retry and fallback logic was exhausted, so honor Retry-After when it's there rather than retrying in a tight loop.
Provider errors
When the provider behind a request rejects it and Tempr's retries and fallbacks don't recover, you get the provider's HTTP status and its error message rather than one of the codes above. Tempr tidies a JSON error body into an error object with a message; a provider error that isn't JSON comes back as the provider sent it.
Errors during a stream
Once a streaming response has started, its status is already 200, so a failure after that point arrives in the stream. On /v1/chat/completions, a stream the provider cuts off before it finishes ends with this chunk, then data: [DONE]:
data: {"choices":[{"delta":{},"finish_reason":"error","index":0}],"error":{"message":"upstream stream closed prematurely","type":"stream_truncated"}}
Treat a finish_reason of "error" as a failed request: the output before it is incomplete.