Gateway API

Errors

Last updated

Errors come back in the format of the API you called: OpenAI's error object from /v1/chat/completions, /v1/responses, /v1/embeddings, and /v1/models, and Anthropic's from /v1/messages. Error-handling code written for either SDK keeps working unmodified.

Error shape

OpenAI format, used by every endpoint except /v1/messages:

{
  "error": {
    "message": "This virtual key is not permitted to use model 'openai/gpt-4o'.",
    "type": "invalid_request_error",
    "code": "model_not_allowed"
  }
}

Anthropic format, used by /v1/messages. It has no code: error.type follows the HTTP status, as it does on Anthropic's own API, and the message is the same one the OpenAI format carries.

{
  "type": "error",
  "error": {
    "type": "permission_error",
    "message": "This virtual key is not permitted to use model 'openai/gpt-4o'."
  }
}
StatusAnthropic error.type
400invalid_request_error
401authentication_error
402billing_error
403permission_error
404not_found_error
429rate_limit_error
503overloaded_error
Other 5xxapi_error

Every response, errors included, carries an x-tempr-request-id header. It matches the request in your request logs; include it when contacting support about a specific request.

Error codes

Every error the gateway itself returns, by what it concerns. An error from the provider behind a request is covered under provider errors.

Authentication

StatusTypeCodeMeaning
401authentication_errormissing_api_keyNo virtual key was sent. Send it as Authorization: Bearer tvk_…, or on /v1/messages in x-api-key.
401authentication_errorinvalid_api_keyThe key doesn't exist or has been revoked.
401authentication_errorkey_expiredThe key is past its expiry date. Create a new key, or extend its expiry in the Portal.
403permission_errorip_not_allowedThe key has an IP allowlist, and the request came from an address that isn't on it.
402billing_errorgateway_inactiveThe key's owner has no active Gateway subscription: a personal key's account has none, or an organization's plan is past due or cancelled.
403invalid_request_errorunsupported_key_kindThis kind of key can't be used on the Gateway API.

The request

StatusTypeCodeMeaning
400invalid_request_errorinvalid_bodyThe body couldn't be read, or isn't valid JSON in the endpoint's format.
400invalid_request_errormissing_modelThe body has no model.
400invalid_request_errorinvalid_reasoning/v1/chat/completions with a malformed reasoning_effort or reasoning: an effort that isn't one of none, minimal, low, medium, high, xhigh, max; a max_tokens that isn't a positive integer; or effort and max_tokens together. See Reasoning.
400invalid_request_errorguardrail_blockedThe request matched a guardrail set to block, so it wasn't sent on. The message says what matched: a sensitive-content pattern (email, card, jwt, bearer-token, aws-key, secret) or, where it's enabled, the prompt-injection check.
400invalid_request_errorprevious_response_id_unsupported/v1/responses with previous_response_id on a model that isn't served by OpenAI's or Azure OpenAI's own Responses API. Tempr doesn't store responses; send the conversation in input instead.

Embeddings

From /v1/embeddings, all before any quota is spent.

StatusTypeCodeMeaning
400invalid_request_errormissing_inputThe body has no input.
400invalid_request_errorinvalid_inputinput isn't a non-empty string, an array of non-empty strings, an array of token ids, or an array of token-id arrays, or it holds more than 2,048 inputs.
400invalid_request_errorinvalid_encoding_formatencoding_format is neither float nor base64.
400invalid_request_errorinvalid_dimensionsdimensions isn't a positive integer.
400invalid_request_errorembeddings_unsupportedThe model's provider has no embedding models Tempr can call (Anthropic has none at all), or on AWS Bedrock the model isn't an Amazon Titan Text or Cohere Embed model. The message lists the providers that work.
400invalid_request_errorunsupported_inputThe input is token ids and the provider (Google, Cohere, AWS Bedrock) takes text only.

Prompt templates

StatusTypeCodeMeaning
400invalid_request_errorprompt_invalid_headerThe x-tempr-prompt header isn't a JSON object with a slug.
400invalid_request_errorprompt_not_foundThe key's owner has no prompt template with that slug, or no such version of it.
400invalid_request_errorprompt_variable_missingThe template uses variables the header's variables didn't supply. The message lists them.

Models & providers

StatusTypeCodeMeaning
400invalid_request_errorunsupported_providerThe provider/ prefix isn't a provider the gateway supports.
410invalid_request_errorprovider_retiredThe model's provider has shut down, for example Meta's Llama API (meta-llama/). The message names the providers that still serve the same models. See Supported providers.
400invalid_request_errormodel_not_foundNone of the providers you have a key for serves this model. GET /v1/models lists the ones that do.
400invalid_request_errorambiguous_modelThe model has no provider/ prefix and more than one of your providers serves it. Add the prefix.
403invalid_request_errormodel_not_allowedThe key's model allowlist doesn't include this model.
400invalid_request_errorendpoint_not_configuredThe provider has no API endpoint set up on the gateway yet.
402billing_errorprovider_key_requiredNo usable API key is set up for the model's provider, or its Azure OpenAI, Azure AI Foundry, or AWS Bedrock configuration is incomplete. Add it in the Portal.
402billing_errorprovider_auth_failedTempr couldn't sign in to Azure with the Microsoft Entra ID credentials you configured.

Limits & billing

StatusTypeCodeMeaning
429rate_limit_errorrate_limit_exceededThe requests-per-minute ceiling was reached: the key's or your plan's, whichever is lower.
429rate_limit_errortpm_limit_exceededThe key's tokens-per-minute ceiling was reached.
429rate_limit_errorquota_exceededThe plan's monthly request allotment is used up and metered overage isn't on.
402billing_erroroverage_cap_exceededMetered overage has reached your overage cap for the month.
402billing_errorvirtual_key_budget_exceededThe key's monthly budget is spent. If the budget is strict, also when this request's worst case doesn't fit in what's left; the message says what the request could cost. See strict budgets.
402billing_errororganization_budget_exceededThe organization's monthly budget is spent, and it's set to a hard cap. If the hard cap is strict, also when this request's worst case doesn't fit in what's left.
403permission_errorend_user_blockedThe request's end user is blocked on this key.
429rate_limit_errorend_user_rate_limit_exceededThe request's end user reached their requests-per-minute limit on this key: their own, or the key's default for each end user. Retry-After says when to try again.
402billing_errorend_user_budget_exceededThe request's end user has spent their monthly budget on this key: their own, or the key's default for each end user. If the key's budget is strict, also when this request's worst case doesn't fit in what's left. Other end users of the key aren't affected.

Fallback chains & backup keys

StatusTypeCodeMeaning
503upstream_errorno_healthy_candidatesNo model in the fallback chain is usable right now: each is cooling down after failures, not on the key's allowlist, or missing a provider key.
502upstream_errorall_candidates_failedEvery model in the fallback chain was tried and failed. The message names the last provider and the status it returned.
502upstream_errorall_keys_failedEvery API key set up for the provider was tried and failed. The message gives the last status.

Remote MCP proxy

StatusTypeCodeMeaning
403invalid_request_errormcp_server_not_allowedThe key isn't allowed to reach this MCP server.
404invalid_request_errormcp_server_not_foundThe key's owner has no remote MCP server with this slug.
403invalid_request_errormcp_connection_requires_userThe server signs each member in with their own account, and this key belongs to an organization with no member behind it. Use a member's key, or give the server a static header.
401authentication_errormcp_sign_in_requiredThe server signs each member in, this member hasn't, and the request body isn't JSON-RPC, so there's nothing to answer in its place. The message has the Portal link to sign in.

Some refusals come back as JSON-RPC errors inside a 200, so MCP clients show them as tool errors:

  • -32602: a tool call's arguments were blocked by the MCP guardrails. The message names what was found, for example "an AWS access key". The request log records mcp_guardrail_blocked.
  • -32042 (MCP's URLElicitationRequiredError): the member needs to sign in first. data.elicitations[0].url is the Portal sign-in link. The log records mcp_sign_in_required.
  • -32001: the server refused the call because the member's sign-in lacks a permission (insufficient_scope). The message names it. The log records mcp_insufficient_scope.

A withheld tool result is a normal result with isError: true and the reason as its text; the log records mcp_guardrail_withheld.

See remote MCP proxy for what comes back from the server itself.

Members' own requests

An organization's members reach models through Tempr's IDE extensions, the CLI and agent runs, not with a virtual key. Their requests answer these errors, as { "error": "…", "message": "…" }, plus currentSpendUsd and limitUsd for a budget. See Group budgets and models.

CodeStatusWhen
model_not_allowed403The model isn't among the member's own allowed models, or their groups' (the message names the groups).
member_rate_limit_exceeded429The member's own requests-per-minute limit is reached.
member_tpm_limit_exceeded429The member's own tokens-per-minute limit is reached.
member_budget_exceeded402The member's own monthly budget is spent.
group_budget_exceeded402One of the member's groups has a hard-capped monthly budget, and its members' spend this month has reached it. The message names the group.
organization_budget_exceeded402The organization's hard-capped monthly budget is spent.

Every budget applies, and the first one used up refuses the request. In an agent run on Tempr's servers, the run ends with the same message.

Limit headers

A request refused by a limit carries headers saying which limit and where you stand:

HeaderSent withValue
x-tempr-rpm-limitrate_limit_exceeded, end_user_rate_limit_exceededThe requests-per-minute ceiling that applied.
x-tempr-quota-usedquota_exceededRequests counted this month.
x-tempr-quota-limitquota_exceededThe plan's monthly allotment.
x-tempr-overage-usedoverage_cap_exceededOverage requests counted this month.
x-tempr-overage-limitoverage_cap_exceededHow many overage requests your cap allows.
x-tempr-quota-resetsquota_exceeded, overage_cap_exceededWhen the monthly counters reset, as an ISO 8601 UTC timestamp: the end of your billing period on a paid plan (a month from the day it started), or the start of next month on Free.
x-tempr-budget-limitvirtual_key_budget_exceeded, end_user_budget_exceededThe key's monthly budget, or the end user's, in US dollars.
x-tempr-org-budget-limitorganization_budget_exceededThe organization's monthly budget, in US dollars.
x-tempr-budget-remainingvirtual_key_budget_exceeded, organization_budget_exceededWhat's left of that budget this month, in US dollars, to four decimals.
x-tempr-end-user-budget-remainingend_user_budget_exceeded, and every successful request whose end user has a budgetWhat's left of the end user's monthly budget on this key, in US dollars, to four decimals.

Strict budgets

A budget checks what's already been spent, so requests that run at the same time can each pass the check and together spend past it. A key's budget, or an organization's hard cap, can be made strict: in the Portal, or with hold_worst_case in the management API. A strict budget sets aside each request's worst case while it runs, and refuses a request whose worst case doesn't fit in what's left:

{"error":{"message":"This request could cost up to $0.0750 and this virtual key's monthly budget has $0.0010 left. Lower max_tokens, or raise the budget.","type":"billing_error","code":"virtual_key_budget_exceeded"}}
  • The worst case is the request's input at its model's price, plus max_tokens of output. Without max_tokens, Tempr assumes at most 8,192 output tokens, so set it when you need the true worst case held. A model with no known price sets nothing aside.
  • A strict budget adds about 5 milliseconds to each request, up to about 30 when many requests share it at once. Near the limit, requests that start together can all be refused when one of them would have fit; retrying gets through.
  • What's set aside is released as soon as the request ends; only what it actually cost counts toward the budget.

Retries & fallback

Tempr already retries transient upstream 429/5xx errors against the same provider before giving up, and — if you've configured a fallback chain on the key — falls back to the next candidate model automatically. When a provider asks for a wait longer than ten seconds, Tempr doesn't sleep through it: the 429 or 503 comes back at once with a Retry-After header carrying the provider's wait. A rate-limited model also sits out fallback chains until then, while that provider's other models stay available. A client-visible error means Tempr's own retry and fallback logic was exhausted, so honor Retry-After when it's there rather than retrying in a tight loop.

Provider errors

When the provider behind a request rejects it and Tempr's retries and fallbacks don't recover, you get the provider's HTTP status and its error message rather than one of the codes above. Tempr tidies a JSON error body into an error object with a message; a provider error that isn't JSON comes back as the provider sent it.

Errors during a stream

Once a streaming response has started, its status is already 200, so a failure after that point arrives in the stream. On /v1/chat/completions, a stream the provider cuts off before it finishes ends with this chunk, then data: [DONE]:

data: {"choices":[{"delta":{},"finish_reason":"error","index":0}],"error":{"message":"upstream stream closed prematurely","type":"stream_truncated"}}

Treat a finish_reason of "error" as a failed request: the output before it is incomplete.