Chat Completions
Last updated
POST /v1/chat/completions — the same request and response shape as OpenAI's API, so any OpenAI-compatible SDK or library works against Tempr Gateway without a rewrite.
Request
POST https://api.temprhq.io/v1/chat/completions
Authorization: Bearer tvk_...
Content-Type: application/json
Body parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | provider/model, e.g. openai/gpt-4o. Must be on your virtual key's allowlist. |
messages | array | Yes | Standard OpenAI message array (system/user/assistant/tool roles). |
stream | boolean | No | Default false. See Streaming. |
temperature, top_p, max_tokens | number | No | Passed through to the upstream provider as-is. |
tools, tool_choice | array / string | No | Function/tool-calling definitions, forwarded to providers that support them. See Tool calling. |
reasoning | object | No | How much the model reasons: effort (none to max), max_tokens (a thinking budget), enabled, exclude. Works on every lab. See Reasoning. |
reasoning_effort | string | No | OpenAI's name for reasoning.effort; the nested value wins if you send both. |
stream_options | object | No | If you set stream: true, Tempr automatically injects include_usage: true so streaming responses still carry a final usage chunk. |
Request headers
| Header | Description |
|---|---|
Authorization | Bearer tvk_... — required. See Authentication. |
x-tempr-metadata | Optional JSON object, size-capped, stored against the request log for filtering — e.g. {"customer_id":"acct_123"}. The Portal's Metadata page groups your requests by any of its keys, with each value's cost, error rate and latency. |
x-tempr-cache-ttl | Optional integer seconds (1–86400). Opts this request into response caching. |
x-tempr-prompt-cache | Optional: on, off, 5m, or 1h. Overrides your account's prompt caching setting for this request; 5m/1h also choose the cache lifetime. |
x-tempr-session-id | Optional string, up to 256 characters. Groups related requests, such as one conversation or agent run, into a session in the Portal's Sessions view. A longer value is ignored. |
x-tempr-prompt | Optional JSON object naming one of your prompt templates, e.g. {"slug":"support-bot","version":3,"variables":{"name":"Ada"}}. The template's content, with its {{name}} variables filled in, goes in front of the request as the first system message. version is optional and defaults to the template's current version. A malformed header or an unknown template is rejected with 400; see prompt template errors. |
Response (non-streaming)
The raw OpenAI-shaped completion object — Tempr does not wrap or reshape it:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "openai/gpt-4o",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hi there!" },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 4,
"total_tokens": 16
}
}
Every response also carries an x-tempr-request-id header, which correlates 1:1 with the request in your request logs, and a Server-Timing header with Tempr's own time and the provider's (see timing). When the guardrails matched anything, x-tempr-guardrails says what they redacted, flagged or blocked, as counts by category.
Streaming
Set "stream": true and Tempr forwards the upstream provider's chat.completion.chunk Server-Sent Events straight through — non-OpenAI providers (Anthropic, Gemini, Cohere) are translated into the same OpenAI-shaped chunks per-line, so your SSE parsing code never needs to know which provider actually served the request.
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hi"}}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there!"}}]}
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":12,"completion_tokens":4,"total_tokens":16}}
data: [DONE]
If the provider's stream ends before a finish reason arrives, Tempr closes it with a chunk whose finish_reason is "error" and whose error.type is stream_truncated, then data: [DONE]. The status is already 200 by then, so check for it; see errors during a stream.
Refusals
A model that declines to answer still returns 200. On every lab, a declined answer has finish_reason "content_filter", as in OpenAI's API. When the lab says why, the reason comes in OpenAI's refusal field: message.refusal, or delta.refusal when streaming. Claude's refusals carry it (its stop_details.explanation); Gemini's safety, recitation and blocked-prompt stops don't.
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": null, "refusal": "This request was declined because it could enable cyber harm." },
"finish_reason": "content_filter"
}
]
Treat any text streamed before a refusal as incomplete. A refusal is never stored in the response cache, so retrying works, and Anthropic's advice for Claude's refusals is to retry on another model. On /v1/messages a refusal is stop_reason "refusal" with stop_details. On /v1/responses it is a refusal content part in an incomplete response.
Tool calling
Pass tools and (optionally) tool_choice exactly as you would to OpenAI's API. Tempr forwards them to providers that support function calling and returns tool_calls on the response message in the same shape:
{
"model": "openai/gpt-4o",
"messages": [{ "role": "user", "content": "What's the weather in Boston?" }],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}
]
}
Tempr forwards tool definitions and tool-call results — it doesn't execute tools itself. You own the loop: read the model's tool_calls, run the function on your side, and send the result back as a tool message on the next request.
Reasoning
Set how much the model thinks before it answers with one shape on every model: OpenRouter's reasoning object, or OpenAI's reasoning_effort for just the level. Tempr sends it on in the lab's own format: reasoning_effort for OpenAI, xAI, and Mistral; adaptive thinking and output_config.effort for Claude (a thinking budget before Claude 4.6); thinkingConfig for Gemini; thinking for Cohere, DeepSeek, Z.AI, Moonshot, MiniMax, and Xiaomi; enable_thinking for Qwen.
A model served by an inference host gets the host's format rather than its lab's, because that is what the host reads, and the levels come from that host's own catalog entry — so the same model can offer different levels, or none, on different hosts. DeepInfra, AWS Bedrock's OpenAI-compatible models, Fireworks AI, Groq and Cerebras take reasoning_effort; Bedrock's Converse models take the fields their model family takes, thinking with output_config.effort for Claude and reasoningConfig for Amazon Nova; Together AI takes its own reasoning object; Novita AI takes reasoning_effort for levels and thinking to switch thinking on or off, which is the only switch it honours. A few models report no levels because their host can't steer them reliably: Together AI's GLM reads reasoning_effort backwards, so it is left to reason at its own depth. Hugging Face reports none at all — it routes one model id to whichever backing host it picks, and those hosts disagree about what a level means, so a picker there could not be honoured. Perplexity reports none either: each of its models is an Agent API preset that sets its own reasoning, so a deeper preset is the way to ask for more (see Perplexity).
{
"model": "anthropic/claude-sonnet-4-6",
"messages": [{ "role": "user", "content": "Plan the migration." }],
"reasoning": { "effort": "high" }
}
| Field | What it does |
|---|---|
effort | none, minimal, low, medium, high, xhigh, or max. Each model supports some of these; a level it doesn't have becomes the nearest one it does, rounding up on a tie. |
max_tokens | A thinking-token budget, for models that take one (Claude before 4.6, Gemini 2.5, Cohere, Qwen). On a model that only takes levels it becomes the closest level. Can't be combined with effort. |
enabled | false turns reasoning off, the same as effort: "none". true turns it on at the model's own depth. |
exclude | Passed on to OpenRouter models. |
Leave both fields out and every model reasons at its own default. Some models can't turn reasoning off (Claude Fable, Gemini 3, GLM-5.3, and Kimi K3 on some inference hosts); none gets them their lowest level instead. Claude Opus 5 gets low too, because Anthropic recommends a low effort over switching its thinking off. The response's x-tempr-reasoning-effort header says what was applied: the level sent, none, on, budget, or default when the model has no reasoning control. If the provider still rejects the level, Tempr retries once at the nearest level the provider lists, or with no reasoning setting when it lists none, and the header reports what the retry sent (default for no setting). If you also send a lab's own reasoning fields, such as Claude's thinking or DeepSeek's thinking, they go through as written and reasoning is ignored for that lab. One case sets its own level rather than yours: Amazon Nova 2 refuses high effort alongside max_tokens, so a request that caps its output runs at medium, which the header reports. A key can also have a reasoning policy: a default level and a range, which x-tempr-reasoning-policy reports when it changes a request's level.
A model's reasoning comes back as reasoning text and reasoning_details on the message (on delta when streaming), as it does from OpenRouter. Where a provider names it reasoning_content (DeepSeek, Kimi, MiMo and GLM among them), that field comes back too. When a Claude model calls a tool, send the assistant message back with its reasoning_details unchanged: Tempr passes them to Claude as its thinking blocks, which a tool-calling turn needs.
Responses API
POST /v1/responses serves OpenAI's Responses API — the format Codex CLI and newer OpenAI SDK code use — with the same virtual key, allowlist, limits, and request logs as chat completions. Streaming sends the standard Responses events: response.created, response.output_text.delta, and so on through response.completed.
- OpenAI models (
openai/…) and Azure OpenAI deployments (azure-openai/…) pass straight through to their own Responses API, so reasoning items, hosted tools, andprevious_response_idall work as they do against OpenAI directly. - Every other model is translated to its provider's API and back.
instructions,inputmessages, function tools, custom (freeform) tools,tool_choice,max_output_tokens, andtext.formatJSON schemas carry over, and prompt caching applies as it does to chat completions. So do tools grouped in anamespace, whose calls come back with theirnamespace, and a client-runtool_search("execution": "client"): the model's searches come back astool_search_callitems, and the tools yourtool_search_outputreturns can be called from then on. That's how Codex CLI's sub-agents work on any model.
| On a non-OpenAI model | What happens |
|---|---|
Hosted tools (web_search, file_search, code_interpreter, OpenAI's own tool_search, …) | Dropped, and named in the x-tempr-unsupported-tools response header. The rest of the request still runs. |
previous_response_id | Rejected with 400 previous_response_id_unsupported: Tempr doesn't store responses. Send the conversation in input instead. |
reasoning.effort | Sent on in the model's own format, as on chat completions. |
reasoning.summary, store, include, truncation, service_tier | Ignored. |
Fallback chains and backup provider keys work here the same way they do for chat completions, and a chain can mix OpenAI and other models. x-tempr-cache-ttl response caching doesn't apply to this endpoint yet. For a working client setup, see the Codex CLI guide.
Messages API
POST /v1/messages serves Anthropic's Messages API, the format Claude Code and clients built on Anthropic's SDK use, with the same virtual key, allowlist, limits, and request logs as chat completions. Send the key as Authorization: Bearer tvk_… or in the x-api-key header.
- Anthropic models (
anthropic/…) pass straight through to Anthropic, so extended thinking, server tools, and beta features work as they do against Anthropic directly. - Every other model is translated to its provider's API and back, streaming included, so a Claude Code session can run on any model you have a key for.
thinkingandoutput_config.effortbecome that model's own reasoning settings, and the reasoning a model returns comes back as athinkingblock (thinking_deltaevents when streaming).
When Tempr translates Claude, its thinking blocks keep Claude's own signature. That covers Claude on AWS Bedrock, OpenRouter's anthropic/… models, and Anthropic models when a request asks for response caching, since those requests are translated rather than passed through. Redacted thinking comes back as a redacted_thinking block. On Claude 4.7 and later a thinking block's text is empty, because Anthropic omits it by default, but its signature still holds the reasoning. Send the blocks back unchanged, as Claude Code does, and Claude gets them again on the next turn, so its reasoning carries through a tool call.
Other models don't sign their thinking, so a thinking block from one of them carries the signature tempr:reasoning. Send it back unchanged too. DeepSeek, Z.AI, Moonshot, and Xiaomi MiMo models, and the hosts that serve them, get its text back as reasoning_content, which they need across tool calls. Models that take their reasoning back as their own signed entries instead — Gemini 3, MiniMax and Kimi through OpenRouter — are a known gap: their thinking is readable in the answer, but it isn't sent back, because a thinking block can't rebuild a signed entry. They answer the next turn reasoning afresh rather than failing. Models that don't take earlier reasoning back don't get it. When a later turn goes to a Claude model, Tempr leaves these blocks out, because Claude would reject their signature, and keeps Claude's own.
Tempr's own errors come back in Anthropic's error format; see error shape. Fallback chains, backup provider keys, and response caching apply, though a request that asks for response caching doesn't use fallback chains or backup keys. On a translated live stream, message_start reports usage.input_tokens as 0. For a working client setup, see the Claude Code guide.
Response caching
Add x-tempr-cache-ttl: <seconds> to opt a request into exact-match caching (same model, messages, and sampling params). A cache hit skips the upstream call entirely, streams back identically to a live response, and is logged and billed as a cached request rather than a full-price one. Caching requires content capture to be enabled on your account — see content capture.
It works on /v1/chat/completions and /v1/messages. When a request asked for caching, the response's x-tempr-cache header says whether it was served from the cache: HIT or MISS.