Models, limits & logs
Last updated
Reference for everything around the core Chat Completions and Embeddings calls: listing models, rate limits and quotas, fallback chains, caching, and how request logging and content capture work.
Models endpoint
GET https://api.temprhq.io/v1/models
Authorization: Bearer tvk_...
Returns the standard OpenAI list shape, filtered to the intersection of your account's configured provider keys and this key's model allowlist. It lists chat models, each with "type": "language", so any client can fill a chat model picker from it:
{
"object": "list",
"data": [
{ "id": "openai/gpt-4o", "object": "model", "owned_by": "openai", "type": "language" },
{ "id": "anthropic/claude-opus-4", "object": "model", "owned_by": "anthropic", "type": "language" }
]
}
Add ?type=embedding to list the embedding models of the providers that have them instead, or ?type=all for both. Embedding models have "type": "embedding" and architecture.output_modalities of ["embeddings"]:
{
"id": "openai/text-embedding-3-small", "object": "model", "owned_by": "openai", "type": "embedding",
"architecture": { "input_modalities": ["text"], "output_modalities": ["embeddings"] }
}
Rate limits
Two independent ceilings apply to every request, and the stricter one wins:
- Key-level RPM/TPM — set per virtual key (see virtual key scoping).
- Plan-level RPM — fixed by your subscription tier (10/60/240 req/min on Free/Pro/Scale), shared across every key on the account.
- Organizations without a Gateway plan — organization keys get the Free plan's limits (10 req/min and 10,000 requests a month), shared across the organization.
A rate-limited request returns HTTP 429 with the code rate_limit_exceeded (requests per minute) or tpm_limit_exceeded (the key's tokens per minute); see limit errors. An app can also limit each of its own customers; see end-user limits.
End-user limits
When one key serves an app with many customers, Tempr can hold each customer, the app's end user, to their own budget and rate, so one heavy user can't use up the whole key.
Naming the end user
Send the end user's id in the x-tempr-end-user header, up to 256 characters (leading and trailing spaces are trimmed; a longer value is ignored). Without the header, Tempr uses the request's own end-user field: safety_identifier, else user, on /v1/chat/completions, /v1/responses and /v1/embeddings, and metadata.user_id on /v1/messages. A request that names no end user isn't limited by any of this.
curl https://api.temprhq.io/v1/chat/completions -H "Authorization: Bearer $TEMPR_KEY" -H "x-tempr-end-user: customer-4821" -d '{"model": "openai/gpt-4o-mini", "messages": [{"role": "user", "content": "Hi"}]}'
Limits per key
Limits belong to a key, so two apps' end users never share them. On a key, set a monthly budget in US dollars and a requests-per-minute limit that apply to each end user: in the Portal's key editor, or with end_user_limits in the management API. Then, for a particular end user, set their own budget or rate, or block them, on the Portal's End users page or with PUT …/end-users/{end_user}/limits. An end user's own value replaces the key's default; 0 means no limit for that end user. End users aren't created anywhere: they appear as requests name them.
- A blocked end user gets
403 end_user_blocked. - Past their requests per minute:
429 end_user_rate_limit_exceededwithRetry-After. - Past their monthly budget:
402 end_user_budget_exceeded. Spend counts from the start of the calendar month, UTC. If the key's budget is strict, the end user's is too. - A request that goes through while a budget applies carries
x-tempr-end-user-budget-remaining: what's left of the end user's budget, in US dollars.
The key's own limits, your plan's and your organization's budget still apply on top. See limit errors.
What providers see
Providers never see your end users' ids. Before a request is sent on, Tempr replaces each end-user field in it (user and safety_identifier, or metadata.user_id on /v1/messages) with a hash such as tu_3f9a…: 32 hex characters of an HMAC-SHA256 keyed with a secret Tempr derives for your account. The same id always gives the same hash for your account, so a provider's abuse monitoring still tells your end users apart, and another account's hash of the same id is different. Nothing else in the body changes, so signed thinking blocks on /v1/messages stay valid. The x-tempr-end-user header itself is never sent on. Tempr's own request logs and the End users page keep the ids as you sent them.
Monthly quota & overage
Each plan includes a monthly request allotment (Free 10,000 / Pro 250,000 / Scale 2,000,000). What happens past it depends on your plan:
| Plan | Past allotment |
|---|---|
| Free | Hard block at quota — no metering, no surprise charge. |
| Pro / Scale | Keeps serving on metered overage ($0.30/1k and $0.20/1k requests respectively) until you hit your own overage cap (default $50), then hard-blocks too. |
A request refused at the allotment returns 429 quota_exceeded with x-tempr-quota-used, x-tempr-quota-limit, and x-tempr-quota-resets. One refused at the overage cap returns 402 overage_cap_exceeded with the overage headers. Both counters reset each month: on a paid plan at the end of your billing period, which runs a month from the day the plan started (a new subscription starts a fresh period), and on Free at the start of the next calendar month, UTC. x-tempr-quota-resets gives the exact time; see limit headers. To hear about it sooner, set up budget alerts on the Portal's Gateway page: they fire at 80% and 100% of your quota, and at 80% of a metered overage cap.
Fallback chains
Define an ordered list of provider/model candidates for a model alias on a virtual key. If a candidate answers with a 429, 401, 402, 403 or 5xx, or can't be reached, Tempr tries the next one in the chain before anything reaches your client. For a streaming request that happens only before the first byte is sent: once a stream has started, an error during it is passed on as it is. Candidates with no provider key, outside the key's model allowlist, or resting after repeated failures are skipped.
# Fallback chain DSL (configured per key in the Portal)
fast: openai/gpt-4o-mini, anthropic/claude-haiku, z-ai/glm-4.6
Call the alias (fast above) as the model value; Tempr resolves it against the chain instead of a single fixed model.
Chain order: as written, cheapest or fastest
A chain tries its models in the order you wrote them, unless you say otherwise after the alias:
bulk (cheapest): groq/openai/gpt-oss-120b, cerebras/qwen-3.8-27b, openai/gpt-6-luna
chat (fastest): anthropic/claude-sonnet-5, openai/gpt-6-sol
| Order | Tries first |
|---|---|
| (none) | The models in the order written, as before. |
(cheapest) | The model that would cost least for this request: its prompt at each model's input price, plus its max_tokens (or the key's usual output length) at each model's output price. A model with no known price goes last. Models within 5% of each other keep the order you wrote. |
(fastest) | The model that has been quickest over the last 15 minutes: the median time to the first token for a streamed request, or the median time to the whole answer for one that isn't, across Tempr's traffic. Models within 10% of each other keep the order you wrote, and a model with too little recent traffic goes after the measured ones. One request in twenty uses the written order, so every model keeps being measured. |
Either way only your models are used, and a model that fails still hands over to the next. The response's x-tempr-fallback-order header says which order the request used (configured, cheapest or fastest), and the request log shows the ranking. Your quota is checked against the model tried first. In the Portal, the key editor shows each chain's current ranking. Through the management API, a chain with an order is { "order": "cheapest", "models": [...] }; a plain list still means as written.
Chains work on embeddings too. There, make the candidates the same model on different providers (say azure-openai/text-embedding-3-small then openai/text-embedding-3-small): vectors from different models can't be compared, so falling back to another model would put incompatible vectors in the same index.
Response caching
Opt a single request into exact-match caching with the x-tempr-cache-ttl header (see Chat Completions). Caching is scoped per account (or per organization for org-owned keys) — a hit on one key can be served from another key's cache write under the same account, but never across accounts. It doesn't apply to /v1/embeddings, which ignores the header.
Prompt caching
Unlike response caching, prompt caching still calls the model — but the provider serves the repeated part of your prompt (tools, system prompt, earlier turns) from its own cache, and on Anthropic models that cached input bills at about a tenth of the normal input price. OpenAI, Gemini, DeepSeek, and xAI already cache long prompts automatically with no request changes; Anthropic only caches at explicit cache_control breakpoints, which Tempr places for you.
With prompt caching on, a request to an Anthropic model gets:
- Breakpoints on the stable prefix — the last tool definition and your system prompt, so the part every request shares is written once and read back after that.
- A breakpoint on the conversation tail once a conversation has history, so each turn reads everything the previous turn wrote. A first-turn request isn't marked there — a one-off question would only pay the write premium.
- Deterministic tool order — tools are sent sorted by name, so the same tool set always forms the same prefix.
- Key affinity — with more than one key for a provider, requests sharing a model, tools, and system prompt all go to the same key. Provider caches are separate per account, so spreading one prompt across keys would pay to write it several times.
A request that already sets its own cache_control anywhere is sent exactly as written — clients that manage their own breakpoints, like Claude Code on /v1/messages, keep them intact whether or not the setting is on.
| Control | Where |
|---|---|
| Account default — Off, On (5-minute cache), or On (1-hour cache) | Portal → Gateway → Settings |
| Per-key override — inherit, on, or off | Portal → Gateway → Keys |
| Per request | x-tempr-prompt-cache: on | off | 5m | 1h (see Chat Completions) |
Writing to the cache costs more than plain input — 1.25× for the 5-minute cache, 2× for the 1-hour cache — so it pays off as soon as a prefix is reused. Pick 1 hour when requests sharing a prefix arrive more than 5 minutes apart; the conversation tail always uses 5 minutes. Prompts shorter than the model's minimum cacheable length (512–4,096 tokens, depending on the model) simply aren't cached, and cost nothing extra.
Each request log shows the tokens read from and written to the cache and the estimated savings after write premiums, and Analytics rolls them up for the account. Token counts in usage and logs cover the whole prompt, cached tokens included — key-level TPM limits count them the same way.
Remote MCP proxy
POST https://api.temprhq.io/mcp/proxy/{serverSlug}
Authorization: Bearer tvk_...
Relays MCP Streamable HTTP requests to a remote MCP server you've added in the Portal, named by its slug, so an MCP client can reach that server with a virtual key while Tempr connects with the details saved for it. If the key has an MCP server allowlist, the server has to be on it. The key's rate limits, quota, and budgets apply, and each call appears in your request logs with mcp as the provider and the slug as the model.
The server's response comes back as the server sent it, status included, except where a guardrail below changes it. A server Tempr can't reach returns 502 with the body {"error": "Could not reach the configured MCP server. Try again in a moment."}. Tempr's own refusals use the error codes in the OpenAI format.
MCP guardrails
The relay checks tool traffic in both directions. Set the checks in the Portal under MCP credentials, for your account or your whole organization; a virtual key can override them. The defaults are in brackets:
- Tool-call arguments (flag): secrets and personal data (bearer tokens, JWTs, AWS and
sk-keys, email addresses, card-like numbers) in what's sent to a tool. Flag records them in the request log; Block answers the call with a JSON-RPC error and never sends it. - Tool results (redact): the same kinds of data in what a tool sends back. Redact masks them before the model reads the result, as
[REDACTED-SECRET]and so on; Block withholds the whole result and returns a tool error instead. - Planted instructions (flag): text in tool results and tool descriptions shaped like instructions to the model ("ignore previous instructions", fake system prompts), and hidden characters (Unicode tag characters, zero-width spaces, bidi overrides), which are removed unless this check is off. Block withholds the result, or drops the tool from the server's tool list.
Results are checked in JSON or as a server-sent event stream, up to 256 KB per message; a larger message passes unchecked unless results are set to Block, in which case it's withheld. These checks catch common shapes. They aren't a data-loss-prevention product, and they don't make prompt injection impossible.
To see what these settings do before you rely on them, use the guardrail tester on the MCP guardrails card: it shows the result as the model would get it, with hidden characters, redactions and planted instructions marked.
Tempr also pins each tool's name, description and input schema the first time it sees them. When a server changes a tool's definition, the change is recorded in your organization's audit log and sent to your Gateway webhook as mcp.tool_changed, with the old and new description. The tool keeps working.
Per-member sign-in
A remote server can sign each member in with their own account instead of a header everyone shares. Choose As each member, with their own account under the server's Sign-in settings in the Portal. Tempr then registers itself with the server's OAuth provider: by a Client ID Metadata Document where the provider supports one, by dynamic client registration otherwise, or with an OAuth app you register yourself (GitHub, for example). Each member signs in from the Portal, or from the link their editor shows on the first tool call, and their calls run as them.
- Tokens are encrypted and stay on Tempr's servers. The relay sends the member's token and refreshes it when it expires.
- Until a member signs in, the relay still answers the server's tool list (as last seen by any member), and a tool call asks them to sign in with MCP's own
-32042URL elicitation. - Owners and admins see who's signed in, with which account and scopes, and can revoke one member or everyone. Members manage their own sign-ins under MCP sign-ins. Leaving the organization or closing the account revokes them.
- A key that belongs to an organization with no member behind it can't reach such a server (
mcp_connection_requires_user).
Tool Packs
An organization can give each person only the MCP tools their job needs. A Tool Pack is a named set of tools: for each of the organization's servers, all of its tools, all except some, or only the ones you pick. Owners and admins build packs under Tool Packs in the Portal, and give them to groups, to members, or to the organization's keys.
- Nothing changes until the first pack. An organization with no pack keeps today's access: everyone can use every tool. Turning packs on offers a default pack with every tool, so nobody loses access until you narrow it.
- A member's tools are every tool in every pack given to them, directly, through their groups, or through the key they use. A default pack covers members and keys with no other pack. There are no deny rules: "all tools except" leaves a tool out of that one pack, and another pack can still give it.
- Groups are made in the Portal, or mirrored from your company directory (directory groups), so someone who changes team gets the new team's tools without anyone editing them.
- Where it holds: a tool outside someone's packs is left out of the server's tool list, and a call to it is refused before it reaches the server, with the JSON-RPC error
-32601"Tempr: this tool isn't in your organization's Tool Packs for you." (logged asmcp_tool_not_allowed). A server with none of its tools in someone's packs is refused like the key's server allowlist (mcp_server_not_allowed) and isn't handed to their Tempr clients at all, local servers' settings included. In agent turns that run on Tempr's servers, those tools are never offered to the model. IDE agent turns run on the server for organizations that use packs. - Servers your organization doesn't manage (a member's own) are allowed unless you switch on Only allow your organization's servers.
- A key's own MCP server allowlist still applies on top: it can narrow what its packs give, never widen it.
- Packs, groups and assignments can also be managed through the management API. Every change is in the audit log.
Asking for a tool
A refusal doesn't have to be the end. When a tool on one of the organization's servers isn't in a member's packs, the refusal carries a link to a Portal page where they can ask for it, with a reason of up to 500 characters, signed in to the Portal as themselves. The link is in the refusal itself, so any MCP client can show it; Tempr's own clients add a Request access button that opens it (the CLI prints the link) from their next releases.
- Through the relay, the JSON-RPC error keeps its code and message and adds
"data": { "requestAccessUrl": "https://portal.temprhq.io/toolpacks/request?server=github&tool=create_pull_request" }. In agent turns that run on Tempr's servers, the tool's result ends with aRequest access:line. - Owners and admins are emailed about each request, at most once an hour per organization: requests that arrive in between come together in the next email. The Tool Packs menu item shows how many are waiting.
- Deciding: the Tool Packs page's Requests tab lists each request with the member's groups and packs. Approve it for 1 hour, 1 day, 1 week or with no limit, or decline it with an optional note; the member is emailed either way. An approval gives that member that one tool beside their packs, which stay as you designed them, and stops working at its time. Approvals in effect can be revoked at any time. When three or more members ask for the same tool within 30 days, the tab suggests adding it to a pack.
- A member has one pending request per tool; asking again updates the reason. Every request, approval, decline, revocation and expiry is in the audit log, and the requests and grants can be read and decided through the management API.
Built-in tools such as the shell, file edits and web fetch aren't part of Tool Packs; personas and approval modes govern them.
Request logs
Every Gateway call is logged with: timestamp, virtual key, model, provider, HTTP status, latency (time-to-first-byte and total), token counts, cost, and finish reason — metadata is always captured. Prompt and response content is captured only when content capture is enabled. Logs are searchable and filterable in the Portal by key, model, provider, status, latency range, date, and any x-tempr-metadata keys you sent. An embeddings request logs its input as the request content, and its response with each vector replaced by a note of its size.
Two more views group the same logs. Send x-tempr-session-id to see a conversation's or agent run's requests together under Sessions, and an end user, in x-tempr-end-user or the request's own field, to see usage per customer and key under End users, with each end user's spend this month against their limit.
Analytics
The Portal's Dashboard charts requests, cost, error rate, latency (median and p95), time to first byte and tokens over the range you pick, and breaks them down by key, provider and model, from the same logged data. Sessions, End users and Metadata group the same traffic by session, by customer, and by each value of an x-tempr-metadata key.
Alerts & reports
Every plan is emailed at 80% and 100% of its monthly requests, and at 80% of the overage cap on a metered plan. On Pro and up, alert rules watch your traffic and tell you when it crosses a line:
- A rule watches the error rate, cost, cost against the usual (the window's cost as a multiple of the average for a window that long over the previous 7 days), p95 latency, or the number of requests. It looks back over the last 15 minutes, 1 hour, 6 hours or 24 hours, at all your keys or at one key, provider or model.
- It alerts when the value is at least its threshold, or, for requests, when it falls below, which catches traffic that stopped. Error rate and p95 latency wait for a minimum number of requests (20 unless you change it).
- It notifies once when it's crossed and once when it recovers, by email (to you, or to a team's owners and admins), by your webhook, or both. Playground requests aren't counted. You can have up to 20 rules.
Below Pro, two fixed webhook alerts can be turned on in Settings instead: error rate (25% or more of at least 20 requests in 15 minutes, at most hourly) and cost anomaly (an hour's spend at 3 times the usual hourly spend and at least $1, at most every 6 hours). On Pro and up they become the first two rules.
The weekly report emails each week to Monday 00:00 UTC: requests, error rate, cost, tokens and latency, the top models and keys by cost, and, where the plan's log retention still covers it, the change from the week before. The daily digest emails the last 24 hours. Turn either on or off in Settings.
Alert webhooks
Each event is a POST of JSON to the webhook URL in your Gateway settings, with an x-tempr-webhook-signature: sha256=<hex> header: an HMAC-SHA256 of the raw body, keyed with the signing secret shown when you set the URL. A 5xx or a network error is retried twice; a 4xx isn't.
{
"event": "gateway.alert.triggered",
"subscriptionId": "…",
"planTier": "Pro",
"timestamp": "2026-09-30T19:16:02.4817730Z",
"data": {
"ruleId": "…",
"name": "Checkout latency",
"condition": "p95 latency is at least 2,500 ms over 1 hour",
"metric": "p95_latency_ms",
"direction": "at_least",
"threshold": 2500,
"windowMinutes": 60,
"value": 3180,
"valueText": "3,180 ms",
"filters": { "virtualKeyId": null, "provider": null, "model": "anthropic/claude-sonnet-5" },
"dashboardUrl": "https://portal.temprhq.io/gateway?range=24h&model=anthropic%2Fclaude-sonnet-5",
"alertsUrl": "https://portal.temprhq.io/gateway/alerts"
}
}
| Event | When |
|---|---|
gateway.alert.triggered | A rule is crossed. metric is error_rate_percent, cost_usd, cost_multiple_of_usual, p95_latency_ms or requests; direction is at_least or below. |
gateway.alert.resolved | It recovers. The same data; value is null when there wasn't enough traffic left to measure. |
gateway.quota_alert | 80% (stage 1) and 100% (2) of the monthly requests, and 80% of the overage cap (3), with quotaCurrent and quotaLimit. |
gateway.error_rate_alert, gateway.cost_anomaly_alert | The fixed alerts, below Pro. |
gateway.test | The Send test webhook button in Settings. |
Guardrails on requests
Before a request goes to a provider, the gateway scans it for secrets and personal data: bearer tokens, JWTs, AWS and sk- keys, email addresses and card-like numbers. Set the mode in the Portal under Gateway settings, for your account or your organization; a key can override it. The modes:
| Mode | What happens |
|---|---|
| Off | Nothing is checked. |
| Flag only | The request goes as sent, and the categories found are recorded in the request log. |
| Redact secrets | JWTs, bearer tokens, AWS access keys and sk- secrets in the request's message text are replaced with placeholders before the model sees them, and the rest goes through. See redaction. |
| Redact secrets and personal data | The same, and email addresses and card-like numbers too. |
| Block | The request is refused with guardrail_blocked. |
The optional LLM guardrail also asks the request's own model, on your own provider key, whether the request tries to override its instructions; under Block an UNSAFE answer refuses the request, and a check that fails lets it through. Under the other modes an UNSAFE answer is recorded in the request log.
Redaction
A redact mode masks matches with the same placeholders the request log uses when redaction of stored logs is on: [REDACTED-JWT], Bearer [REDACTED-TOKEN], [REDACTED-AWS-KEY], [REDACTED-SECRET], [REDACTED-EMAIL] and [REDACTED-CARD]. The same text always becomes the same bytes, so prompt caching keeps working across turns. Values aren't put back into the answer: the model works with the placeholder. A pasted .env no longer refuses an agent's turn; the keys in it just never leave Tempr.
Only message text is masked, never model ids, tool definitions, ids, images or signatures:
| Endpoint | Message text |
|---|---|
/v1/chat/completions | Each message's content (a string, or its text parts), for every role including tool results; tool_calls[].function.arguments. |
/v1/messages | system; text blocks; the string values inside a tool_use block's input (not its keys); tool_result content; text documents and search results. Never thinking or redacted_thinking blocks: Anthropic signs them, and a changed one is refused. |
/v1/responses | instructions; input as a string or its message items' text parts; function_call arguments and function_call_output output; custom tool calls' input and output. Never reasoning items. |
/v1/embeddings | The input strings. |
The masked request is what everything after the check sees: the response cache key, the LLM guardrail, the provider, and the request log. A secret left outside message text, in a tool's schema for example, is flagged in the log rather than masked. Redaction is pattern matching: it catches the common shapes of secrets and personal data, not every one, and it isn't a data-loss-prevention product. The card pattern matches any 13 to 19 digit run, so long numbers in message text are masked under Redact secrets and personal data.
The same setting covers your organization's members' own turns in the IDE extensions, the CLI, Tempr Code and agent runs on the server, which don't use a Gateway key: see the prompt guardrail for members.
What the guardrail did: x-tempr-guardrails
When anything matched, the response carries x-tempr-guardrails with counts by category, never the values: redacted=aws-key:1,email:2; flagged=jwt:1. redacted lists what was masked, flagged what was found and let through, and blocked what refused the request. Categories are jwt, bearer-token, aws-key, secret, email and card. The request log records the same categories.
Test them first. Each guardrail setting in the Portal (your Gateway settings, your organization's, the MCP guardrails and each key's edit form) has a Test your guardrails panel. Paste a request, a tool call, a tool result or a tool description and see every match highlighted, the redacted text and the verdict in words, for the settings as they are on the page, saved or not. It runs the same checks the gateway and the MCP relay run. The text is checked and thrown away: it isn't stored or logged. The LLM check is off in the tester until you tick it, since it's a real call to the model you pick on your own provider key, billed by your provider; that call sends x-tempr-guardrail-test, which the gateway honours only from the Portal, so the text isn't captured. For CI, the management API has the same test without the LLM check.
Timing: Server-Timing
Every Gateway response carries a standard Server-Timing header, which browser developer tools show as a timeline: tempr;dur=12.4, guardrails;dur=1.1, upstream;dur=834.0, in milliseconds.
| Metric | What it measures |
|---|---|
tempr | Tempr's own time before the provider was called: authentication, limits, guardrails and routing. For a response no provider was called for, such as an error or a cache hit, the time until it went out. |
guardrails | The guardrails' share of that: the pattern check, redaction and the LLM guardrail. Present only when one ran. |
upstream | From the first provider call until the response's headers went out. On a stream that's the provider's first byte; on a JSON answer, the whole answer, since Tempr reads it before replying. Present only when a provider was called; after a fallback it includes the attempts before the one that answered. |
The header is written as the response starts, so it holds what's known at that moment; the request log has the full duration and time to first byte.
Reasoning policy per key
A key can set how much its requests reason, in the key's edit form in the Portal or reasoning_policy in the management API: a default level for a request that names none, and a lowest and highest level. The levels are the Gateway's own, none, minimal, low, medium, high, xhigh and max (see reasoning). A request below the lowest level or above the highest moves to the nearest allowed one, and then, as always, to the nearest level the model supports. A request's level is its effort, none when it switches reasoning off, or the level a thinking budget roughly corresponds to; reasoning switched on with no level is only moved when the highest level is none.
x-tempr-reasoning-effort reports the level sent, and x-tempr-reasoning-policy says why it differs from the request: default when the key's default was used, or moved; requested=xhigh when the request's level was outside the range. Requests that go to a provider's own API as sent (OpenAI and Azure OpenAI on /v1/responses, Claude on /v1/messages) get the policy in their own fields: reasoning.effort, or Claude's output_config.effort, thinking budget or thinking switch. A request to Claude that sets no thinking keeps Claude's own default, and a Responses request gets the default only on a model that reasons.
Content capture & privacy
Content capture is on by default so real request/response detail is available for debugging: 7-day retention on Free, longer on paid plans, truncated at 256 KB per request. It's a per-account toggle — switch to metadata-only mode at any time and nothing but metadata is stored going forward. For a Teams organization's keys the setting is the organization's, and its owners and admins can read captured content on the Team request logs page; each request they open is recorded in the organization's audit log. This is separate from the Chat product, whose synced chats have their own switch (see the Privacy Policy).
| Plan | Log retention |
|---|---|
| Free | 7 days |
| Pro | 30 days |
| Scale | 90 days |
| Enterprise | Set by your contract |
Log streaming
On the Enterprise plan, an organization can stream its Gateway request logs, agent traces and audit log to its own tools as they happen. Owners and admins set it up in the Portal under Organization → Log streaming, with up to five destinations. A destination gets what happens from the moment it's added, not what came before.
| Destination | What it takes |
|---|---|
| OpenTelemetry | Request logs and agent traces as spans, and the audit log as log records, over OTLP/HTTP with JSON bodies, gzipped. Anything that accepts that works, including the OpenTelemetry Collector, which can forward to any backend. |
| Signed webhook | The audit log only, as signed JSON batches to any HTTPS address. |
Headers you add to a destination, usually your backend's API key, are sent with every request and stored encrypted. You can replace them without losing the destination's place. Addresses that resolve to private or internal networks are refused.
OpenTelemetry
The endpoint is a base URL, as with OTEL_EXPORTER_OTLP_ENDPOINT: spans go to {endpoint}/v1/traces and log records to {endpoint}/v1/logs. Every record carries the resource attributes service.name (tempr) and tempr.organization.id. Names follow the OpenTelemetry GenAI conventions where one fits.
| Source | Shape | Main attributes |
|---|---|---|
| Request logs | One span per request, in its own trace whose id is the request id from x-tempr-request-id. Named chat {model} or embeddings {model}; an MCP proxy call is mcp {server}. | gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.input_tokens and output_tokens, gen_ai.conversation.id (your session id), http.response.status_code, error.type, tempr.cost_usd, tempr.cache_hit, tempr.end_user.id, tempr.metadata (your x-tempr-metadata) |
| Agent traces | A span tree per agent run, whose trace id is the run id: the run, each model call and each tool call, with a delegated sub-agent's calls under it. | gen_ai.operation.name (invoke_agent, chat, execute_tool), gen_ai.agent.name, gen_ai.request.model, gen_ai.tool.name, token counts, tempr.cost_usd, user.id |
| Audit log | One log record per entry, with event.name tempr.audit and log.record.uid audit-{id}. | tempr.audit.action, tempr.audit.actor.kind, tempr.audit.actor.email, tempr.audit.target.type and target.id, tempr.audit.details |
A failed request's span has an error status with its message. Request and response bodies are left out unless you choose to include them when adding the destination, and then only where content capture kept them, as tempr.request.body and tempr.response.body, after any redaction.
Signed webhook
Each batch is a POST of JSON like this:
{
"type": "tempr.audit_log",
"batch_id": "audit-1041-1043",
"organization_id": "7d2c…",
"events": [
{
"id": 1041,
"created_at": "2026-09-29T14:03:11.208Z",
"action": "VirtualKeyCreated",
"actor": { "kind": "user", "id": "4b1e…", "email": "ana@example.com", "label": null },
"target": { "type": "VirtualKey", "id": "a90f…" },
"details": { "label": "ci" }
}
]
}
When you add the destination, the Portal shows its signing secret once. Every batch has a Tempr-Signature header, t={unix seconds},v1={signature}, where the signature is the hex HMAC-SHA256 of {t}.{raw body} keyed with the whole secret. Check it against the raw body before parsing it, and turn away batches whose t is more than five minutes old:
import hashlib, hmac, time
def verify(secret: str, header: str, body: bytes) -> bool:
parts = dict(p.split("=", 1) for p in header.split(","))
expected = hmac.new(secret.encode(), f"{parts['t']}.".encode() + body, hashlib.sha256).hexdigest()
fresh = abs(time.time() - int(parts["t"])) <= 300
return fresh and hmac.compare_digest(expected, parts["v1"])
Answer with any 2xx once you've stored the batch. "Test" in the Portal sends an empty batch with "type": "tempr.test".
Delivery
- Timing: within seconds of the activity, in batches of up to 500 (20 with bodies included).
- At least once: a batch can arrive twice, if an answer is lost. It keeps its ids when it's resent (trace and span ids,
log.record.uid, each event'sidand theTempr-Batch-Idheader), so deduplicate on them. - Failures: Tempr retries with a growing delay, up to 15 minutes, and waits longer if you send
Retry-After. A413makes the batches smaller. An OpenTelemetry endpoint's400means the data can't be taken, so, as the OTLP spec says, that batch is skipped and counted in the Portal. - Stopping: if a destination answers
401,403,404,405or415five times in a row, or keeps failing for a day, Tempr stops sending to it and emails your owners and admins. Once it's fixed, resume it in the Portal: it catches up from where it stopped, as far back as your log retention reaches. - After it's sent: what a destination received stays there. Removing the destination, or closing the organization, doesn't delete it from your tools.