One endpoint. Your keys. Full observability.
An OpenAI-compatible API in front of the provider keys you already have. Point your existing SDK at one endpoint and get virtual keys, request logs, and analytics — no IDE license required.
Drop-in for your existing SDK
Same request and response shape as OpenAI's API — swap the base URL and key, nothing else in your code changes. Python, Node, and every other OpenAI SDK work exactly the same way.
# Point any OpenAI SDK at Tempr Gateway
curl https://api.temprhq.io/v1/chat/completions \
-H "Authorization: Bearer tvk_..." \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o",
"messages": [{"role": "user", "content": "Say hi"}]
}'
Full Python and Node examples in the Gateway quickstart.
See it in action
The Gateway side of the Tempr Portal: what your keys are doing, what each request cost, and what every key is allowed to do.
Screenshots of the Tempr Portal running against a demo account.
Built for teams who already bring their own keys
Virtual keys, request logs, and analytics — BYOK all the way through, so Tempr never marks up or resells a token.
Bring your own keys
Connect your own first-party provider keys — OpenAI, Anthropic, Google, xAI, and more. Tempr fronts none of its own: you pay your provider directly, and Tempr for the gateway itself.
Virtual keys
Issue scoped keys per app or environment with a model allowlist, a budget, and RPM/TPM limits — revoke instantly, no rotating your real provider key.
Request logs
Every call is searchable by key, model, provider, status, and latency, with a per-request detail view — metadata always included, prompt/response when you opt in.
Analytics
Requests, tokens, cost, latency (p50/p95), and error rate — broken down by key, model, and day.
Quotas, not surprises
Each plan includes a monthly request quota. Paid tiers keep serving past it on metered overage, capped at a limit you set — never a silent bill shock.
Fallback chains
Define an ordered list of provider/model candidates per key, so one provider's outage or rate limit doesn't take your app down with it. A chain can try the cheapest model for each request first, or the fastest right now, and still hand over when one fails.
Prompt caching
Switch it on and Anthropic requests get cache breakpoints on your tools, system prompt, and conversation history, so repeated context bills at about a tenth of the input price. Requests sharing a prompt stay on one provider key so they read the same cache, and every request shows its cached tokens and savings.
Response caching
Opt in per request to exact-match caching for repeat prompts — cached hits are logged and billed as cached, not full price.
Content capture, your call
Request/response logging is on by default (7-day retention, 256 KB truncated) so you can debug real traffic — toggle it off, or switch to metadata-only, anytime.
OpenAI-compatible, end to end
POST /v1/chat/completions, POST /v1/responses, POST /v1/embeddings, and GET /v1/models, with tool calling passed straight through and streaming in the same event shapes OpenAI sends. Point any OpenAI client — or Codex CLI, on any model — at Tempr and your existing code, streamed UI included, works unchanged.
Model routing
Provider-prefixed model ids like openai/gpt-4o pick the provider per request — and per-key allowlists pin exactly which models each key is allowed to serve.
Custom models on your own servers
Call a model on your own OpenAI-compatible server — vLLM, SGLang, a LiteLLM proxy — as custom/your-name, with the same allowlists, fallback chains, request logs, and optional pricing as any other model. Public HTTPS with an API key, on every plan.
Built for teams
Org-managed virtual keys pool provider keys centrally with per-member budget, rate, and model limits — one provider bill, hard caps per seat, SSO on the account.
How it works
1. Point your SDK at Tempr
Set the base URL to https://api.temprhq.io/v1 in any OpenAI client — Python, Node, curl, whatever you already use.
2. Connect your provider key(s)
Add your own first-party provider keys — OpenAI, Anthropic, Google, xAI — in the Portal. Every request goes out on your key, at your provider's rates — Tempr never fronts tokens.
3. Issue virtual keys
Create a tvk_… key per app or environment with its own budget, rate limits, and model allowlist. Rotate or revoke instantly without touching the real provider key.
4. Ship with observability
From the first request you get logs, analytics by key and model, fallback chains, plus prompt and response caching — no SDK changes, no instrumentation code.
The governed MCP gateway
Every tool call your agents make runs through one place, and that place checks it, signs it in as the right person, and writes it down.
Every call runs as the person asking
Turn on member sign-in for a remote server and each person connects their own GitHub, Linear or Notion account. The agent acts with their permissions, and the tool's own audit trail shows who did what. No shared bot token. Tempr registers with the server's OAuth provider for you, or uses an app you register.
Secrets never reach the laptop
Server URLs, headers and every member's tokens are stored encrypted on Tempr's servers and attached by the relay on the way out, so laptops, CI runners and IDEs never hold them. Tokens are refreshed for you, and revoked when someone signs out, is removed, or leaves.
Tool traffic is checked both ways
Secrets and personal data in a tool call's arguments are flagged or blocked. The same in a tool's results are redacted before the model reads them, or the result is withheld. You choose how strict, per account, organization or key, and can test them first: paste text and see every match, the redacted text and the verdict.
Planted instructions and changed tools are flagged
Text in a result or a tool's description that's shaped like instructions to the model is flagged or withheld, and hidden characters are removed. Each tool's definition is pinned, so a server quietly rewriting a tool shows up in your audit log and webhooks.
Admins see it all, and can stop it
Each key has its own MCP allowlist and limits. Every call is in the request log with its server, tool and what was flagged. Owners see who's signed in to each server, with which account and scopes, and can revoke one person or everyone.
Least privilege, per tool
Bundle the tools a job needs into a Tool Pack and give it to a group, synced from your identity provider, or to a key. Anything outside a member's packs is hidden from the agent and refused if it's called anyway, in every Tempr client and through the gateway.
Your playbook, in every agent
Write your organization's rules once and publish your team's skills in the standard SKILL.md format. Every Tempr agent follows the rules and loads a skill when the task calls for it, in the terminal, the desktop app and every IDE. Required rules can't be switched off. How it works
Every Tempr client, or your own
VS Code, JetBrains, Visual Studio, the CLI and Tempr Code all go through the relay, and show a sign-in link when a server needs one. Any other MCP client can use /mcp/proxy/{server} with a virtual key. Private-network addresses and redirects are refused, and paid tools are never paid for.
The checks catch common shapes of secrets and injected instructions; they aren't a data-loss-prevention product, and no filter makes prompt injection impossible. How the guardrails work.
What teams use it for
The same gateway layer, whatever your reason for wanting keys, logs, and limits in front of your models.
Stop sharing raw provider keys
Teammates and services get scoped virtual keys with their own limits — nobody pastes the production OpenAI key into a Slack thread ever again.
Survive provider outages
Ordered fallback chains move traffic to the next candidate the moment a provider rate-limits or goes down — your app never knows.
Debug real production traffic
Logs searchable by key, model, status, and latency — with per-request detail and optional prompt/response capture when you need the full picture.
Never get a surprise bill
Every plan has a request quota; paid tiers serve past it on metered overage capped at a limit you set. A runaway script tops out quietly instead of billing endlessly.
Stop paying full price for repeated context
Coding agents resend the same tools, instructions, and codebase context every turn. Prompt caching bills that repeat at a fraction of the input price — teammates sending identical context read one shared cache — and exact-match response caching skips the call entirely for identical requests.
Change models without a deploy
The model is a string in the request, and allowlists keep production pinned — try a new model on a dev key, promote it by changing one line.
Bring a key from any of these
The same ~30-provider allowlist as Tempr Chat, enforced server-side — your key, routed straight to the provider. A few of the most popular:
What early users are saying
Illustrative quotes for preview purposes — replace with real customer feedback before launch.
"We swapped the base URL, kept our OpenAI client code exactly as-is, and suddenly had request logs we never had before. Took ten minutes."
— Platform engineer, devtools startup
"Fallback chains saved us during a provider outage — traffic just moved to the next candidate instead of paging anyone at 2am."
— Backend lead, e-commerce team
"No token markup, no shared pool — it's our own key with a much better dashboard in front of it. Exactly what we wanted from a gateway."
— Independent developer
Security & privacy
Bring-your-own-key all the way through — Tempr fronts no tokens and resells none.
Your keys, your account
Provider keys are stored encrypted at rest and never pooled or shared across customers. Every request goes out on a key you supplied.
Content capture, your call
Request/response logging is on by default (7-day retention, truncated at 256 KB) so you can debug real traffic — flip it to metadata-only or off entirely, per account.
Analytics without content
Even in metadata-only mode you keep the full dashboard — requests, tokens, cost, latency, and error rate by key, model, and day — with zero content stored.
Gateway capabilities, plan by plan
Free, Pro, and Scale side by side — every plan is bring-your-own-key.
| Capability | Free | Pro | Scale |
|---|---|---|---|
| Requires your own provider API key(s) | ✔️ | ✔️ | ✔️ |
| Requests included / month | 10,000 | 250,000 | 2,000,000 |
| Rate limit | 10 req/min | 60 req/min | 240 req/min |
| Log retention | 7 days | 30 days | 90 days |
| OpenAI-compatible chat completions, Responses API & models endpoints | ✔️ | ✔️ | ✔️ |
| Tool / function calling passthrough | ✔️ | ✔️ | ✔️ |
| Virtual keys (model allowlist, budget, RPM/TPM) | ✔️ | ✔️ | ✔️ |
| Request logs & per-request detail | ✔️ | ✔️ | ✔️ |
| Analytics (tokens, cost, latency, error rate) | ✔️ | ✔️ | ✔️ |
| Fallback chains | ✔️ | ✔️ | ✔️ |
| Custom models on your own OpenAI-compatible endpoints | ✔️ | ✔️ | ✔️ |
| Prompt caching (automatic breakpoints, key affinity, savings reporting) | ✔️ | ✔️ | ✔️ |
| Response caching | ✔️ | ✔️ | ✔️ |
| Content capture (per-account toggle) | ✔️ | ✔️ | ✔️ |
| Streaming responses (SSE) | ✔️ | ✔️ | ✔️ |
| Budget cap per virtual key | ✔️ | ✔️ | ✔️ |
| Model allowlist per virtual key | ✔️ | ✔️ | ✔️ |
| RPM / TPM limits per virtual key | ✔️ | ✔️ | ✔️ |
| Metadata-only logging mode | ✔️ | ✔️ | ✔️ |
| Remote MCP relay (encrypted server credentials, MCP allowlist per key) | ✔️ | ✔️ | ✔️ |
| Per-member OAuth sign-in to remote MCP servers | ✔️ | ✔️ | ✔️ |
| MCP guardrails (arguments, results, planted instructions) | ✔️ | ✔️ | ✔️ |
| MCP tool definition pinning | ✔️ | ✔️ | ✔️ |
| Overage beyond included requests | Hard block | Metered, capped | Metered, capped |
| Overage rate | — | $0.30 / 1,000 req | $0.20 / 1,000 req |
Gateway FAQ
The questions specific to the API. Account, BYOK, and billing questions live on the general FAQ page.
POST /v1/chat/completions and POST /v1/responses (each streaming and non-streaming), POST /v1/embeddings, and GET /v1/models, with the same request and response shapes the OpenAI SDKs already expect. Point the SDK's base URL at Tempr Gateway and use a virtual key in place of your OpenAI key. The Responses API works on non-OpenAI models too, which is how Codex CLI runs on Claude or DeepSeek through the gateway, and embeddings work with Gemini, Mistral, Cohere and Bedrock models as well as OpenAI's, which is how Roo Code and Kilo Code index a codebase through it.custom/your-name like any other model, with allowlists, fallback chains, and request logs. The server has to be reachable over public HTTPS and require an API key. Custom models aren't priced unless you set a price, so budgets only count their usage once you do. See the custom models guide.POST /mcp/proxy/{server} with a virtual key, or use it from any Tempr client once you're signed in. A server that signs people in with OAuth can sign each member in with their own account; see per-member sign-in. Local servers that run as a process on your machine don't go through the Gateway; the Portal can still store their environment variables for Tempr's clients. See the MCP docs.x-tempr-prompt-cache and the gateway places the cache breakpoints for you; OpenAI, Gemini, DeepSeek, and xAI already cache long prompts automatically. Clients that set their own breakpoints, like Claude Code, keep them intact, and each request's log shows its cached tokens and savings. See the docs.Also on your account
Pair the Gateway with the same agent where you actually write code — one login covers everything.
