Chat (IDE)

Providers & BYOK

Last updated

Tempr is bring-your-own-key (BYOK): you add your own API key for the providers you want, and Tempr relays requests to that provider directly using your key. This applies identically across Chat, the CLI, and the Gateway — same allowlist, same key storage — with one exception: aggregator keys (OpenRouter) work in Chat and the CLI only, not the Gateway.

How BYOK works

There's no shared usage pool. Every model request goes out on a key you supplied, billed by that provider at their own rates. Tempr's fee is for the product wrapped around it — the IDE tooling for Chat, or the virtual keys/logs/analytics for Gateway — never a markup on tokens.

  • Keys are encrypted at rest and never logged in plaintext.
  • You can remove a key at any time, which immediately stops future requests from using it.
  • A key can be scoped to a personal account or, on Teams, pooled at the organization level with per-member budget, rate, and model limits.

Supported providers

This is the allowlist Tempr enforces server-side today, not a wishlist. If a provider isn't listed, requests naming it are rejected before ever reaching a provider.

First-party model providers

The model creator's own API.

AI21 Labs, Aleph Alpha, Alibaba Cloud (DashScope), Anthropic, Codestral, Cohere, DeepSeek, Google Gemini, Inception Labs, Meta AI, MiniMax, Mistral, Moonshot AI, Morph, OpenAI, Perplexity, Public AI, Sarvam AI, xAI, Xiaomi MiMo, Z.ai.

Meta's Llama API (model ids starting meta-llama/) shut down on July 6, 2026, and Tempr no longer supports it. Requests naming it fail with provider_retired. Llama models are still available through Together AI, DeepInfra, Hugging Face, Nscale, AWS Bedrock and OpenRouter: prefix the model with that provider, for example together-ai/meta-llama/Llama-3.3-70B-Instruct-Turbo. Meta AI, Meta's current API, is unaffected.

Inference providers

Hosts other companies' open-weight models on its own infrastructure.

Together AI, Fireworks AI, Hugging Face, DeepInfra, Nscale, Groq, Novita AI, Cerebras.

Cloud providers

A cloud platform hosting first-party models under its own managed service.

AWS Bedrock, Azure OpenAI, Azure AI Foundry.

Gateways & aggregators

One key, routed to many providers and models through a single API. Available in Chat and the CLI only — not the Gateway.

OpenRouter.

Missing a provider?

More providers are added to the allowlist regularly. Let us know and we'll look at adding it.

Model naming

Models are addressed as provider/model (for example openai/gpt-4o, anthropic/claude-opus-4 or google/gemini-3.6-flash). On Gateway, GET /v1/models returns the exact set your account can currently use — see Models endpoint. Custom models on your own endpoints are called as custom/your-name — see Custom models.

Google's Gemini models were previously addressed as gemini/…. That spelling still works, so nothing you have configured needs changing, but google/… is the name they're listed under now — the same prefix OpenRouter, Vercel's AI Gateway and models.dev use.

Perplexity

Perplexity replaced its Sonar models with the Agent API, and Tempr calls it there. Instead of a fixed model, you pick a preset: a model, reasoning effort, search budget and set of tools that Perplexity tunes for a depth of research. Tempr lists the presets as models:

ModelGood forReplaces
perplexity/fastSingle facts and quick summaries, one searchperplexity/sonar
perplexity/lowEveryday research with light multi-step lookupsperplexity/sonar-pro
perplexity/mediumMulti-hop research across many sourcesperplexity/sonar-reasoning-pro
perplexity/highExhaustive, expert-level researchperplexity/sonar-deep-research
perplexity/xhighOpen-ended work that runs code and long tool-use loops—

Your Perplexity key also reaches the other labs' models that the Agent API serves directly, such as perplexity/openai/gpt-5.6-luna, perplexity/anthropic/claude-opus-5 or perplexity/google/gemini-3.8-flash. They're billed to your Perplexity account at Perplexity's rates. Tempr gives them Perplexity's web search tool, and they search when the question calls for it. GET /v1/models lists what your key can use.

The Sonar names still work everywhere, including saved model choices and virtual-key model allowlists, and run as the preset in the same row. A key whose allowlist names perplexity/sonar-pro can use perplexity/low.

  • Reasoning — the preset sets it, so these models have no reasoning picker (nor do the other labs' models through Perplexity), and a reasoning or reasoning_effort you send is not forwarded. Pick a deeper preset for deeper reasoning.
  • Citations — answers cite sources inline ([1] from fast, [web:1] from the others). On chat completions the sources come back as citations (URLs, where citations[n-1] is [n]) and search_results, as Sonar returned them; on /v1/responses they're Perplexity's own search_results output item.
  • Cost — tokens are billed at the preset's underlying model rate (twice that for fast, which runs at priority), plus $2.50 per 1,000 web searches and $0.50 per 1,000 page fetches the preset makes. Tempr records the total Perplexity reports for each request, searches included, and returns it as usage.cost. The prices on the models page are the token rates only.
  • Tools — your own function tools work as usual. Perplexity requires each one to have a description, so a tool without one is sent with its name as its description. The web searches Perplexity runs itself are never handed to you as tool calls.

Azure OpenAI, Azure AI Foundry & AWS Bedrock

These are configured slightly differently from a plain API key, since they're cloud-hosted deployments rather than a single global endpoint:

  • Azure OpenAI — you provide your resource endpoint, deployment name, and API key (or Entra ID credentials, where supported).
  • Azure AI Foundry — you provide your Foundry resource endpoint, API key, and the model id(s) you've deployed (for the non-OpenAI part of the catalog — Phi, Llama, Mistral, and others; GPT models on the same resource still go through Azure OpenAI above).
  • AWS Bedrock — you provide an IAM access key/secret (or role) scoped to Bedrock model invocation in your chosen region.

Custom models

Custom models — your own OpenAI-compatible endpoints — are included with Chat Pro and Teams and on every Gateway plan. Point Tempr at a server you run or rent (vLLM, SGLang, a LiteLLM proxy, or a hosted OpenAI-compatible service) with its base URL and API key, pick its models, and call each one as custom/your-name. See Custom models for setup, pricing, and server notes.