Guides

Kilo Code

Last updated

Kilo Code's custom provider setting lets you supply any OpenAI-compatible base URL. Point it at Tempr Gateway and one provider entry gives you every model on your allowlist — Claude, GPT, Gemini, whatever you've configured — instead of switching providers per model.

  1. Create a virtual key

    In the Portal, create a Gateway virtual key. Keys are prefixed tvk_ and shown once at creation.

  2. Add a provider key

    Gateway is BYOK: add a key for at least one provider from the supported list.

  3. Add a custom provider

    Open Kilo Code's settings (gear icon) → Providers tab → Custom provider. Set Base URL to https://api.temprhq.io/v1 and paste your virtual key as the API key (no Bearer prefix — Kilo Code adds that automatically).

  4. Pick a model

    Kilo Code fetches the available model list automatically from GET /v1/models; choose one from your allowlist. The list holds chat models only; embedding models, for indexing below, are entered by id.

Codebase indexing

Kilo Code's codebase indexing embeds your code so it can search it by meaning. It can embed through Tempr too, with the same virtual key, so it needs no separate embedding provider account.

  1. Open the indexing settings

    Go to Settings → Indexing, or click the indexing indicator at the bottom of the prompt input.

  2. Point the embedder at Tempr

    Choose the OpenAI-Compatible embedding provider, set the base URL to https://api.temprhq.io/v1, and paste your virtual key as the API key.

  3. Choose an embedding model

    Enter an embedding model from a provider you have a key for, such as openai/text-embedding-3-small, google/gemini-embedding-001 or mistral/codestral-embed, and, if Kilo asks, its dimension from the embeddings models table (1536 for text-embedding-3-small). The model has to be on your virtual key's allowlist.

  4. Pick a vector store and start

    Keep the default LanceDB store, which needs no server, or connect Qdrant, then start indexing.

Each batch Kilo sends is one Gateway request, billed at the model's input price with your own provider key, and appears in your request logs. Changing the embedding model or its dimension means re-indexing.

What's next