Running Codex CLI on Claude, and every other model

  • Gateway
  • Codex

Codex CLI is OpenAI's coding agent for the terminal. It's a good agent, and out of the box it only talks to OpenAI models. That's partly a matter of format: Codex speaks OpenAI's Responses API, not the chat completions format most other tools and providers share. Anthropic, Google and DeepSeek don't serve the Responses API, so you can't simply point Codex at them.

Tempr Gateway serves the Responses API at POST /v1/responses, with the same virtual key, model allowlist, limits and request logs as the rest of the gateway. Add it as a model provider in Codex's config, and Codex can run any model on your allowlist.

toml
model = "anthropic/claude-sonnet-5"
model_provider = "tempr"

[model_providers.tempr]
name = "Tempr Gateway"
base_url = "https://api.temprhq.io/v1"
env_key = "TEMPR_API_KEY"
wire_api = "responses"

Put your virtual key in TEMPR_API_KEY and start codex as usual. The rest of this post is about what happens to those requests.

Two paths: passed through, or translated

What Tempr does with a Responses request depends on the model it's for.

OpenAI models and Azure OpenAI deployments pass straight through. An openai/… model goes to OpenAI's own Responses API, and azure-openai/your-deployment-name goes to your Azure resource's. Tempr sends the body as it came and relays the answer untouched, so reasoning items, hosted tools and previous_response_id work as they do against OpenAI directly. OpenAI's codex models, which OpenAI serves only on the Responses API, work this way too.

Every other model is translated. The request becomes the provider's own API call, and the answer comes back as Responses events: response.created, response.output_text.delta and so on through response.completed. Instructions, input messages, function tools, custom (freeform) tools, tool_choice, max_output_tokens and text.format JSON schemas carry over. reasoning.effort does too, mapped onto the model's own reasoning settings. Codex's freeform apply_patch tool becomes a function tool taking one string, and calls to it go back to Codex as the custom tool calls it expects.

What doesn't carry over

Some of the Responses API only makes sense on OpenAI's side, and we'd rather say so than pretend.

  • Hosted tools, such as web search, file search and the code interpreter, run on OpenAI's servers. Another provider can't run them. Tempr drops them rather than failing the request, and names what it dropped in the x-tempr-unsupported-tools response header. A Codex session with web search turned on still works on Claude, without the web search.
  • Encrypted reasoning is OpenAI's own, and isn't sent to other providers.
  • previous_response_id continues a response the server stored. Tempr doesn't store responses, so on a translated model it's refused with 400 previous_response_id_unsupported, and the message says to send the conversation in input instead. Codex already does that on every turn, so it isn't affected.
  • reasoning.summary, store, include, truncation and service_tier are ignored.

Sub-agents on every model

Codex can start sub-agents, and it gives the model the tools to do it in two newer shapes of the Responses API. Tools can be grouped in a namespace, and a call to one comes back carrying its namespace, which Codex dispatches on. Or the tools can sit behind a client-run tool_search ("execution": "client"): the model searches, Codex answers with a tool_search_output listing the tools, and from then on the model can call them.

Neither shape existed in chat completions, so at first Codex's sub-agent tools didn't reach other models. Since October 1 they do. A namespace's tools become ordinary function tools for the provider, and their calls go back out with their namespace. A client-run tool_search becomes a function tool whose calls go back as tool_search_call items, and the tools a tool_search_output returned join the request's tools, the way OpenAI keeps the tools a search loaded. The result is that Codex can start sub-agents on Claude and the rest, as it does on GPT.

Telling Codex about your models

There was a second, quieter gap, and it was on Codex's side. Codex decides which tools to offer from its own catalog of models, and the models you use through Tempr aren't in it, not even OpenAI's codex models under their openai/… ids. For a model it doesn't know, Codex warns Model metadata for … not found, assumes a 272K-token context window, and leaves out apply_patch, so edits fall back to shell commands.

The fix is to give Codex a catalog of your models. The Codex CLI guide has a short command that copies the details of a model Codex ships with, gpt-5.5, under each of your model ids, with the context windows you set. Then point Codex at the file by its full path:

toml
model_catalog_json = "/home/you/.codex/tempr-models.json"

The catalog replaces Codex's own, so list every model you use through Tempr, and run the command again after upgrading Codex. The copied details include reasoning levels, so /model offers low to xhigh effort on every model, and Tempr maps the level onto each provider's settings.

Logging the last event

When we tested Codex through Tempr at the start of October 2026, we found requests in the logs with no status, tokens or cost. The cause was timing. Codex disconnects as soon as a Responses stream's last event, response.completed, arrives. OpenAI keeps the connection open a moment longer, and Tempr waited for that close before finishing the request's record. When Codex hung up first, the record never got its numbers.

Streams now end at their last event: response.completed on the Responses API, Anthropic's message_stop, or [DONE] on chat completions. A request a client hangs up on at the end is logged in full.

Fallback chains still apply

/v1/responses tries the next model in a fallback chain, and the next key for a provider, the same way chat completions does. A chain can mix OpenAI and other models, passed through and translated in one chain. A request that uses previous_response_id only goes to the models in the chain that can take it as-is, OpenAI's and Azure's.

Getting started

You need a Gateway virtual key and a key for at least one provider. The Codex CLI guide walks through the setup, and the Responses API section of the reference lists exactly what carries over to other models. Errors come back in OpenAI's format, so Codex reads them as it would OpenAI's; the codes are in Errors. If you also run Claude Code, the same gateway serves Anthropic's Messages API at /v1/messages: see the Claude Code guide.