OpenAI's newest models, on the API everything else speaks

  • Gateway
  • OpenAI

OpenAI has two APIs for talking to its models. Chat completions is the older one, and nearly every tool speaks it: IDE extensions, agent frameworks, scripts, and most gateways, Tempr's included. The Responses API is newer. OpenAI builds its new features there first, and more and more of its models can do some things only there.

For a tool that speaks chat completions, that shows up as errors that are hard to explain. A model is in the model list, but every request fails. A request works until you turn reasoning up, then fails. An agent works until it tries to call a tool. We hit all three, so the Tempr Gateway now sends exactly those requests through the Responses API and translates the answer back.

Three ways a request can fail

We checked each model live, with real requests. There are three kinds of gap:

  • The whole model is Responses-only. OpenAI's pro and codex models, such as gpt-5-pro, o3-pro and gpt-5.3-codex, answer every chat completions request with a 404.
  • The deepest reasoning level is Responses-only. gpt-5.6 with its luna, sol and terra variants, and gpt-6-astra, gpt-6-luna, gpt-6-sol and gpt-6.1-sol, reject reasoning effort max on chat completions. Every other level works on both APIs.
  • Function tools with reasoning are Responses-only. On those same models, chat completions rejects any request with function tools unless reasoning is set to none, and that includes requests that don't set reasoning at all. gpt-6-astra has no none level, so on chat completions it can't call a function tool at all.

The third one matters most. Calling tools is what an agent does: read a file, run a command, edit code. A model that can't call tools on chat completions is no use to a coding agent that speaks chat completions.

A bridge for each request

Tempr decides which API to use for each request, not for each model. A gpt-6-sol request with tools and reasoning goes to the Responses API. The same model's plain requests stay on chat completions, where parameters like stop and seed still apply. The Responses API has no equivalent of those, so a bridged request goes without them.

On the way there, the chat request becomes a Responses request. Messages become input items in the same order. Assistant tool calls and their results become function-call items. Your reasoning setting becomes the Responses API's own. Two details took some care:

  • Tool schemas keep chat's behaviour. Chat completions doesn't enforce a function's schema strictly unless you ask it to, and the Responses API does by default. A loose tool schema, as many IDE tools have, would be held to stricter rules than it was written for, so Tempr keeps chat's default.
  • Reasoning survives a tool loop. Tempr doesn't store responses, so it asks OpenAI for the model's reasoning in encrypted form, the only way to get it back without storage. It comes back as reasoning_details on the assistant message, which Tempr's clients already return on the next request, and goes back to OpenAI from there. A model reasoning through a sequence of tool calls keeps its earlier reasoning.

On the way back, the Responses answer becomes a chat completions answer, streaming included, with text, tool calls, finish reasons and usage where chat clients expect them. Tempr doesn't ask for reasoning summaries, though: OpenAI rejects the whole request when the API key's organization hasn't been verified, and we didn't want a request to fail over an extra.

All this applies to every path into Tempr: the Gateway's /v1/chat/completions and /v1/messages, the IDE extensions, the CLI and agent runs.

Knowing which requests to send

OpenAI's model list doesn't say which API a model needs, so Tempr works it out in three ways:

  1. By name. An OpenAI model with -pro at the end or codex in its name goes straight to Responses, which saves a rejected request.
  2. From the error. When chat completions answers that a model or a request only works on the Responses API, Tempr sends that request there instead, and remembers.
  3. From a list. For the models whose gaps we've checked, the reasoning levels and tool requests chat completions rejects are listed, so those requests go to Responses the first time.

The second step needed a fix. Chat completions rejects a function-tool request on these models with an error that also says to use the Responses API. The first version of the bridge read that as "this model is Responses-only" and sent all of the model's requests there, so its plain requests lost stop and seed for no reason. Tempr now remembers what it learns separately for each model and kind of request: a model that only needs Responses for tools keeps using chat completions for everything else.

New models

New OpenAI models arrive often. Tempr picks up their reasoning levels on its own, from a catalog it refreshes every 12 hours, but not which requests they reject on chat completions. Checking that is quick and doesn't cost anything. Chat completions checks the model before the parameters, so a request with a reasoning level that doesn't exist, such as "banana", is rejected without generating any tokens, and the error says whether chat completions serves the model. A level the model doesn't have, such as max, gets an error listing the levels it does have. That's how we found the same gaps on gpt-6-luna, gpt-6-sol and gpt-6.1-sol as on gpt-6-astra, and they joined the list on October 5.

You don't need to do anything to use any of this: send your chat completions request as usual. If you'd rather use the Responses API yourself, Tempr's /v1/responses passes OpenAI requests straight through. See the Responses API section of the docs, and the Codex CLI guide for a client that speaks it.