One reasoning setting for every provider, and what it took

  • Gateway
  • Reasoning

Most models worth using for code can think before they answer, and you can usually choose how much. More thinking costs more tokens and takes longer, and it often gets you a better answer on hard problems. What you can't choose is how you ask for it, because no two labs ask the same way.

OpenAI, xAI and Mistral take reasoning_effort. Claude takes adaptive thinking with an effort level (or a token budget, before Claude 4.6). Gemini takes thinkingConfig. DeepSeek, Z.AI, Moonshot, MiniMax, Xiaomi and Cohere each take a thinking field of their own, and Qwen takes enable_thinking. The levels don't line up either: one model offers low, medium and high, another adds minimal or xhigh or max, and some can only be switched on or off.

If you use one model, that's a detail. If you switch models, or send a request to whichever model fits it, the detail becomes a lot of code.

One setting in, the right format out

The Tempr Gateway takes one shape for every model: OpenRouter's reasoning object, or OpenAI's reasoning_effort if you only want to set the level.

json
{
  "model": "anthropic/claude-sonnet-4-6",
  "messages": [{ "role": "user", "content": "Plan the migration." }],
  "reasoning": { "effort": "high" }
}

Tempr sends that on in the format the model's provider reads. The same goes for /v1/messages (Claude's thinking and output_config.effort) and /v1/responses (reasoning.effort), so a tool built for one of those APIs can run on any model. We made four decisions along the way:

  • Ask for nothing, get the model's default. Leave reasoning out and each model thinks as much as its lab intends. Tempr doesn't switch thinking off behind your back to save tokens.
  • A level the model doesn't have becomes the nearest one it does, rounding up on a tie. Asking for xhigh on a model that stops at high gets you high, not an error.
  • The response says what was applied. The x-tempr-reasoning-effort header gives the level that was sent, so you can see when your request was adjusted.
  • none on a model that can't stop thinking gets its lowest level. Claude Opus 5 gets low as well, because Anthropic recommends a low effort over switching its thinking off.

In the IDE extensions this is a reasoning picker next to the model picker, showing only the levels the selected model has, and your choice is remembered per model. In the CLI it's /reasoning, also saved per model, or --reasoning for a single run.

The same model, on three hosts

Labs were the easy part. Open-weight models are served by many inference hosts, and each host decides for itself which reasoning fields it reads. We tested each host with real keys in September 2026, and the documentation didn't always match what the APIs did. The clearest case was GLM-5.3, one model on three hosts:

  • On DeepInfra, the levels worked as expected. Thinking couldn't be switched off, and trying to made the model's reasoning show up in its visible answer.
  • On Together AI, the levels behaved the opposite way: low and high produced no thinking, while max and none did. A picker there would switch thinking off for anyone who chose "high".
  • On Fireworks AI, the levels formed a proper ladder, and none was rejected with a clear error.

So Tempr reads reasoning levels from each host's own catalog entry, not from the model's lab, and the same model can offer different levels, or none, depending on where it runs. GLM on Together AI offers no levels at all, because the host can't steer it reliably. Hugging Face offers none either: it's a router that picks a host for each request, and of the 20 reasoning models we checked there, none had only one host behind it.

The hosts differ in other ways too. Fireworks rejects any field it doesn't recognise, including the reasoning object, so Tempr sends it only reasoning_effort. AWS Bedrock's OpenAI-compatible endpoint ignores the object and reads only reasoning_effort, while Claude and Amazon Nova on Bedrock's Converse API each take their own fields and reject each other's.

Testing it without paying for it

Checking every model on every host would be expensive if every check generated tokens. Most don't need to. A host checks your parameters only after it has accepted the model, and a rejected request costs nothing. Send a deliberately invalid level, such as "reasoning_effort": "banana", and the error says whether your key can reach the model and, often, which levels it accepts. Our whole Together AI investigation cost less than a tenth of a cent.

Keeping it right

Model catalogs change. Tempr takes reasoning levels from models.dev's open catalog, refreshed every 12 hours, and corrects it where a provider's API behaves differently. That correction has already mattered. On October 1 the catalog started listing Kimi K3 as a model that always reasons, and for a few days none ran it at its lowest level. Moonshot's API still switches it off, so Tempr does again, and behaviour we've checked live now overrides the catalog.

The full reference, including every field and how each provider maps it, is in the Reasoning section of the docs. Each model's levels are in GET /v1/models under reasoning_options, and in the model catalog.