<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Tempr blog</title>
  <subtitle>Notes from the team building Tempr: how the Gateway, the IDE extensions and the CLI work, and what we learn running them across every AI provider.</subtitle>
  <id>https://temprhq.io/blog</id>
  <link rel="alternate" type="text/html" href="https://temprhq.io/blog"/>
  <link rel="self" type="application/atom+xml" href="https://temprhq.io/blog.xml"/>
  <icon>https://temprhq.io/assets/favicon.svg</icon>
  <author><name>The Tempr team</name><uri>https://temprhq.io/</uri></author>
  <updated>2026-10-07T00:00:00Z</updated>
  <entry>
    <title>Every MCP tool call, through one place you control</title>
    <id>https://temprhq.io/blog/every-mcp-tool-call-through-one-place</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/every-mcp-tool-call-through-one-place"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="MCP"/>
    <category term="Gateway"/>
    <category term="Security"/>
    <summary>MCP gives coding agents real tools, along with tokens on every laptop, shared bot accounts and tool output the model trusts. How Tempr's Gateway relays MCP traffic so it runs as the right person, gets checked both ways, and leaves a record.</summary>
    <content type="html">&lt;p&gt;MCP servers are how coding agents get real tools: open a GitHub issue, update a Linear ticket, search Notion, query a database. Connecting one is easy. Connecting one for a whole team raises the questions a security review will ask:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Where do the credentials live?&lt;/strong&gt; Usually in a config file on every developer's laptop and in every CI runner, often as one shared token.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Who does the call run as?&lt;/strong&gt; With a shared token, every agent acts as the same bot account, so the tool's own audit trail can't say which person asked.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;What goes out, and what comes back?&lt;/strong&gt; An agent can paste a secret into a tool call. A tool's result goes straight into the model's context, and it can carry text written to steer the model.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;What happened?&lt;/strong&gt; Nobody has a record of which tools were called, by whom, with what.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Tempr's Gateway now answers those in one place. Remote MCP servers are added once, in the Portal, and every call Tempr's IDE extensions and CLI make to a remote server goes through the Gateway's MCP relay on its way to the server. Any other MCP client can use the same relay at &lt;code&gt;/mcp/proxy/{server}&lt;/code&gt; with a Gateway key.&lt;/p&gt;
&lt;h2 id=&quot;credentials-stay-on-the-server&quot;&gt;Credentials stay on the server&lt;/h2&gt;
&lt;p&gt;A server's address, its headers and every member's tokens are stored encrypted on Tempr's servers. The relay adds them to each request on the way out, so laptops, CI runners and editors never hold them. There's nothing to copy into a dotfile, nothing to leak from a laptop, and nothing to rotate on fifty machines when someone leaves.&lt;/p&gt;
&lt;h2 id=&quot;each-call-runs-as-the-person-asking&quot;&gt;Each call runs as the person asking&lt;/h2&gt;
&lt;p&gt;For servers that sign people in with OAuth, such as GitHub, Linear or Notion, you can turn on &lt;strong&gt;member sign-in&lt;/strong&gt;. Each person then connects their own account, and the agent acts with that person's permissions. A ticket the agent opens shows who actually asked for it, and someone without access to a private repository can't reach it through the agent either.&lt;/p&gt;
&lt;p&gt;Setting that up is usually the hard part of OAuth, so Tempr does it for you. It registers itself with the server's OAuth provider, using a Client ID Metadata Document where the provider supports one and dynamic registration otherwise, or uses an OAuth app you register yourself where a provider requires that, as GitHub does. Members' tokens are stored encrypted and refreshed when they expire, and they're revoked when someone signs out, is removed, or leaves the organization.&lt;/p&gt;
&lt;p&gt;A member who hasn't signed in yet still sees the server's tools. Their first tool call asks them to sign in, using MCP's own request for it, and every Tempr client shows that as a link. The CLI prints it and opens it in your browser, and the IDE extensions show a sign-in link on the tool call, so signing in takes one click in the middle of a task.&lt;/p&gt;
&lt;h2 id=&quot;tool-traffic-is-checked-both-ways&quot;&gt;Tool traffic is checked both ways&lt;/h2&gt;
&lt;p&gt;The relay reads every tool call and every result, and three checks run on them. You choose how strict each one is, for your account or your whole organization, and a key can override them:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;What goes out.&lt;/strong&gt; Secrets and personal data in a tool call's arguments, such as bearer tokens, JWTs, AWS keys, &lt;code&gt;sk-&lt;/code&gt; style API keys, email addresses and card-like numbers, are flagged in the request log by default. Set it to block, and the call is refused and never sent.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;What comes back.&lt;/strong&gt; The same kinds of data in a tool's result are redacted by default, masked before the model ever reads them. Set it to block, and the whole result is withheld.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Planted instructions.&lt;/strong&gt; Text in a tool's result or description that's shaped like instructions to the model, such as &amp;quot;ignore previous instructions&amp;quot; or a fake system prompt, is flagged by default, or can be withheld. Hidden characters that can carry instructions a person can't see, such as zero-width spaces, Unicode tag characters and direction overrides, are removed.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;We're careful about what this is. The checks catch the common shapes of secrets and injected instructions. They aren't a data-loss-prevention product, and no filter makes prompt injection impossible. What they give you is a first line of defence on the path every tool call takes, and a record of what they caught.&lt;/p&gt;
&lt;h2 id=&quot;tools-can-t-change-quietly&quot;&gt;Tools can't change quietly&lt;/h2&gt;
&lt;p&gt;An MCP server describes its tools to the model: a name, a description and the input it expects. Those descriptions are instructions the model follows, and a server can change them at any time, including after your team has started relying on a tool that looked harmless.&lt;/p&gt;
&lt;p&gt;So Tempr pins each tool's definition the first time it sees it. When a server changes a tool's name, description or input schema, the change goes to your organization's audit log, with the old and new description, and to your Gateway webhook as &lt;code&gt;mcp.tool_changed&lt;/code&gt;. The tool keeps working, because most changes are ordinary updates, but nothing changes without your team being told.&lt;/p&gt;
&lt;h2 id=&quot;what-admins-see&quot;&gt;What admins see&lt;/h2&gt;
&lt;p&gt;Every tool call is in the request log with its server, its tool, and anything that was flagged or redacted. Each Gateway key has its own list of allowed MCP servers and its own limits. For servers with member sign-in, owners and admins see who's signed in, with which account and scopes, and when it was last used, and they can revoke one member or everyone at once.&lt;/p&gt;
&lt;p&gt;A few protections apply to every server, whatever its settings. The relay refuses private-network addresses and redirects, so a server entry can't be used to reach your internal network. And when a tool asks to be paid per call, the Gateway refuses it and never pays; the &lt;a href=&quot;https://temprhq.io/blog/paid-mcp-tools-refused-never-paid&quot;&gt;paid MCP tools post&lt;/a&gt; covers why.&lt;/p&gt;
&lt;h2 id=&quot;getting-started&quot;&gt;Getting started&lt;/h2&gt;
&lt;p&gt;Add a remote server in the Portal under MCP servers, choose its sign-in setting, and set the guardrails under MCP credentials. Once you're signed in, every Tempr client can use it. The &lt;a href=&quot;https://temprhq.io/gateway&quot;&gt;Gateway page&lt;/a&gt; has the overview, and the reference covers &lt;a href=&quot;https://temprhq.io/docs/gateway-reference#mcp-guardrails&quot;&gt;the guardrails&lt;/a&gt; and &lt;a href=&quot;https://temprhq.io/docs/gateway-reference#mcp-sign-in&quot;&gt;member sign-in&lt;/a&gt; in detail. Connecting servers in the IDE is in the &lt;a href=&quot;https://temprhq.io/docs/mcp#remote&quot;&gt;MCP docs&lt;/a&gt;.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>A sandbox for the agent's commands, with nothing to install</title>
    <id>https://temprhq.io/blog/a-sandbox-with-nothing-to-install</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/a-sandbox-with-nothing-to-install"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="CLI"/>
    <category term="Security"/>
    <summary>How the Tempr CLI lets an agent run any build or test command without asking, by limiting what those commands can touch. Why we built it on Linux's Landlock instead of the usual tools, and what it took to make real builds work inside it.</summary>
    <content type="html">&lt;p&gt;A coding agent is most useful when it can run things: build the project, run the tests, try a fix, run the tests again. It's most worrying for the same reason. Every command it runs can do anything you can do on that machine.&lt;/p&gt;
&lt;p&gt;Approval prompts are the usual answer, and they work while you're watching. They don't work for a run nobody is watching, in CI or a long task you've left alone. An allow-list helps, but only for the commands you thought of in advance. The agent's next idea, a script it just wrote or a tool it wants to try, stops the run.&lt;/p&gt;
&lt;p&gt;The Tempr CLI's sandbox turns that around. Instead of deciding which commands may run, it limits what any command can do, and then lets the agent run whatever it needs.&lt;/p&gt;
&lt;h2 id=&quot;what-the-sandbox-allows&quot;&gt;What the sandbox allows&lt;/h2&gt;
&lt;p&gt;Inside the sandbox, the agent's commands:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;can change files only in the current folder, the temp folders and the package caches&lt;/strong&gt; builds write to: npm's, pip's, NuGet's, Go's, Cargo's, Gradle's and Maven's. Everything else, your home folder, &lt;code&gt;/etc&lt;/code&gt;, other projects, is read-only to them.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;can read anything&lt;/strong&gt;, as a build needs to.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;can't open network connections&lt;/strong&gt;, if you choose. Cutting off the network blocks outgoing TCP connections, while name lookups still work.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Because the damage a command can do is limited, any command runs without asking, in every approval mode, even one with no allow-list entry. A few still ask: commands that delete, publish or force-push, because those can do harm inside the project too. When the sandbox blocks something a command really needed, such as a download, the agent can ask to run that command outside the sandbox, and that always asks you first. A run with nobody to ask refuses it.&lt;/p&gt;
&lt;p&gt;You turn it on with &lt;code&gt;--sandbox&lt;/code&gt;, or for every run with &lt;code&gt;tempr config set-sandbox on&lt;/code&gt;. Since CLI 0.9.8 it's also the default for runs where nobody decides command by command, &lt;code&gt;--yolo&lt;/code&gt; and autopilot in CI, with the network left on there. Cutting the network also blocks connections to &lt;code&gt;localhost&lt;/code&gt;, and a test suite that starts a local server needs those.&lt;/p&gt;
&lt;h2 id=&quot;why-not-bubblewrap&quot;&gt;Why not bubblewrap&lt;/h2&gt;
&lt;p&gt;The usual tool for sandboxing a process on Linux is bubblewrap, which builds an isolated view of the system using the kernel's namespaces. We tried it first, on the systems our users actually run, and it failed on two of them:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Stock Ubuntu 24.04&lt;/strong&gt;, a very common Linux on developer machines and CI runners, restricts the unprivileged user namespaces bubblewrap needs. As a normal user, it couldn't start.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;A default Docker container&lt;/strong&gt;, which is where many CI jobs and our own container recipe run, doesn't allow creating new namespaces either.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;A sandbox that's missing in the two places it matters most isn't much of a sandbox. So we built ours on &lt;strong&gt;Landlock&lt;/strong&gt;, a security feature in the Linux kernel itself. With Landlock, a process restricts itself: it declares which folders it may write to and whether it may connect out, and from then on neither it nor anything it starts can do more. It needs no install, no root and no namespaces. It works on stock Ubuntu 24.04 and inside a default Docker container. Restricting files needs Linux 5.13 or later, and restricting the network needs 6.7 or later. Inside WSL2 on Windows, it works too.&lt;/p&gt;
&lt;h2 id=&quot;the-parts-landlock-can-t-do&quot;&gt;The parts Landlock can't do&lt;/h2&gt;
&lt;p&gt;Landlock grants access to whole folders. The project folder has to be writable, or nothing could build. But a few files inside it can do harm later, outside the sandbox: git hooks run on your next commit, git's config can point at a program to run, and &lt;code&gt;.agentcommands.json&lt;/code&gt; is the agent's own allow-list. A command inside the sandbox could change any of them, and the change would take effect after the sandbox was gone.&lt;/p&gt;
&lt;p&gt;So Tempr takes a copy of &lt;code&gt;.git/hooks&lt;/code&gt;, &lt;code&gt;.git/config&lt;/code&gt; and &lt;code&gt;.agentcommands.json&lt;/code&gt; before each sandboxed command, and puts back anything the command changed. The agent is told when that happens, so it doesn't build on a change that's been undone.&lt;/p&gt;
&lt;h2 id=&quot;making-real-builds-work&quot;&gt;Making real builds work&lt;/h2&gt;
&lt;p&gt;A sandbox that blocks a normal build is one people turn off, so we spent most of our time on what builds actually need to write.&lt;/p&gt;
&lt;p&gt;Package caches were the obvious part. A restore or install writes to a per-user cache outside the project, so those folders are writable: npm's, Cargo's and Maven's from the first release, NuGet's and Go's module cache from 0.9.8, and pip's and Gradle's too.&lt;/p&gt;
&lt;p&gt;.NET took more work. The &lt;code&gt;dotnet&lt;/code&gt; tool needs to write small marker files in &lt;code&gt;~/.dotnet&lt;/code&gt;, and fails if it can't. The documented switch to skip that didn't help in our tests on .NET 10. Making &lt;code&gt;~/.dotnet&lt;/code&gt; writable wasn't an option either, because it holds global tools, programs that run outside the sandbox later. So inside the sandbox, &lt;code&gt;dotnet&lt;/code&gt; gets its own folder for those files, in &lt;code&gt;~/.cache/tempr/sandbox-dotnet&lt;/code&gt;. Then it turned out NuGet follows that same setting for its packages and configuration, so Tempr points NuGet back at your real &lt;code&gt;~/.nuget/packages&lt;/code&gt; and package sources. The result is that &lt;code&gt;dotnet&lt;/code&gt; builds, restores and tests inside the sandbox, even on a machine where it has never run before, while your global tools stay read-only.&lt;/p&gt;
&lt;h2 id=&quot;checked-on-real-machines&quot;&gt;Checked on real machines&lt;/h2&gt;
&lt;p&gt;Both releases that changed the sandbox were checked on fresh cloud machines running stock Ubuntu 24.04, as a normal user, on both Intel and Arm (AWS Graviton) processors. Each check confirmed that writes to the project and the caches work, writes to the home folder and &lt;code&gt;/etc&lt;/code&gt; are denied, the network is blocked when it should be and open when it should be, and pipes still work. One lesson from that: an emulated Arm machine reported no Landlock at all, so it couldn't prove anything. Only real hardware could.&lt;/p&gt;
&lt;h2 id=&quot;where-it-runs-and-where-it-doesn-t-yet&quot;&gt;Where it runs, and where it doesn't yet&lt;/h2&gt;
&lt;p&gt;The sandbox runs on Linux. On macOS and Windows it isn't available yet. What happens there depends on whether you asked for it:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;If you asked for it with &lt;code&gt;--sandbox&lt;/code&gt; and it can't run, nothing runs. The CLI exits with code 7 instead of running your commands unsandboxed.&lt;/li&gt;&lt;li&gt;The default for &lt;code&gt;--yolo&lt;/code&gt; and CI wasn't something you asked for, so where the sandbox can't run, those runs go ahead as they did before it existed, and an interactive session tells you so.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;&lt;code&gt;tempr doctor&lt;/code&gt; tells you what the sandbox can do on your machine. It covers the commands the agent runs, the .NET test tools and the agent's diagnostics. It doesn't cover commands you type yourself with &lt;code&gt;!command&lt;/code&gt;, or MCP servers.&lt;/p&gt;
&lt;p&gt;The details are in the &lt;a href=&quot;https://temprhq.io/docs/cli-reference#sandbox&quot;&gt;sandbox section of the CLI reference&lt;/a&gt;, and &lt;a href=&quot;https://temprhq.io/docs/cli-scripts#ci&quot;&gt;Scripts, pipes and CI&lt;/a&gt; shows it in a pipeline. For more on running the CLI unattended, see &lt;a href=&quot;https://temprhq.io/blog/piping-into-the-tempr-cli&quot;&gt;Pipe it to the agent&lt;/a&gt;.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Pipe it to the agent: the Tempr CLI in scripts and CI</title>
    <id>https://temprhq.io/blog/piping-into-the-tempr-cli</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/piping-into-the-tempr-cli"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="CLI"/>
    <summary>The Tempr CLI reads stdin, writes stdout and exits with a code that means something, so an AI agent fits into a shell pipeline or a CI job like any other command. Patterns that work, and what we fixed so they do.</summary>
    <content type="html">&lt;p&gt;The quickest way to give an agent context is often the one you already use for everything else on the command line: a pipe. The output you're looking at, a diff, a build log, a list of commits, goes straight into the agent along with what you want done with it.&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;bash&lt;/span&gt;&lt;pre&gt;&lt;code&gt;git diff | tempr &amp;quot;Review this change for bugs and missing tests&amp;quot;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The Tempr CLI is built to fit into that kind of workflow. It reads stdin, writes to stdout, reports in JSON when you ask it to, and exits with a code a script can act on. Here are the patterns we use most, and what had to be right underneath for them to work.&lt;/p&gt;
&lt;h2 id=&quot;the-prompt-says-what-the-pipe-says-what-with&quot;&gt;The prompt says what, the pipe says what with&lt;/h2&gt;
&lt;p&gt;With a prompt argument, piped text follows the prompt as context. With no prompt, the piped text is the whole prompt, so you can write the instruction into the stream yourself:&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;bash&lt;/span&gt;&lt;pre&gt;&lt;code&gt;# Explain a failure from a build log
cat build.log | tempr &amp;quot;Why did this fail? Cite the relevant lines.&amp;quot;
# Put the instruction in the stream
{ echo &amp;quot;Summarize these commits for the release notes:&amp;quot;; git log --oneline v1.2..; } | tempr
# A prompt too long or too quote-heavy for the command line
tempr --prompt-file task.md&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;A piped run takes one turn and exits, so it works in a script. For a single file, you don't need a pipe at all: &lt;code&gt;@path&lt;/code&gt; in the prompt attaches it, as in &lt;code&gt;tempr &amp;quot;add doc comments to @src/Payment/PaymentRequest.cs&amp;quot;&lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&quot;what-we-fixed-in-0-9-11&quot;&gt;What we fixed in 0.9.11&lt;/h2&gt;
&lt;p&gt;Two things got in the way of this until CLI 0.9.11, released on October 5.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Piped input was ignored whenever there was a prompt.&lt;/strong&gt; &lt;code&gt;git diff | tempr &amp;quot;Review this&amp;quot;&lt;/code&gt; sent the prompt and dropped the diff, so the model reviewed nothing. Now the prompt goes first and the piped input after it, as it does in Claude Code and Codex. The CLI waits up to 3 seconds for piped input to start, so a slow &lt;code&gt;git diff&lt;/code&gt; on a large repository still makes it. If nothing arrives, which happens when a parent process leaves stdin open without writing to it, the prompt runs on its own and the CLI notes that on stderr. Redirect stdin from &lt;code&gt;/dev/null&lt;/code&gt; (&lt;code&gt;NUL&lt;/code&gt; on Windows) to skip the wait.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Whitespace was being collapsed.&lt;/strong&gt; Blank lines, indentation, and lines holding only spaces were squeezed into single spaces before the prompt reached the model. For prose that hardly matters. For a diff it matters a lot: lines ran together, and a review could report formatting problems that weren't in the code. Prompts now reach the model exactly as written, blank lines and indentation included, and that fix applies to &lt;code&gt;tempr acp&lt;/code&gt; in editors too.&lt;/p&gt;
&lt;h2 id=&quot;asking-it-to-act-not-just-answer&quot;&gt;Asking it to act, not just answer&lt;/h2&gt;
&lt;p&gt;Piping is for context. When the agent should change something, give it the goal and let it work:&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;bash&lt;/span&gt;&lt;pre&gt;&lt;code&gt;tempr -y &amp;quot;fix the failing test in CartServiceTests and run the suite&amp;quot;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;code&gt;-y&lt;/code&gt; lets file edits and allow-listed commands run without stopping to ask. Commands that aren't on the allow-list, and anything that deletes, publishes or force-pushes, still need approval. When there's no terminal to approve them on, because input is piped, output is JSON, or the run is in CI, they're refused, and the model is told why so it can find another way or stop. You add the commands a job needs with &lt;code&gt;tempr config allow-command &amp;lt;prefix&amp;gt;&lt;/code&gt;, or an &lt;code&gt;.agentcommands.json&lt;/code&gt; file at the workspace root.&lt;/p&gt;
&lt;p&gt;On Linux, &lt;code&gt;-y&lt;/code&gt; also runs the agent's commands in a sandbox by default, built on the kernel's Landlock, with no install and no special privileges. Commands can change the checkout, temporary folders and package caches, and nothing else. &lt;code&gt;--sandbox&lt;/code&gt; cuts off the network as well. If you ask for the sandbox and the machine can't provide it, nothing runs.&lt;/p&gt;
&lt;h2 id=&quot;output-a-script-can-read&quot;&gt;Output a script can read&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;--json&lt;/code&gt; turns the run into newline-delimited JSON: a first line describing the session, one line per event as the agent works, and a final result line.&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;bash&lt;/span&gt;&lt;pre&gt;&lt;code&gt;tempr --json -y &amp;quot;list the public types in this project&amp;quot; \
  | tail -n 1 \
  | jq '{status, summary, files: .filesChanged}'&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The result line carries the final answer as &lt;code&gt;summary&lt;/code&gt;, the files the agent changed, token usage, and a &lt;code&gt;status&lt;/code&gt;. The exit code makes the same distinctions, so a script can branch on it without parsing anything:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;code&gt;0&lt;/code&gt;, &lt;code&gt;ok&lt;/code&gt;: the turn finished.&lt;/li&gt;&lt;li&gt;&lt;code&gt;1&lt;/code&gt;, &lt;code&gt;error&lt;/code&gt;: something went wrong: the server, the network, a tool.&lt;/li&gt;&lt;li&gt;&lt;code&gt;2&lt;/code&gt;: the command line didn't make sense.&lt;/li&gt;&lt;li&gt;&lt;code&gt;3&lt;/code&gt;: not signed in, or the license has no plan.&lt;/li&gt;&lt;li&gt;&lt;code&gt;4&lt;/code&gt;, &lt;code&gt;step-limit&lt;/code&gt;: the run hit its step limit with work left. &lt;code&gt;--continue&lt;/code&gt; picks it up.&lt;/li&gt;&lt;li&gt;&lt;code&gt;5&lt;/code&gt;, &lt;code&gt;tool-refused&lt;/code&gt;: a tool call was refused, usually because nothing could approve it.&lt;/li&gt;&lt;li&gt;&lt;code&gt;6&lt;/code&gt;, &lt;code&gt;cost-limit&lt;/code&gt;: the run reached its &lt;a href=&quot;https://temprhq.io/blog/a-price-ceiling-for-agent-runs&quot;&gt;price ceiling&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;&lt;code&gt;7&lt;/code&gt;: you asked for the sandbox and it can't run on this machine, so nothing ran.&lt;/li&gt;&lt;/ul&gt;
&lt;h2 id=&quot;in-ci&quot;&gt;In CI&lt;/h2&gt;
&lt;p&gt;Put those together and an agent becomes a pipeline step. A pull-request review in GitHub Actions:&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;yaml&lt;/span&gt;&lt;pre&gt;&lt;code&gt;- env:
    TEMPR_LICENSE_KEY: ${{ secrets.TEMPR_LICENSE_KEY }}
  run: |
    { echo &amp;quot;Review this PR:&amp;quot;; git diff origin/main...; } \
      | tempr --json --max-cost 1.00 -m anthropic/claude-sonnet-4-6&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;A few habits make this reliable:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Keep the key in a secret.&lt;/strong&gt; The CLI signs in from &lt;code&gt;TEMPR_LICENSE_KEY&lt;/code&gt;, so the key never appears on a command line or in a log.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Name the model.&lt;/strong&gt; &lt;code&gt;-m&lt;/code&gt; keeps the run from depending on a default set on someone's machine.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Cap the spend.&lt;/strong&gt; &lt;code&gt;--max-cost&lt;/code&gt; puts a ceiling on the run, and exit code 6 tells the job it was reached.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Keep tasks read-only&lt;/strong&gt; unless the workspace is disposable. Review, summarize and triage are safe anywhere; edits belong on a runner you can throw away.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;CI systems set the &lt;code&gt;CI&lt;/code&gt; variable, and the CLI notices: a tool call that would have asked for approval is refused instead of waiting forever, and there's no logo, update notice or crash report in your logs.&lt;/p&gt;
&lt;p&gt;The full guide, with a complete workflow file, is &lt;a href=&quot;https://temprhq.io/docs/cli-scripts&quot;&gt;Scripts, pipes and CI&lt;/a&gt; in the docs, and every flag is in the &lt;a href=&quot;https://temprhq.io/docs/cli-reference&quot;&gt;CLI reference&lt;/a&gt;. If you haven't installed the CLI yet, start on the &lt;a href=&quot;https://temprhq.io/cli&quot;&gt;CLI page&lt;/a&gt;.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Paid MCP tools: refused clearly, never paid</title>
    <id>https://temprhq.io/blog/paid-mcp-tools-refused-never-paid</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/paid-mcp-tools-refused-never-paid"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="MCP"/>
    <category term="Chat"/>
    <category term="CLI"/>
    <category term="Gateway"/>
    <summary>Some MCP servers now charge per tool call, through x402 or MPP. What an agent does when a tool asks to be paid, why Tempr doesn't pay, and how it makes the refusal clear instead of a confusing error.</summary>
    <content type="html">&lt;p&gt;MCP servers give an agent tools: search an issue tracker, query a database, read a wiki. Most are free to call, or are paid for by the account you connect with. A newer kind charges for each call. The server answers an unpaid call with a price and payment instructions, and the client is expected to pay, usually a few cents in a stablecoin, and try again. Two protocols for this have appeared, x402 and MPP, and Cloudflare's agents toolkit lets a server make any of its tools a paid one.&lt;/p&gt;
&lt;p&gt;Tempr doesn't pay for tools. This post is about what happens instead, because before this release that was a confusing error.&lt;/p&gt;
&lt;h2 id=&quot;what-used-to-happen&quot;&gt;What used to happen&lt;/h2&gt;
&lt;p&gt;An agent calls a tool. The tool says &amp;quot;pay first&amp;quot;. Each of Tempr's four MCP clients, in the CLI, VS Code, JetBrains and Visual Studio, turned that into some kind of failure, and none of them said what it was:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;The model saw an unexplained error, so it often tried the same call again, and again.&lt;/li&gt;&lt;li&gt;You saw a failed tool call with no reason.&lt;/li&gt;&lt;li&gt;In some clients it didn't even look like an error. The payment instructions came back as if they were the tool's answer, or as an empty result.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;Part of the reason is that there's no single way for a tool to refuse. We found four. The x402 MCP spec puts the payment terms in an error result. Cloudflare's toolkit puts them in the result's metadata. MPP uses a JSON-RPC error code. And a payment gateway in front of a server can answer with a plain HTTP 402, the status code the web reserved for &amp;quot;payment required&amp;quot; decades ago and barely used since. Our clients, and the MCP libraries under them, handled each of these differently.&lt;/p&gt;
&lt;h2 id=&quot;what-happens-now&quot;&gt;What happens now&lt;/h2&gt;
&lt;p&gt;When a tool asks to be paid, in any of the four forms, every Tempr client does the same thing:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;The agent is told plainly.&lt;/strong&gt; The tool's result says it requires payment, at what price when the server gives one, that Tempr doesn't pay for tools so it wasn't run, and not to call it again, but to tell you if the task needs it.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;You see it in the chat.&lt;/strong&gt; The tool call shows &amp;quot;Paid tool, not run&amp;quot;, with the price per call.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;It's asked once.&lt;/strong&gt; Later calls to the same tool are answered straight away, without contacting the server again, until the MCP servers restart.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Paid tools can be marked in advance.&lt;/strong&gt; Where a server lists a tool as paid, as Cloudflare's toolkit does, the agent sees that before its first call, so it can plan around the tool instead of discovering the price by trying it.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;A server that charges just to connect says so.&lt;/strong&gt; Its startup error explains that it requires payment, instead of looking like a broken server.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;The price is shown when the terms give it in USDC, a dollar stablecoin, or when a server lists a tool's price in dollars. When there are several options, the cheapest is shown. Anything else is shown as &amp;quot;a fee&amp;quot;.&lt;/p&gt;
&lt;h2 id=&quot;why-only-the-price-gets-through&quot;&gt;Why only the price gets through&lt;/h2&gt;
&lt;p&gt;A payment request carries text the seller wrote: a description of the resource, an error message, an address to pay. To the agent, that text would look like instructions from a tool it's using, and a seller could write anything there, including instructions meant for the agent. So nothing from a payment request reaches the model or your screen except the price, parsed as a number. The agent gets a message Tempr wrote, with that number in it.&lt;/p&gt;
&lt;p&gt;All four clients, and the Gateway, are tested against the same set of refusals in every form, including one with an injection attempt in its description, and the clients must produce exactly the same messages.&lt;/p&gt;
&lt;h2 id=&quot;through-the-gateway&quot;&gt;Through the Gateway&lt;/h2&gt;
&lt;p&gt;MCP traffic that goes through Tempr's Gateway gets the same treatment on the server side. When a remote MCP server answers with HTTP 402, the headers that carry its terms now reach your client unchanged, and the request log records the call as &lt;code&gt;mcp_payment_required&lt;/code&gt;, with the price when the terms give one in USDC. The Gateway doesn't pay either.&lt;/p&gt;
&lt;h2 id=&quot;why-not-just-pay&quot;&gt;Why not just pay?&lt;/h2&gt;
&lt;p&gt;For a few cents, paying might seem easy. But an agent that can pay is an agent that can spend your money on its own judgement, many times over in one run. Paying means holding a wallet and funding it, deciding who approves which purchases, and handling the support cases that follow when an agent buys something nobody meant it to. No customer has asked us for that from their coding agent.&lt;/p&gt;
&lt;p&gt;If that changes, the work done here still counts. A client that pays has to know a call is a purchase, must never retry one blindly, and should count it against the run's budget, the same &lt;a href=&quot;https://temprhq.io/blog/a-price-ceiling-for-agent-runs&quot;&gt;price ceiling&lt;/a&gt; that limits model spend. Until then, a paid tool is a tool your agent knows it can't use, and says so.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://temprhq.io/mcp&quot;&gt;MCP servers page&lt;/a&gt; lists servers you can connect, and the &lt;a href=&quot;https://temprhq.io/docs/mcp&quot;&gt;MCP docs&lt;/a&gt; show how to set one up.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>OpenAI's newest models, on the API everything else speaks</title>
    <id>https://temprhq.io/blog/openai-newest-models-on-chat-completions</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/openai-newest-models-on-chat-completions"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="Gateway"/>
    <category term="OpenAI"/>
    <summary>OpenAI's pro, codex and newest gpt-6 models do some things, or everything, only on its Responses API. How Tempr bridges chat completions requests to them, one request at a time, so your tools and agents keep working.</summary>
    <content type="html">&lt;p&gt;OpenAI has two APIs for talking to its models. Chat completions is the older one, and nearly every tool speaks it: IDE extensions, agent frameworks, scripts, and most gateways, Tempr's included. The Responses API is newer. OpenAI builds its new features there first, and more and more of its models can do some things only there.&lt;/p&gt;
&lt;p&gt;For a tool that speaks chat completions, that shows up as errors that are hard to explain. A model is in the model list, but every request fails. A request works until you turn reasoning up, then fails. An agent works until it tries to call a tool. We hit all three, so the Tempr Gateway now sends exactly those requests through the Responses API and translates the answer back.&lt;/p&gt;
&lt;h2 id=&quot;three-ways-a-request-can-fail&quot;&gt;Three ways a request can fail&lt;/h2&gt;
&lt;p&gt;We checked each model live, with real requests. There are three kinds of gap:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;The whole model is Responses-only.&lt;/strong&gt; OpenAI's pro and codex models, such as &lt;code&gt;gpt-5-pro&lt;/code&gt;, &lt;code&gt;o3-pro&lt;/code&gt; and &lt;code&gt;gpt-5.3-codex&lt;/code&gt;, answer every chat completions request with a 404.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;The deepest reasoning level is Responses-only.&lt;/strong&gt; &lt;code&gt;gpt-5.6&lt;/code&gt; with its &lt;code&gt;luna&lt;/code&gt;, &lt;code&gt;sol&lt;/code&gt; and &lt;code&gt;terra&lt;/code&gt; variants, and &lt;code&gt;gpt-6-astra&lt;/code&gt;, &lt;code&gt;gpt-6-luna&lt;/code&gt;, &lt;code&gt;gpt-6-sol&lt;/code&gt; and &lt;code&gt;gpt-6.1-sol&lt;/code&gt;, reject reasoning effort &lt;code&gt;max&lt;/code&gt; on chat completions. Every other level works on both APIs.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Function tools with reasoning are Responses-only.&lt;/strong&gt; On those same models, chat completions rejects any request with function tools unless reasoning is set to &lt;code&gt;none&lt;/code&gt;, and that includes requests that don't set reasoning at all. &lt;code&gt;gpt-6-astra&lt;/code&gt; has no &lt;code&gt;none&lt;/code&gt; level, so on chat completions it can't call a function tool at all.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;The third one matters most. Calling tools is what an agent does: read a file, run a command, edit code. A model that can't call tools on chat completions is no use to a coding agent that speaks chat completions.&lt;/p&gt;
&lt;h2 id=&quot;a-bridge-for-each-request&quot;&gt;A bridge for each request&lt;/h2&gt;
&lt;p&gt;Tempr decides which API to use for each request, not for each model. A &lt;code&gt;gpt-6-sol&lt;/code&gt; request with tools and reasoning goes to the Responses API. The same model's plain requests stay on chat completions, where parameters like &lt;code&gt;stop&lt;/code&gt; and &lt;code&gt;seed&lt;/code&gt; still apply. The Responses API has no equivalent of those, so a bridged request goes without them.&lt;/p&gt;
&lt;p&gt;On the way there, the chat request becomes a Responses request. Messages become input items in the same order. Assistant tool calls and their results become function-call items. Your reasoning setting becomes the Responses API's own. Two details took some care:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Tool schemas keep chat's behaviour.&lt;/strong&gt; Chat completions doesn't enforce a function's schema strictly unless you ask it to, and the Responses API does by default. A loose tool schema, as many IDE tools have, would be held to stricter rules than it was written for, so Tempr keeps chat's default.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;Reasoning survives a tool loop.&lt;/strong&gt; Tempr doesn't store responses, so it asks OpenAI for the model's reasoning in encrypted form, the only way to get it back without storage. It comes back as &lt;code&gt;reasoning_details&lt;/code&gt; on the assistant message, which Tempr's clients already return on the next request, and goes back to OpenAI from there. A model reasoning through a sequence of tool calls keeps its earlier reasoning.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;On the way back, the Responses answer becomes a chat completions answer, streaming included, with text, tool calls, finish reasons and usage where chat clients expect them. Tempr doesn't ask for reasoning summaries, though: OpenAI rejects the whole request when the API key's organization hasn't been verified, and we didn't want a request to fail over an extra.&lt;/p&gt;
&lt;p&gt;All this applies to every path into Tempr: the Gateway's &lt;code&gt;/v1/chat/completions&lt;/code&gt; and &lt;code&gt;/v1/messages&lt;/code&gt;, the IDE extensions, the CLI and agent runs.&lt;/p&gt;
&lt;h2 id=&quot;knowing-which-requests-to-send&quot;&gt;Knowing which requests to send&lt;/h2&gt;
&lt;p&gt;OpenAI's model list doesn't say which API a model needs, so Tempr works it out in three ways:&lt;/p&gt;
&lt;ol&gt;&lt;li&gt;&lt;strong&gt;By name.&lt;/strong&gt; An OpenAI model with &lt;code&gt;-pro&lt;/code&gt; at the end or &lt;code&gt;codex&lt;/code&gt; in its name goes straight to Responses, which saves a rejected request.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;From the error.&lt;/strong&gt; When chat completions answers that a model or a request only works on the Responses API, Tempr sends that request there instead, and remembers.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;From a list.&lt;/strong&gt; For the models whose gaps we've checked, the reasoning levels and tool requests chat completions rejects are listed, so those requests go to Responses the first time.&lt;/li&gt;&lt;/ol&gt;
&lt;p&gt;The second step needed a fix. Chat completions rejects a function-tool request on these models with an error that also says to use the Responses API. The first version of the bridge read that as &amp;quot;this model is Responses-only&amp;quot; and sent all of the model's requests there, so its plain requests lost &lt;code&gt;stop&lt;/code&gt; and &lt;code&gt;seed&lt;/code&gt; for no reason. Tempr now remembers what it learns separately for each model and kind of request: a model that only needs Responses for tools keeps using chat completions for everything else.&lt;/p&gt;
&lt;h2 id=&quot;new-models&quot;&gt;New models&lt;/h2&gt;
&lt;p&gt;New OpenAI models arrive often. Tempr picks up their reasoning levels on its own, from a catalog it refreshes every 12 hours, but not which requests they reject on chat completions. Checking that is quick and doesn't cost anything. Chat completions checks the model before the parameters, so a request with a reasoning level that doesn't exist, such as &lt;code&gt;&amp;quot;banana&amp;quot;&lt;/code&gt;, is rejected without generating any tokens, and the error says whether chat completions serves the model. A level the model doesn't have, such as &lt;code&gt;max&lt;/code&gt;, gets an error listing the levels it does have. That's how we found the same gaps on &lt;code&gt;gpt-6-luna&lt;/code&gt;, &lt;code&gt;gpt-6-sol&lt;/code&gt; and &lt;code&gt;gpt-6.1-sol&lt;/code&gt; as on &lt;code&gt;gpt-6-astra&lt;/code&gt;, and they joined the list on October 5.&lt;/p&gt;
&lt;p&gt;You don't need to do anything to use any of this: send your chat completions request as usual. If you'd rather use the Responses API yourself, Tempr's &lt;code&gt;/v1/responses&lt;/code&gt; passes OpenAI requests straight through. See the &lt;a href=&quot;https://temprhq.io/docs/gateway-chat-completions#responses&quot;&gt;Responses API section of the docs&lt;/a&gt;, and the &lt;a href=&quot;https://temprhq.io/docs/guide-codex-cli&quot;&gt;Codex CLI guide&lt;/a&gt; for a client that speaks it.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>One reasoning setting for every provider, and what it took</title>
    <id>https://temprhq.io/blog/one-reasoning-setting-every-provider</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/one-reasoning-setting-every-provider"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="Gateway"/>
    <category term="Reasoning"/>
    <summary>Every AI lab names its reasoning controls differently, and the same model can behave differently on two hosts. How Tempr gives you one setting that works everywhere, and what we found testing it.</summary>
    <content type="html">&lt;p&gt;Most models worth using for code can think before they answer, and you can usually choose how much. More thinking costs more tokens and takes longer, and it often gets you a better answer on hard problems. What you can't choose is how you ask for it, because no two labs ask the same way.&lt;/p&gt;
&lt;p&gt;OpenAI, xAI and Mistral take &lt;code&gt;reasoning_effort&lt;/code&gt;. Claude takes adaptive thinking with an effort level (or a token budget, before Claude 4.6). Gemini takes &lt;code&gt;thinkingConfig&lt;/code&gt;. DeepSeek, Z.AI, Moonshot, MiniMax, Xiaomi and Cohere each take a &lt;code&gt;thinking&lt;/code&gt; field of their own, and Qwen takes &lt;code&gt;enable_thinking&lt;/code&gt;. The levels don't line up either: one model offers &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt; and &lt;code&gt;high&lt;/code&gt;, another adds &lt;code&gt;minimal&lt;/code&gt; or &lt;code&gt;xhigh&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt;, and some can only be switched on or off.&lt;/p&gt;
&lt;p&gt;If you use one model, that's a detail. If you switch models, or send a request to whichever model fits it, the detail becomes a lot of code.&lt;/p&gt;
&lt;h2 id=&quot;one-setting-in-the-right-format-out&quot;&gt;One setting in, the right format out&lt;/h2&gt;
&lt;p&gt;The Tempr Gateway takes one shape for every model: OpenRouter's &lt;code&gt;reasoning&lt;/code&gt; object, or OpenAI's &lt;code&gt;reasoning_effort&lt;/code&gt; if you only want to set the level.&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;json&lt;/span&gt;&lt;pre&gt;&lt;code&gt;{
  &amp;quot;model&amp;quot;: &amp;quot;anthropic/claude-sonnet-4-6&amp;quot;,
  &amp;quot;messages&amp;quot;: [{ &amp;quot;role&amp;quot;: &amp;quot;user&amp;quot;, &amp;quot;content&amp;quot;: &amp;quot;Plan the migration.&amp;quot; }],
  &amp;quot;reasoning&amp;quot;: { &amp;quot;effort&amp;quot;: &amp;quot;high&amp;quot; }
}&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Tempr sends that on in the format the model's provider reads. The same goes for &lt;code&gt;/v1/messages&lt;/code&gt; (Claude's &lt;code&gt;thinking&lt;/code&gt; and &lt;code&gt;output_config.effort&lt;/code&gt;) and &lt;code&gt;/v1/responses&lt;/code&gt; (&lt;code&gt;reasoning.effort&lt;/code&gt;), so a tool built for one of those APIs can run on any model. We made four decisions along the way:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;Ask for nothing, get the model's default.&lt;/strong&gt; Leave reasoning out and each model thinks as much as its lab intends. Tempr doesn't switch thinking off behind your back to save tokens.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;A level the model doesn't have becomes the nearest one it does&lt;/strong&gt;, rounding up on a tie. Asking for &lt;code&gt;xhigh&lt;/code&gt; on a model that stops at &lt;code&gt;high&lt;/code&gt; gets you &lt;code&gt;high&lt;/code&gt;, not an error.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;The response says what was applied.&lt;/strong&gt; The &lt;code&gt;x-tempr-reasoning-effort&lt;/code&gt; header gives the level that was sent, so you can see when your request was adjusted.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;&lt;code&gt;none&lt;/code&gt; on a model that can't stop thinking gets its lowest level.&lt;/strong&gt; Claude Opus 5 gets &lt;code&gt;low&lt;/code&gt; as well, because Anthropic recommends a low effort over switching its thinking off.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;In the IDE extensions this is a reasoning picker next to the model picker, showing only the levels the selected model has, and your choice is remembered per model. In the CLI it's &lt;code&gt;/reasoning&lt;/code&gt;, also saved per model, or &lt;code&gt;--reasoning&lt;/code&gt; for a single run.&lt;/p&gt;
&lt;h2 id=&quot;the-same-model-on-three-hosts&quot;&gt;The same model, on three hosts&lt;/h2&gt;
&lt;p&gt;Labs were the easy part. Open-weight models are served by many inference hosts, and each host decides for itself which reasoning fields it reads. We tested each host with real keys in September 2026, and the documentation didn't always match what the APIs did. The clearest case was GLM-5.3, one model on three hosts:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;On &lt;strong&gt;DeepInfra&lt;/strong&gt;, the levels worked as expected. Thinking couldn't be switched off, and trying to made the model's reasoning show up in its visible answer.&lt;/li&gt;&lt;li&gt;On &lt;strong&gt;Together AI&lt;/strong&gt;, the levels behaved the opposite way: &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;high&lt;/code&gt; produced no thinking, while &lt;code&gt;max&lt;/code&gt; and &lt;code&gt;none&lt;/code&gt; did. A picker there would switch thinking off for anyone who chose &amp;quot;high&amp;quot;.&lt;/li&gt;&lt;li&gt;On &lt;strong&gt;Fireworks AI&lt;/strong&gt;, the levels formed a proper ladder, and &lt;code&gt;none&lt;/code&gt; was rejected with a clear error.&lt;/li&gt;&lt;/ul&gt;
&lt;p&gt;So Tempr reads reasoning levels from each host's own catalog entry, not from the model's lab, and the same model can offer different levels, or none, depending on where it runs. GLM on Together AI offers no levels at all, because the host can't steer it reliably. Hugging Face offers none either: it's a router that picks a host for each request, and of the 20 reasoning models we checked there, none had only one host behind it.&lt;/p&gt;
&lt;p&gt;The hosts differ in other ways too. Fireworks rejects any field it doesn't recognise, including the &lt;code&gt;reasoning&lt;/code&gt; object, so Tempr sends it only &lt;code&gt;reasoning_effort&lt;/code&gt;. AWS Bedrock's OpenAI-compatible endpoint ignores the object and reads only &lt;code&gt;reasoning_effort&lt;/code&gt;, while Claude and Amazon Nova on Bedrock's Converse API each take their own fields and reject each other's.&lt;/p&gt;
&lt;h2 id=&quot;testing-it-without-paying-for-it&quot;&gt;Testing it without paying for it&lt;/h2&gt;
&lt;p&gt;Checking every model on every host would be expensive if every check generated tokens. Most don't need to. A host checks your parameters only after it has accepted the model, and a rejected request costs nothing. Send a deliberately invalid level, such as &lt;code&gt;&amp;quot;reasoning_effort&amp;quot;: &amp;quot;banana&amp;quot;&lt;/code&gt;, and the error says whether your key can reach the model and, often, which levels it accepts. Our whole Together AI investigation cost less than a tenth of a cent.&lt;/p&gt;
&lt;h2 id=&quot;keeping-it-right&quot;&gt;Keeping it right&lt;/h2&gt;
&lt;p&gt;Model catalogs change. Tempr takes reasoning levels from models.dev's open catalog, refreshed every 12 hours, and corrects it where a provider's API behaves differently. That correction has already mattered. On October 1 the catalog started listing Kimi K3 as a model that always reasons, and for a few days &lt;code&gt;none&lt;/code&gt; ran it at its lowest level. Moonshot's API still switches it off, so Tempr does again, and behaviour we've checked live now overrides the catalog.&lt;/p&gt;
&lt;p&gt;The full reference, including every field and how each provider maps it, is in the &lt;a href=&quot;https://temprhq.io/docs/gateway-chat-completions#reasoning&quot;&gt;Reasoning section of the docs&lt;/a&gt;. Each model's levels are in &lt;code&gt;GET /v1/models&lt;/code&gt; under &lt;code&gt;reasoning_options&lt;/code&gt;, and in the &lt;a href=&quot;https://temprhq.io/models&quot;&gt;model catalog&lt;/a&gt;.&lt;/p&gt;</content>
  </entry>
  <entry>
    <title>Putting a price ceiling on an agent run</title>
    <id>https://temprhq.io/blog/a-price-ceiling-for-agent-runs</id>
    <link rel="alternate" type="text/html" href="https://temprhq.io/blog/a-price-ceiling-for-agent-runs"/>
    <published>2026-10-07T00:00:00Z</published>
    <updated>2026-10-07T00:00:00Z</updated>
    <author><name>The Tempr team</name></author>
    <category term="CLI"/>
    <category term="Chat"/>
    <category term="Portal"/>
    <category term="Budgets"/>
    <summary>An agent run is a loop of model calls whose cost you can't know in advance. How Tempr lets you, and your organization, cap what one run may spend, and see which runs cost the most.</summary>
    <content type="html">&lt;p&gt;A chat reply is one model call, and you can roughly guess what it costs. An agent run isn't. You give it a task, and it reads files, runs commands, calls the model again with what it found, maybe hands part of the work to a sub-agent, and keeps going until it's done. One prompt can turn into a few calls or a few dozen, each carrying a longer conversation than the one before.&lt;/p&gt;
&lt;p&gt;Monthly budgets are a good backstop, but they catch a runaway run only after it has spent the money. What you want is a limit for each run: stop this task once it has cost this much.&lt;/p&gt;
&lt;h2 id=&quot;a-limit-for-one-run&quot;&gt;A limit for one run&lt;/h2&gt;
&lt;p&gt;In the CLI, it's a flag:&lt;/p&gt;
&lt;div class=&quot;blog-code&quot;&gt;&lt;span class=&quot;blog-code-lang&quot;&gt;bash&lt;/span&gt;&lt;pre&gt;&lt;code&gt;tempr --max-cost 0.50 &amp;quot;migrate the payment tests to the new fixtures&amp;quot;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The run keeps a total of what its model calls have cost, sub-agents included, from the cost each call reports. Once the total reaches the limit, the run stops before its next model call, the same way Ctrl+C stops a turn. The CLI exits with code 6 and status &lt;code&gt;cost-limit&lt;/code&gt;, so a script or CI job can tell this apart from an error. With &lt;code&gt;--json&lt;/code&gt;, the final result line also carries &lt;code&gt;spentUsd&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;The IDE extensions have the same limit as a setting: &amp;quot;Stop an agent turn after (US dollars)&amp;quot; in VS Code, JetBrains and Visual Studio. It's blank by default.&lt;/p&gt;
&lt;h2 id=&quot;enforced-on-the-server-too&quot;&gt;Enforced on the server too&lt;/h2&gt;
&lt;p&gt;A limit that only the client checks has gaps. The client doesn't see every call billed as it happens, and a run that's stopped on your machine can still have work going on the server. So Tempr's server enforces the limit as well. Each run is sent what is left of its limit, and the server stops the run before any model call that would start once the limit is reached.&lt;/p&gt;
&lt;p&gt;Two things to know about the edges:&lt;/p&gt;
&lt;ul&gt;&lt;li&gt;&lt;strong&gt;The limit is checked between calls&lt;/strong&gt;, so a run can finish a little above it, by the cost of the call that was already running when it crossed the line. Set the limit with that call in mind.&lt;/li&gt;&lt;li&gt;&lt;strong&gt;A model with no known price can't be counted.&lt;/strong&gt; Tempr prices a call from what the provider reports or from the model's listed price. When it has neither, as with a custom model you haven't priced, the call counts as $0, and the CLI warns you rather than pretending the limit still holds.&lt;/li&gt;&lt;/ul&gt;
&lt;h2 id=&quot;a-limit-for-your-whole-organization&quot;&gt;A limit for your whole organization&lt;/h2&gt;
&lt;p&gt;On a team, the person setting the limit often isn't the one running the agent. An organization's owners and admins can set the most one agent run may cost on the Portal's Budget page, and it applies to every member, in the CLI and in the IDE extensions. Members can set a lower limit for themselves, but not a higher one, and &lt;code&gt;--max-cost&lt;/code&gt; can only lower it. Changes to the limit appear in the organization's audit log.&lt;/p&gt;
&lt;p&gt;When a run hits the organization's limit, only that turn stops. The session carries on, and the next run starts with a fresh limit, so a developer can look at what happened and decide whether to continue.&lt;/p&gt;
&lt;h2 id=&quot;seeing-where-the-money-went&quot;&gt;Seeing where the money went&lt;/h2&gt;
&lt;p&gt;A limit stops a run. It doesn't tell you which runs are worth a closer look. The Portal's Usage page now lists this month's most expensive agent runs, with when each started, who ran it, which model it used, how many model calls it made and what it cost, sub-agents included. On a team, the same list appears on the Budget page for owners and admins. Each run links to its trace, so you can see which step took the most.&lt;/p&gt;
&lt;p&gt;One more budget tool shipped alongside this. Monthly budgets on virtual keys, on members and on an organization's hard cap now have a &lt;strong&gt;Strict&lt;/strong&gt; option. Without it, several requests running at the same moment can each see room left in a budget and, together, go past it. With it, they can't. Strict mode adds about 5 milliseconds per request, and up to about 30 when many requests share a budget at once, so it's off unless you turn it on.&lt;/p&gt;
&lt;p&gt;The details are in the docs: &lt;a href=&quot;https://temprhq.io/docs/cli-scripts#max-cost&quot;&gt;capping what a run spends&lt;/a&gt; in the CLI guide, and the &lt;a href=&quot;https://temprhq.io/docs/cli-reference&quot;&gt;CLI reference&lt;/a&gt; for every flag.&lt;/p&gt;</content>
  </entry>
</feed>
