A sandbox for the agent's commands, with nothing to install

  • CLI
  • Security

A coding agent is most useful when it can run things: build the project, run the tests, try a fix, run the tests again. It's most worrying for the same reason. Every command it runs can do anything you can do on that machine.

Approval prompts are the usual answer, and they work while you're watching. They don't work for a run nobody is watching, in CI or a long task you've left alone. An allow-list helps, but only for the commands you thought of in advance. The agent's next idea, a script it just wrote or a tool it wants to try, stops the run.

The Tempr CLI's sandbox turns that around. Instead of deciding which commands may run, it limits what any command can do, and then lets the agent run whatever it needs.

What the sandbox allows

Inside the sandbox, the agent's commands:

  • can change files only in the current folder, the temp folders and the package caches builds write to: npm's, pip's, NuGet's, Go's, Cargo's, Gradle's and Maven's. Everything else, your home folder, /etc, other projects, is read-only to them.
  • can read anything, as a build needs to.
  • can't open network connections, if you choose. Cutting off the network blocks outgoing TCP connections, while name lookups still work.

Because the damage a command can do is limited, any command runs without asking, in every approval mode, even one with no allow-list entry. A few still ask: commands that delete, publish or force-push, because those can do harm inside the project too. When the sandbox blocks something a command really needed, such as a download, the agent can ask to run that command outside the sandbox, and that always asks you first. A run with nobody to ask refuses it.

You turn it on with --sandbox, or for every run with tempr config set-sandbox on. Since CLI 0.9.8 it's also the default for runs where nobody decides command by command, --yolo and autopilot in CI, with the network left on there. Cutting the network also blocks connections to localhost, and a test suite that starts a local server needs those.

Why not bubblewrap

The usual tool for sandboxing a process on Linux is bubblewrap, which builds an isolated view of the system using the kernel's namespaces. We tried it first, on the systems our users actually run, and it failed on two of them:

  • Stock Ubuntu 24.04, a very common Linux on developer machines and CI runners, restricts the unprivileged user namespaces bubblewrap needs. As a normal user, it couldn't start.
  • A default Docker container, which is where many CI jobs and our own container recipe run, doesn't allow creating new namespaces either.

A sandbox that's missing in the two places it matters most isn't much of a sandbox. So we built ours on Landlock, a security feature in the Linux kernel itself. With Landlock, a process restricts itself: it declares which folders it may write to and whether it may connect out, and from then on neither it nor anything it starts can do more. It needs no install, no root and no namespaces. It works on stock Ubuntu 24.04 and inside a default Docker container. Restricting files needs Linux 5.13 or later, and restricting the network needs 6.7 or later. Inside WSL2 on Windows, it works too.

The parts Landlock can't do

Landlock grants access to whole folders. The project folder has to be writable, or nothing could build. But a few files inside it can do harm later, outside the sandbox: git hooks run on your next commit, git's config can point at a program to run, and .agentcommands.json is the agent's own allow-list. A command inside the sandbox could change any of them, and the change would take effect after the sandbox was gone.

So Tempr takes a copy of .git/hooks, .git/config and .agentcommands.json before each sandboxed command, and puts back anything the command changed. The agent is told when that happens, so it doesn't build on a change that's been undone.

Making real builds work

A sandbox that blocks a normal build is one people turn off, so we spent most of our time on what builds actually need to write.

Package caches were the obvious part. A restore or install writes to a per-user cache outside the project, so those folders are writable: npm's, Cargo's and Maven's from the first release, NuGet's and Go's module cache from 0.9.8, and pip's and Gradle's too.

.NET took more work. The dotnet tool needs to write small marker files in ~/.dotnet, and fails if it can't. The documented switch to skip that didn't help in our tests on .NET 10. Making ~/.dotnet writable wasn't an option either, because it holds global tools, programs that run outside the sandbox later. So inside the sandbox, dotnet gets its own folder for those files, in ~/.cache/tempr/sandbox-dotnet. Then it turned out NuGet follows that same setting for its packages and configuration, so Tempr points NuGet back at your real ~/.nuget/packages and package sources. The result is that dotnet builds, restores and tests inside the sandbox, even on a machine where it has never run before, while your global tools stay read-only.

Checked on real machines

Both releases that changed the sandbox were checked on fresh cloud machines running stock Ubuntu 24.04, as a normal user, on both Intel and Arm (AWS Graviton) processors. Each check confirmed that writes to the project and the caches work, writes to the home folder and /etc are denied, the network is blocked when it should be and open when it should be, and pipes still work. One lesson from that: an emulated Arm machine reported no Landlock at all, so it couldn't prove anything. Only real hardware could.

Where it runs, and where it doesn't yet

The sandbox runs on Linux. On macOS and Windows it isn't available yet. What happens there depends on whether you asked for it:

  • If you asked for it with --sandbox and it can't run, nothing runs. The CLI exits with code 7 instead of running your commands unsandboxed.
  • The default for --yolo and CI wasn't something you asked for, so where the sandbox can't run, those runs go ahead as they did before it existed, and an interactive session tells you so.

tempr doctor tells you what the sandbox can do on your machine. It covers the commands the agent runs, the .NET test tools and the agent's diagnostics. It doesn't cover commands you type yourself with !command, or MCP servers.

The details are in the sandbox section of the CLI reference, and Scripts, pipes and CI shows it in a pipeline. For more on running the CLI unattended, see Pipe it to the agent.