Hands Off Development

Isolated VM Environments for Coding Agents

Agents need isolated execution to do real engineering work, not just run constrained scripts.

Contributing Editor, Cloud Economics & Strategy · · 11 min read
Cover illustration for “Isolated VM Environments for Coding Agents”
Cloud Agent Environments · October 6, 2026 · 11 min read · 2,430 words

Coding agents need more from their execution environment than any earlier generation of automated tooling, because they don't just run predetermined scripts. They install packages, start services, read and write files, run tests, and check their own output, all in a single session. If an agent cannot install a dependency, open a port, or run a build pipeline mid-task, it is a constrained autocomplete wrapper with a chat interface bolted on, since a real coding agent requires exactly that freedom from its environment to do engineering work that holds up.

This reframes environment selection as a decision about capability, not just about infrastructure or security policy. If you pick the wrong environment, the agent isn't merely exposed to more risk. It loses the ability to do parts of the job a human engineer would consider basic: spin up a local service, hit an endpoint, confirm a migration applies cleanly, open a browser and check that a page renders. Every one of those actions depends on the environment underneath the agent, and none of them are optional if the goal is code a team can actually merge.

What isolation means in an agent execution context

"Isolated environment" Isolated environment" gets used as a vague safety label more often than it gets defined. Isolation in an agent execution context breaks into four distinct boundaries, and all four have to hold at once for the claim to mean anything. Windmill's sandbox design, when enabled, makes this visible by showing what's missing without it: an agent lacking isolation has access to the host filesystem, to environment variables including credentials, to other jobs running alongside it, and to the network without restriction. Isolation means closing all four of those doors, not just one or two of the more obvious ones.

Resource limits belong in the same bundle. CPU, memory, and disk usage need hard caps per execution, so if one agent task runs away, it can't starve every other job sharing the same infrastructure. An environment that isolates filesystem and network access but lets one process consume all available memory hasn't solved the problem, it's just relocated it.

None of this makes the environment restrictive in the sense of limiting what the agent can build. A fully isolated environment still gives the agent a complete Linux userspace: it can install packages, start background services, and expose ports exactly as it would on an unsandboxed machine. None of those actions reach anything outside the boundary. The agent gets the full range of a real operating system, and everything it does inside that range stays contained to that one execution.

Standard Linux containers and the isolation bar for untrusted agent code

The default instinct for most engineering teams is to reach for containers, since containers already run most production workloads and nobody thinks twice about them day to day. For agent-generated code, that default doesn't hold up. Standard Linux containers share the host kernel. Namespace separation is the mechanism containers use to keep processes apart, but it's a logical boundary enforced in software, not a hardware boundary. A container escape hands an attacker, or a misbehaving agent, direct access to the kernel and to everything else running on that host.

microVMs and VM-based runtimes give each workload a dedicated kernel, while containers rely on namespace separation alone, a categorically weaker boundary with a documented CVE history even in rootless configurations. The distinction isn't theoretical. When each workload gets a dedicated kernel, a compromise inside one workload has no kernel-level path to any other workload on the same machine. Namespace separation gives you no such guarantee, because it depends on every layer of the shared kernel holding correctly, all the time.

The stakes are higher for coding agents specifically because of how the code gets there in the first place. An agent generates code autonomously and then runs it in the same session, so you don't know the code being executed in advance, and you can't statically check it for safety before it runs. The isolation boundary has to be strong enough to contain the worst case the agent could produce, not the average case.

A team that has already run containers in production for years without incident might reasonably ask what the fuss is about. The absence of an incident so far isn't evidence the boundary is sufficient; it's evidence the bad case hasn't landed yet. Risk here is tail-heavy: a container escape is rare but severe, and an agent generating and running code at volume, across many sessions, changes the odds in a way that a small number of human-reviewed deploys does not. gVisor intercepts system calls in user space, cutting down the kernel attack surface without requiring a fully separate kernel per workload. It trades some performance for meaningfully stronger isolation than a plain container, without the full overhead of a microVM. For GPU workloads, where VM passthrough can carry real overhead, container-based sandboxes may be a reasonable compromise. For the core coding-agent workload, where the agent executes code it just wrote, you should start from hardware-enforced boundaries, not treat them as the exception.

The three isolation architectures teams choose between

Diagram: Three Isolation Architectures: What Each Trades Away. Visualizes: Show a ranked comparison of the three isolation architectures for coding-agent workloads: microVMs (Firecracker, Kata Containers), gVisor (user-space kernel interception)…

Once containers are off the table as the default, three real architectures remain, and the choice among them comes down to workload shape, not just security posture. These are microVMs (Firecracker, Kata Containers), user-space kernel interception (gVisor), and hardened containers, and each trades differently on startup speed, operational overhead, and what kinds of workloads it fits.

MicroVMs give each workload a dedicated kernel, which is the strongest form of isolation on this list: a kernel-level exploit inside one workload has no boundary to cross because there's no shared kernel to exploit. The cost people expect to pay for that is slow startup, and it turns out not to be true at this level. Firecracker and Kata Containers achieve sub-second boot times. The overhead a team might assume comes with "a full virtual machine per task" is close to negligible for most agent task durations.

gVisor takes a different approach: instead of a dedicated kernel per workload, it intercepts system calls in user space, cutting the kernel attack surface without the cost of standing up a separate kernel for every job.

Hardened containers are rootless configurations layered with seccomp or other kernel-level security policies; they reduce the risk a plain container carries but still share the host kernel, which remains exploitable. That makes them a fit for workloads where the code being run can be partially trusted, but not the right default for general coding-agent use, where the agent writes code and immediately executes it with no human review in between.

What a properly configured VM environment makes possible

Isolation up to this point has been described in terms of what it blocks. The more interesting question for a team evaluating infrastructure is what it opens up. A properly isolated VM environment doesn't restrict an agent's capabilities, it completes them. Inside that boundary, the agent can install arbitrary packages, start and talk to background services, expose and test ports, run a full build pipeline, drive a browser, and check its own output, all within one session, without any of that activity touching the host or any other tenant on the platform.

Getting that right across a single task is only half the problem. If an agent loses its working state between tool calls, it can't reliably finish anything that spans multiple steps, so the environment needs persistent volumes that carry state across executions, so the agent doesn't have to rebuild its working directory from nothing each time. Windmill's volume design shows one concrete way to handle this: before each execution, files sync down from object storage into the sandbox; after execution, changed files sync back; the agent reads and writes normally in between. So an agent can carry state across runs through that three-phase cycle, and the orchestration layer never has to manage it by hand.

Pre-loaded environments, where a team's actual dependencies, tooling, and configuration already exist in the sandbox before the agent starts, remove a common source of agent failure: setup steps that eat context, fail quietly, or drift from run to run until the environment no longer matches what the team actually runs in production.

Self-verification deserves particular attention here, because it's the capability most directly tied to whether an agent's output is actually correct, and it only exists where the agent can run the code it just wrote. Running the test suite, starting a server and hitting an endpoint, confirming a migration applies without error: these aren't conveniences layered on top of the core job. They are the mechanism by which an agent confirms its own work before anyone else looks at it. Modal's infrastructure guide describes how Ramp uses sandboxed environments for exactly this purpose, to power background coding agents that generate changes and write them back as commits or pull requests. The output that reaches a human reviewer is a reviewable artifact that already passed through verification, not code that ran loose in a shared environment and happened to work.

How environment configuration affects agent reliability

The environment an agent runs in is a direct input into the quality of what that agent produces, not a separate concern bolted on for compliance reasons. An agent running in an environment that doesn't match a team's actual stack will install the wrong dependency versions, miss platform-specific behavior, and hand back output that fails in the real codebase even after it passed every check the agent ran on itself.

This gap between the sandbox and the production environment is the coding-agent version of "works on my machine." An agent that verified its output against a generic Linux baseline has not verified the same thing as an agent that verified its output against a configured replica of the team's actual stack, down to the dependency versions and the linter configuration.

Research on scalable oversight of coding agents gives this a structural frame: unconstrained agents erode codebase scalability and make human review more expensive over time, and the methods that have managed large human engineering teams for decades, access control, network policies, and coding conventions enforced by tooling, carry over directly to coding agents. That same research also found you can run these constraints cheaper, measured in tokens, than recent agentic scaffolding built to achieve similar ends through prompting alone. The environment is what enforces these constraints in practice; an agent cannot reliably impose them on itself from inside a session.

If you pre-load the sandbox with a team's actual dependencies, linters, type checkers, and test runners, an agent's self-verification reflects real conditions instead of a generic baseline that tells the team nothing about how the code will behave once merged. This is the same principle that makes reproducible build environments valuable for human engineers, and the agent gains from that discipline for the same reasons a human team does. Network policy inside the sandbox carries the same weight for reliability as it does for security: an agent free to make arbitrary outbound calls during a task might fetch a package version that doesn't match the team's lock file, producing environment-specific behavior that's hard to reproduce and harder to debug later.

The oversight problem isolation solves on the review side of the agent workflow

Isolated VM environments are what make the background-agent model work at all. Because the agent's work is fully contained, its output can be packaged as a pull request, a discrete artifact a human can read and review, rather than as changes applied straight into a shared environment where tracing what happened after the fact becomes guesswork.

The oversight research cited above makes a direct claim that applies here: coding agents introduce security risk and erode codebase scalability when they run unconstrained, and the enforced environment around them is what makes review at scale possible in the first place. Recall rose from 54.5% to 90.9% when a constrained substrate plus a 200-line-of-code documentation CLI were added on top of an unconstrained agent with no tools, with the two interventions contributing independently. That's a near doubling of a reviewer's ability to catch planted problems, driven by changes to the substrate the agent runs on.

The isolation boundary also creates a natural point of attribution. Every artifact that leaves the sandbox, a commit, a pull request, a log, is the product of one bounded, recorded execution, which makes it possible to trace what the agent had access to and what it produced from that access.

Isolation solves environment consistency, but it does nothing about how fast agent-generated pull requests pile up against a human review queue that can only move so fast. That constraint sits outside what environment architecture can fix. The answer to it is better review tooling and process, not a weaker environment boundary. Loosening the sandbox to move PRs through faster trades a solved problem for an unsolved one.

Diagram: Recall Nearly Doubled With a Constrained Substrate. Visualizes: Show a before/after magnitude contrast: an unconstrained coding agent with no tools achieved 54.5% recall; adding a constrained substrate plus a 200-line-of-code documentation…

What teams evaluating sandbox infrastructure should check

A team evaluating sandbox infrastructure for coding agents has five concrete things to test, in this order: isolation model, environment configurability, persistence model, startup speed at scale, and compliance posture.

Isolation model comes first because it determines what kind of boundary a team is actually buying. Does the platform use microVMs or gVisor for CPU workloads, or does it lean on standard container namespace separation dressed up with extra policy? That answer decides whether the boundary is enforced by hardware or by software convention, and the difference matters most in exactly the worst-case scenario a team hopes never to see.

Environment configurability comes next: can a team's actual dependencies, tooling, and configuration be loaded into the sandbox image before the agent ever starts a task? When the baseline is generic, an agent's self-verification can't mean anything close to production behavior.

Persistence model matters just as much. Does the platform support volumes that survive across executions, and does it manage leases correctly so that concurrent agent sessions don't corrupt shared state? Windmill's exclusive-lease model, where a volume locks to one job at a time and other jobs wait for the lease to release, is one concrete answer to this problem and a useful pattern to understand even when the platform under evaluation handles it differently.

Startup speed and concurrency round out the list. Sub-second boot times matter for interactive agent workflows, because a slow sandbox turns into a slow agent from the user's point of view, and if you run many concurrent agent sessions, you need to check a platform's stated concurrency limits and creation-rate limits against your own actual workload before committing to it.

Sources

  1. Steerability via constraints: a substrate for scalable oversight of coding agents

More in Cloud Agent Environments