Network Access Controls for Sandboxed Agent Environments
Sandboxed agents need network controls designed for their specific task, not blanket policies.

Network access controls are the most under-configured layer in sandboxed agent environments, and the reason has nothing to do with negligence. A coding agent needs to install packages, call APIs, clone repositories, and sometimes drive a browser, and every one of those legitimate functions requires outbound network access. A blanket egress block would stop the agent from doing its job, so teams end up leaving the network layer looser than every other part of the sandbox, often without fully deciding to.
The threat surface here differs categorically from a chatbot's. A chatbot that misbehaves produces bad text. These agents run shell commands, clone repositories, and call APIs, making network isolation a first-class security concern. Teams running agents in cloud environments like Replicas, where each task runs in its own isolated Linux VM with full package installation and outbound capability, face this tension directly: the same infrastructure that lets agents do real work also opens multiple exit vectors unless network boundaries are enforced with care.
Two incidents from production make the stakes concrete. A sandbox can have strong file permissions, careful credential scoping, and a hardened execution environment, and still leak through a network path nobody thought to close.
Why the zero-trust model is the right frame
Zero-trust is the correct design principle for agent network policy: no outbound connection succeeds unless it sits on an explicit allowlist, and everything else is blocked by default. For most enterprise network architecture, this principle is well established and reasonably straightforward to apply, because the set of systems a given workload needs to reach is fixed and known in advance. Agents break that assumption in a specific way.
An agent with shell access runs with the ambient authority of its process. A package install through npm, pip, or cargo can touch dozens of registry hosts and CDN endpoints, many of them not obvious until the install actually runs. An agent calling a team's internal APIs needs live access to endpoints that may not be enumerated in advance, especially as those APIs evolve.
Default deny is hard to implement correctly for reasons that have nothing to do with teams skipping the step. Building a complete allowlist requires knowing every network dependency the agent will ever have, and that is itself an open-ended problem that shifts every time a new package gets installed or a new integration gets added. The practical resolution is to scope the allowlist to what the agent actually needs for its specific job and to treat any excess authority as attack surface to be closed. An agent fixing a frontend bug needs a different network policy than one auditing infrastructure, and because the policy matches that specificity, it closes far more unnecessary attack surface than a single team-wide list ever could.
How isolation technology choices constrain and enable network policy
The isolation technology underneath an agent sandbox sets the ceiling on how finely network policy can be enforced, and the three main approaches differ sharply in what they offer. Standard Docker containers share the host kernel. That gives it isolation stronger than a standard container, and the syscall interception point doubles as a natural chokepoint for enforcing network policy. Firecracker microVMs go further still, giving each workload a dedicated kernel. An attacker has to escape both the guest kernel and the hypervisor to reach anything else, and because network interfaces to the guest VM are explicitly provisioned, "no network by default" becomes architecturally natural.
OS-level policy primitives operate at the process level and shape what any of these three approaches can enforce. AppArmor restricts files, network, and processes through policy profiles, while bubblewrap restricts them through namespace isolation, with its security boundary determined entirely by the arguments passed to it at launch.
The choice between standard containers, gVisor, and microVMs is not just a performance decision. At the syscall layer or the kernel boundary, network policy can be enforced directly; otherwise it requires external tooling. Teams running agents in plain containers who want strict egress filtering have to bolt network policy on from the outside, using iptables rules, CNI plugins, or a proxy layered on top of infrastructure that wasn't designed with that boundary in mind. Teams running microVMs can design the network interface from scratch, and the guest simply has no route unless one is explicitly provisioned. Platforms that run agents in dedicated microVMs, where each task gets its own kernel and the network interface is explicitly provisioned, can enforce zero-trust architecture at the hardware boundary. Teams using shared-kernel containers are left layering external policy on top, which makes the security model harder to reason about and easier to misconfigure.
Harness choice adds another variable. Codex ships OS-level isolation primitives by default, using Seatbelt on macOS and bubblewrap, Landlock, and seccomp on Linux, making it one of a small number of major harnesses whose codebase shows built-in OS-level network restrictions. Most other harnesses leave the outer boundary entirely to the operator running them.
The two-phase network model: setup vs. execution
Splitting an agent's lifecycle into a setup phase and an execution phase is the most practical architectural pattern available for making default deny workable without crippling what the agent can do. Package installation through pip, npm, cargo, or apt has legitimately broad and partially unpredictable network needs, touching CDNs, mirrors, and transitive dependencies that are difficult to enumerate in full ahead of time. Trying to allowlist all of that in advance is error-prone and tends to produce either a list full of holes or one so broad it defeats the purpose.
Execution-phase needs look different. Splitting the two phases means a team doesn't need to solve the hard problem of allowlisting all of package-install-land in order to implement the control that actually matters: restricting what the agent can reach while it is doing the work it was asked to do. Codex's own cloud environments implement a version of this, running behind an HTTP and HTTPS network proxy for security and abuse prevention that applies across both phases, with all outbound internet traffic passing through it.
A proxy sitting at the execution-phase boundary turns a distributed, hard-to-audit ruleset into a single inspectable, deniable chokepoint. Every outbound call during execution passes through one place rather than being scattered across a complex set of rules spread across the infrastructure. A proxy configured this way can enforce host-level allowlists, log every request with its timestamp and target, and block requests to unknown IPs over any port. An agent writing a Python script has no legitimate reason to call an unknown IP address over port 443, and a proxy that blocks this behavior by default catches exfiltration attempts that would otherwise succeed quietly.
Designing an egress allowlist that doesn't break the agent and doesn't leave holes
A well-designed egress allowlist starts from the agent's actual task scope rather than from a generic list of domains presumed safe, and building it properly surfaces network dependencies a team didn't know it had. The agent's actual job should narrow and specify what network calls it needs to complete that particular task. A coding agent fixing a test typically needs the repository host, whether GitHub or GitLab, possibly an internal API it is testing against, and the package registry if installs happen during execution. An agent integrating with Slack, Linear, or an internal ticketing system needs those specific API endpoints, which tend to be enumerable, stable, and well documented, making them the easiest category to allowlist cleanly.
Running the agent in a logging-only mode, capturing every outbound host it contacts, reviewing that list, and then converting to deny-by-default using it as the allowlist reveals the dependencies that would otherwise cause breakage later: the registry mirror nobody remembered, the CDN sitting behind a package, the auth endpoint hiding behind an API call.
Three mistakes appear repeatedly once teams start building these lists. Allowlisting by IP address rather than hostname creates fragility, since IP ranges for major CDNs and cloud providers rotate over time, while hostname-based rules stay more stable and easier to audit. Because the setup and execution phases have genuinely different requirements, each needs its own allowlist; merging them back into one list erases the distinction that made the two-phase architecture useful.
Browser-driving agents need separate handling, because by design they have to reach arbitrary URLs, so allowlisting at the URL level is a non-starter. An unsandboxed browsing agent can be prompted by hidden content on a page to download and execute a binary and connect to a command-and-control server, an attack chain that works because the agent's network access was left unrestricted.
What the audit log needs to capture
If an egress allowlist has no immutable audit log, nobody can verify it, and an unverified control tends to fail silently rather than loudly, which is worse. Every sandbox should emit a log of every network request, every shell command, and every file write, and the network request log specifically is what tells a team whether the allowlist is being enforced and what the agent is actually reaching in practice.
Each network request in that log needs to carry a timestamp and the agent run identifier, so a specific network call can be traced back to the specific task that triggered it. It needs the destination hostname and port, and a clear record of whether the request was permitted or denied by the allowlist. It needs to identify the process or tool inside the agent that made the call, whether that was a shell command, a browser, or an SDK call, since that detail is often what makes a denial actionable.
The log earns its keep over time. Patterns across many runs help identify which kinds of agent tasks generate unexpected network calls, which is often the clearest signal that a task deserves its own narrower allowlist rather than sharing a global one with every other job the agent performs.
Attribution down to the task level is what makes all of this usable, since knowing that a network call came from "a coding agent run" tells a team very little. Knowing it came from a specific harness and model, triggered by a specific Linear issue, using a specific credential, turns remediation into something targeted. Replicas provides per-run analytics attributable to source, person, harness, model, credential, and the skills and MCP servers the agent reached for, the same attribution model that makes cost tracking useful making network audit trails something a team can act on.
How isolation architecture distributes the blast radius
Even a carefully designed egress allowlist will fail eventually, whether through misconfiguration, a zero-day in a proxying layer, or an allowlisted host that itself gets compromised. Because that failure is a matter of when rather than whether, the architecture around the network control has to limit how much damage one failure can cause. Blast radius is the measure of how much can go wrong when one part of the system fails, and network controls reduce that blast radius only when the execution environment underneath them is also isolated.
If a network policy runs on a shared container host, where multiple agent runs coexist on the same kernel, it limits outbound calls but does nothing to stop lateral movement between those runs if one of them is compromised. A network policy enforced on a per-run microVM limits both outbound calls and lateral movement at once, so if that VM's network control fails, the damage stays scoped to that one run's environment.
Per-run isolation is the structural answer to this problem. Firecracker microVMs give each workload a dedicated kernel, so an attacker has to escape both the guest kernel and the hypervisor, not merely the network policy layer sitting on top. Each run gets its own network namespace, so there is no shared network state that a compromised run could read from or poison for another run sharing the same infrastructure.
A practical defense-in-depth stack for a coding agent combines execution in an isolated VM per run rather than a shared one, a network interface provisioned explicitly with no default route and only the allowlisted routes added, execution-phase traffic routed through a proxy that enforces the allowlist and logs every request, an immutable audit log capturing both denials and permits, and no credentials present inside the execution environment beyond what the specific task actually requires. Its Team tier includes access to a static egress IP, which makes it possible to allowlist Replicas' outbound traffic on the destination side, against a team's internal APIs or its registry, without resorting to wildcarding the internet to cover every possible case.
The network boundary only holds when both layers are right at once. The policy layer (egress filtering, allowlists, and proxy) has to be correctly scoped, and the isolation layer (per-run VMs with no shared state) has to actually contain whatever gets through. Getting one of these right while leaving the other weak simply moves the single point of failure from one layer to the other.
Sources
- Methods and apparatus for control and detection of malicious content using a sandbox environment
- Methods and apparatus for control and detection of malicious content using a sandbox environment
- Methods and apparatus for control and detection of malicious content using a sandbox environment
- Methods and apparatus for control and detection of malicious content using a sandbox environment


