Dependency Preloading Strategies for Agent VMs
Preloading dependencies before agent tasks start dramatically cuts cold-start time overhead.

A batch of agent tasks goes out, and for the first ninety seconds, sometimes longer, nothing resembling work happens. Pip resolves dependencies, npm walks a lockfile, apt pulls packages from a mirror, and the agent sits idle because the environment underneath it has not yet installed what the task requires. That dead interval, repeated on every fresh boot, is the dominant cost in an agent VM pipeline, not the speed of the model doing the reasoning once it finally starts. The install sequence that runs before inference can begin is the bottleneck, and it scales the wrong way: a single agent session can absorb a two-minute setup without anyone noticing, but a team running dozens of parallel tasks multiplies that same two minutes across every task, turning a tolerable delay into a structural drag on throughput.
The readiness ceiling of an agent VM is set by the environment, not by which model or harness is driving it. Point an identical harness at an identical task: the run on a well-preloaded VM starts producing useful work far sooner than the same run on a bare one, because the model has nothing to do until the filesystem, the dependencies, and the runtime are in place. The infrastructure layer, snapshot-based bootstrapping, dependency installation, secrets injection, terminal session management, is what separates a 30-second cold start from one that runs into the minutes. The model sits downstream of that constraint. It cannot compensate for an environment that isn't ready, and improving it does nothing to shorten the wait.
The isolation technology floor for boot performance
The isolation primitive sets a boundary that every preloading strategy operates inside, and the isolation spectrum runs from OS-level primitives, which add the least overhead but offer the weakest boundary, through container runtimes, to microVMs and full virtual machines, which add their own kernel and therefore their own boot cost. Each step along that spectrum trades raw boot speed for a stronger guarantee that one tenant's workload can't touch another's.
OS-level primitives sit at the fast end of that trade. Tools like macOS Seatbelt, Linux Landlock, bubblewrap, and seccomp-bpf constrain a normal host process rather than booting a separate kernel, so the overhead they add is minimal. They serve as the backend for several CLI-local sandboxes, where the isolation model fits a single developer running a single agent on their own machine.
Container runtimes, chiefly Docker and the broader OCI ecosystem along with Podman, occupy the middle of the spectrum. They form the de-facto base for many agent sandbox products because they start faster than a full VM while still giving each workload its own filesystem and process namespace. The trade-off is a shared kernel: containers on the same host are still tenants of the same underlying operating system, which is a weaker boundary than a microVM provides.
Firecracker microVMs sit further along that trade, booting their own kernel in exchange for a stronger isolation guarantee. Fly.io's Sprites product runs on Firecracker, and Vercel Sandbox, also built on Firecracker and generally available, adds filesystem snapshots as a platform feature on top of that isolation layer. Daytona takes a different route to a similar guarantee, using Sysbox containers by default to get VM-level isolation without the overhead of hardware virtualization, with Kata Containers available as an option when full microVM-level isolation is required.
The self-hosted frontier shows what's possible when a team configures a primitive by hand. A NixOS setup built on microvm.nix can create each microVM with source code repositories and build dependencies already present inside it, using Cloud Hypervisor (or QEMU) as the hypervisor, with each VM allotted an 8 GB disk overlay, 8 vCPUs, and 4 GB of RAM. The VM itself is defined through a Claude skill, a markdown instruction file that directs Claude Code to create the microVM, which makes the environment definition declarative and reproducible.
The isolation primitive sets the security and performance floor a team is working within, but it does not determine where they actually land on boot time. Preloading strategy is the variable that moves a deployment from that floor toward sub-30-second boots, and it applies regardless of which primitive is underneath.
Two configuration paths that determine what is already in the VM before an agent runs
A team has to decide how it will express what the VM is supposed to contain before any snapshot or layer-caching technique can help. That choice shapes how reproducible, portable, and fast every subsequent boot will be, and it's a genuine design decision with real trade-offs on both sides rather than a question with one obviously correct answer.
The first path is agent-driven setup. The team points the agent at a repository, supplies secrets, and lets the agent discover and run the install sequence on its own; once that sequence succeeds, the resulting VM state is saved as a snapshot for reuse on later boots. The advantage is that it requires almost no upfront environment authoring: the agent figures out what the project needs without a human writing a Dockerfile or a Nix expression first. The disadvantage is that the first run is always slow, since the agent is discovering the install sequence rather than replaying a known one, and the resulting snapshot can encode implicit state that's hard to audit or reproduce from scratch if the snapshot is ever lost or needs to be rebuilt.
The second path, declarative configuration through Nix, takes the opposite stance. A NixOS setup built on microvm.nix declares the entire VM as code in a flake, including which source repositories and build dependencies are pre-populated inside it, producing ephemeral VMs where nothing persists except what is explicitly shared with the host. Each VM is reproducible from its declaration alone, which removes the failure mode of a stateful VM that drifts over time as patches and manual fixes accumulate on top of an original build. The cost is that this path requires Nix expertise, a specialized skill set that most engineering teams haven't invested in and that raises the bar for anyone maintaining the configuration later.
Teams that want fast iteration with minimal upfront investment tend to lean toward agent-driven setup paired with snapshotting. Teams that need auditability and reproducible builds across a large volume of runs tend to lean toward the declarative path, or toward a Dockerfile-based build that sits between the two in terms of effort and guarantees.
A fourth design axis cuts across both paths and changes what "preloading" even means for a given deployment: the two-phase network model used by Codex cloud's runtime. Setup, including dependency installation, runs with the network on; the agent phase that follows runs with the network off by default, though that's configurable per environment. That ordering means all network-dependent preloading has to complete during the setup phase. If the agent discovers a missing package mid-task, it can't simply fetch it; the environment has to be reconfigured to allow that, which makes the setup phase the single point where a team's preloading decisions either hold up under the task or don't.
Snapshot-based bootstrapping and the elimination of the install sequence on every boot
Saving VM state after a dependency install succeeds, then restoring from that saved state on every later boot, is the most direct technique available for pulling the install sequence out of the critical path. It converts what was a variable, network-dependent operation, one that could take thirty seconds or five minutes depending on registry latency and package count, into a deterministic disk restore that takes roughly the same amount of time every time.
The sequence after a snapshot restore is bounded and predictable. The platform restores the snapshot, or builds from a Dockerfile if no snapshot exists yet, and then runs the install command again. That second step, running the install command on an already-restored disk, carries a constraint that's easy to miss and expensive to get wrong: the install command has to be idempotent, because it runs on every fresh boot and only disk state persists across the snapshot boundary, not running processes. An install step that creates a database schema on first run will fail the second time it runs against a disk where that schema already exists, because the step wasn't written to check for existing state before acting. Teams that write install steps assuming a clean system every time will find the snapshot either breaks outright or silently re-runs expensive work it didn't need to repeat.
Platforms building around Firecracker have started treating snapshotting as infrastructure rather than a hidden optimization buried in a managed service. Vercel Sandbox exposes filesystem snapshots as a first-class, generally available platform feature. Fly.io Sprites offers checkpoint and restore with persistent NVMe storage and scale-to-zero, giving teams a primitive to build their own bootstrapping logic on top of.
The obvious objection to relying on snapshots is reproducibility: a snapshot can encode implicit state that quietly diverges from the team's declared configuration over time, particularly when the snapshot was produced by an agent-driven setup rather than a controlled Dockerfile build. The discipline that answers this objection is to treat snapshots the way a team treats any other build artifact: version them, rebuild them on a schedule or whenever the dependency manifest changes, and validate the result against the declared configuration rather than trusting that an old snapshot still matches what the team thinks it contains.
Secrets injection runs alongside this restore sequence and deserves a brief mention without becoming the section's focus: credentials and environment-specific configuration need to reach the VM at boot without being baked into the snapshot itself. That is why secrets injection is treated as a separate step, kept apart from the install sequence.
Layer design: what to put in the image versus what to defer to runtime
Separating what changes rarely from what changes often is the design principle that governs image construction. System-level dependencies, compilers, language runtimes, native libraries, change on a timescale of months. Project code, lockfiles, and generated assets change on a timescale of minutes. Only the former belongs in the image layer that's meant to benefit from caching; put the latter there and the cache stops paying for itself.
The practical rule that falls out of this principle is a simple constraint: do not copy the full project into the image. The image layer's job is to capture system dependencies, nothing more, while the agent runtime handles checking out the actual workspace at the correct commit. That division keeps a single image reusable across every branch and every commit in a repository's history, instead of forcing a rebuild each time the code changes.
Layer ordering inside a Dockerfile is what determines the cache hit rate in practice. Dependencies that change rarely, OS packages, language runtimes, belong in the earliest layers, where they sit untouched across builds. Dependencies that change with the project, package lock files, build tool versions, belong in later layers that can be invalidated and rebuilt without discarding the stable base underneath them. The common mistake runs in the opposite direction: copying the entire repository early in the Dockerfile invalidates every layer that follows on every single code change. This erases the performance benefit layer caching was supposed to provide. The better pattern copies only the lock file or dependency manifest early, runs the install against that, and lets the agent runtime supply the actual source tree afterward.
The NixOS declarative approach takes this same separation to its logical endpoint. Because the VM declaration includes the source repositories and build dependencies explicitly, the environment is always derivable from the declaration itself, and there's no image layer left that can drift away from what's declared. For monorepos, the same separation principle extends to configuration: naming conventions for secrets, scoping by service name, keep service-specific configuration out of shared image layers while still making it available at runtime to the correct agent working on that service.
Warm pools as the operational pattern that decouples boot latency from task arrival
Even a well-designed snapshot leaves a gap between the moment a task arrives and the moment an agent is actually ready to work on it. Warm pools close that gap by doing the boot and setup work ahead of time, before any task is waiting for it; the boot sequence runs in advance, so the agent is already running when the task shows up.
The pattern itself is straightforward: a platform maintains a pool of pre-started sandboxes that have already gone through the full startup sequence, restored from snapshot, run their idempotent install command, launched a dev server and a test watcher inside a persistent terminal session. When a task arrives, it gets assigned to one of these already-ready environments.
What this buys a team is a separation between perceived latency and actual boot latency. The boot work still happens, and it still takes however long the snapshot restore and idempotent install take; it happens ahead of demand. The user or the system triggering the task experiences only the assignment step, handing a ready environment to a waiting task, rather than waiting through the restore-install-start sequence that produced it.

