A snapshot in this system is a checkpoint: an immutable capture of a
sandbox’s state that new sandboxes branch from. Taking one is
POST sandboxes/{name}/snapshot (or e2b’s POST /sandboxes/{id}/snapshots);
branching one produces a fresh sandbox with its own id and lease.
This document is about the part that is not obvious: a checkpoint is created on one node’s local disk, and the node that eventually has to serve a branch of it may not be that node.
Cloning a sandbox is fast — 28–75 ms — because nothing is copied:
| step | mechanism |
|---|---|
| guest memory | os.Link hardlink of the snapshot’s memory-range-* files |
| writable disks | reflink (FICLONE), concurrent, NoSync |
| restore | mmap copy-on-write — the guest maps the file, pages privatize on write |
All three require the checkpoint and the new VM’s run directory to be on the
same local filesystem, and that filesystem to support reflink (XFS/btrfs).
Hardlinks do not cross filesystems. FICLONE does not work on NFS, Lustre, or
any FUSE object-storage mount. mmap over a network filesystem turns a page
fault into a network round trip.
So the fast path is not an optimization layered on top of the storage choice — it is the storage choice. Any design that puts checkpoints on shared network storage has already given up the number that makes the system worth using.
That is the whole reason the design below prefers moving the placement to the data over moving the data to the placement.
A checkpoint stays on the node that took it. A branch of it issued to that node reads it from the local store and clones on the local reflink path.
Nothing in this tier is new; it is what a single-node deployment already does. It is listed because it is the case that must stay fast, and the two tiers below exist only to avoid degrading it.
Nodes gossip the checkpoint ids they hold, alongside the warm-pool counts and
promoted-template hashes they already gossip. A node asked to branch a
checkpoint it does not hold answers with a redirect to a node that does —
the same 200 + redirect: [addrs] contract a warm-miss claim already uses,
and the client retries there with no_redirect: true.
The record does not move. The clone still happens on a node whose disk already holds the data, on its local fast path. Cross-node correctness is bought with one extra round trip and zero bytes transferred.
This is the tier that makes “snapshot on node A, branch from anywhere” work.
Redirect cannot answer when no owner can serve: the owning node is gone, draining, or out of capacity. The choice then is between paying one transfer and failing the request.
With checkpoint_peer_heal enabled, the node pulls the record from an owner
over sandboxd’s own peer transport, publishes it into its local store through
the same atomic staging path a locally-created checkpoint takes, and then
serves the branch locally.
The cost is paid once: afterwards this node is itself an owner, gossips the record, and serves later branches of it at L1 speed.
It is off by default. Trading a transfer for availability is an operator’s decision, not a default.
The peer transport is deliberately not cocoon’s image/snapshot mover. cocoon’s transfer is bound to its VM store layout and its own addressing; reusing it would tie a checkpoint’s mobility to the engine’s release cycle and to a layout that has nothing to do with the record format being moved.
sandboxd’s transport moves exactly one thing — a store record: an export/
directory plus its meta.json — between two sandboxd nodes that already
authenticate to each other with the fleet api_token, over the control-plane
port they already share. It is a tar stream (GET /v1/checkpoints/{id}/blob),
so a pull is one round trip and never materializes the record twice. Entry
names are validated against path traversal and the stream is size-capped: a
peer is authenticated, but an authenticated peer is still not allowed to write
outside the destination this node chose.
Two independent reasons, either of which is sufficient:
It destroys the fast path. JuiceFS does not support reflink. Putting checkpoints on it turns every clone from “hardlink + reflink, 28 ms, zero bytes moved” into a full read over the network. The system’s headline number is the thing being traded away.
It does not solve the stated problem. JuiceFS is a metadata engine plus
an object store. With no S3 available, running JuiceFS means first running
MinIO/SeaweedFS — at which point object storage exists, and the
already-implemented s3 store backend is strictly simpler than adding a
POSIX layer on top of it.
The same reasoning rules out NFS and Lustre mounts for the checkpoint directory.
If a shared backend is ever required, use the s3 backend (which keeps a local
cache generation and therefore preserves the local-clone step) rather than a
network POSIX mount.
A checkpoint lives on one node. If that node’s disk is lost, the checkpoint is lost. L3 heals a node that cannot reach a record; it does not replicate one.
This is a deliberate, scoped decision, not an oversight — checkpoints here are branch points for agent workloads, not backups. Making them survive node loss needs asynchronous replication to N peers and a placement policy that tracks replica sets; that work is recorded in ROADMAP.md and is explicitly not built.
Operationally: treat a checkpoint as durable only while its node is. If a
checkpoint matters beyond that, promote it to a template
(POST /v1/sandboxes/{id}/promote) and distribute it as one, or export it out
of the fleet.
| setting | default | effect |
|---|---|---|
checkpoint_dir |
<data_dir>/checkpoints |
where records live. Keep it on the same local reflink filesystem as the VM run directory. |
checkpoint_store.kind |
dir |
s3 switches to object storage with a local cache generation. |
checkpoint_peer_heal |
false |
enables L3. Requires a mesh; ignored without one. |
L1 and L2 need no configuration: gossip and redirect are active whenever a mesh is configured.