cocoon sandbox

Deploying sandboxd

sandboxd is a single static binary per node. It drives VM lifecycle through the cocoon CLI and needs a template image with silkd baked in.

Prerequisites

Prebuilt static linux/amd64 and linux/arm64 binaries (sandboxd, sandboxd.dbg, sandbox-mcp, silkd, with checksums.txt) ship with every GitHub release; the boot artifact and the base/rt/python/python-rt/node/node-rt images are multi-arch manifests (browser, android, e2b-rt and e2b-ci remain amd64-only). Build from source with make sandboxd (produces dist/sandboxd); either way sandboxd -version reports what you are running.

Upgrading

This release does not convert VM or snapshot state from older releases. Drain old claims and use fresh data_dir and checkpoint_dir locations when upgrading; older checkpoints and promoted templates must not be reused.

The scalar egress-attachment keys are retired: rename "bridge": "br0" to "bridges": ["br0"] and "network": "cni" to "networks": ["cni"] before starting the new binary — config loading rejects the old spellings loudly rather than silently dropping the egress lane.

Dataset volumes require a lockstep rollout. Upgrade every sandboxd node and cocoon to the required version before enabling the catalog or shipping an SDK that requests volumes. Mixed-version serving is unsupported. Once a volume claim has finalized, do not roll a node back to an older sandboxd until all volume claims are gone; the older daemon cannot preserve the volume claims’ admission holds and marker bookkeeping recorded in claims.json.

The guest images and sandboxd move together on the egress lane: sandboxd tells a NIC-bearing guest how its traffic leaves through /etc/silkd-lane, and a guest image from another release reads a different file or none, which on a locked bridge lane leaves its execs routing into the lock. Roll the images with the daemon. A config that uses the socks5 or ports policy keys does not load on an older sandboxd (unknown keys are rejected), so a rollback takes the config with it. The SOCKS5 door is opt-in and the pool policy owns it: a pool that relied on the door binding for any bare host rule adds "socks5": true to its policy or loses the door silently, since its config still loads; a tenant policy’s flag counts only on a claim outside any configured pool, where it is the whole policy.

Configuration

sandboxd reads one JSON file (-config, default /etc/sandboxd/config.json):

{
  "listen": ":7777",
  "data_dir": "/var/lib/sandboxd",
  "cocoon_bin": "cocoon",
  "restore_mode": "mmap",
  "no_direct_io": true,
  "advertise_addr": "10.0.0.5:7777",
  "bridges": ["br0"],
  "volumes": [
    {"name": "imagenet", "path": "/srv/datasets/imagenet.img"},
    {"name": "weights-llama", "path": "/srv/datasets/llama.img", "directio": "on"},
    {"name": "acme-corpus", "path": "/srv/datasets/acme.img", "tenants": ["acme"]}
  ],
  "api_token": "…",
  "mesh": {
    "node_id": "node-a",
    "bind": "10.0.0.5:7946",
    "join": ["10.0.0.6:7946"],
    "cluster_key": "base64…"
  },
  "pools": [
    {"template": "base:24.04", "net": "none",   "size": "small", "warm": 4},
    {"template": "base:24.04", "net": "egress", "size": "small", "warm": 2}
  ]
}

As written the egress pool’s guests reach nothing: on a bridge lane the NIC is locked default-deny and the proxy door is bound only for a claim that has a policy. Give the pool an egress block before expecting traffic; a tenant’s claim of it (acme above) needs the tenant’s own block as well and gets the intersection of the two.

field default meaning
listen :7777 control- and data-plane HTTP listener
data_dir /var/lib/sandboxd golden snapshot exports, the claims journal, the usage/audit journals (usage.jsonl, audit.jsonl + .1 backups), and checkpoints/ by default
cocoon_bin cocoon cocoon CLI binary
restore_mode unset clone and wake-restore memory mode: copy, ondemand, or mmap; use mmap for dense pools
no_direct_io false use buffered writable disks for Cloud Hypervisor cold boots and clones; recommended for dense ephemeral pools to avoid direct-I/O CoW journal contention
sync_claims false fsync claims.json and its directory on every commit (claim, release, renew, hibernate, wake, fork, archive, reap), about 0.3 ms each on NVMe; without it a host power loss can leave an empty journal, which fails the next start’s reconcile until the file is removed, after which hibernate snapshots are reaped as orphans and archive checkpoints stay stranded. Off by default; turn it on when running claims must come back from a power loss (see vmm_restart)
cocoond_socket unset cocoon daemon’s read-only API socket (by default <cocoon run_dir>/cocoond.sock); set, the node detects VMM exits from its event stream and polls only as a 60 s backstop. See dead VMMs
vmm_restart cold what the node does with a running claim whose VMM exited outside a transition (a VMM crash, or a host restart that took it): cold cold-boots the claim’s own VM from its disk under the same sandbox id and token (guest memory, processes, and the instance-metadata document are lost), none keeps the claim listed as failed for the caller to release or wake. Egress-lane and volume claims, a VM record gone from cocoon, and a VMM that exits more than 3 times in 10 minutes are marked failed under either value. See dead VMMs
no_balloon false boot pool and template VMs without the virtio-balloon (cocoon otherwise returns 25% of guest memory to the host); clones inherit it from the golden. A guest that thrashes before deflate-on-OOM fires — a 16G build tier running a large typecheck — needs its whole memory
advertise_addr = listen internal HTTP host:port for peer traffic and preview forwarding; also used by clients when client_advertise is unset. Must be routable when listen is a wildcard; a node with mesh set refuses to load while it names an unspecified host
client_advertise unset client-facing HTTP(S) origin for this node, e.g. https://node-a.sandbox.example.com; published in owner, redirect, and peer responses. No path, query, fragment, or userinfo
bridges / networks unset egress-lane attachment: a list of host bridge devices, or a list of CNI conflist names. Mutually exclusive; with neither set the node serves only the no-network lane. A Linux bridge holds at most 1024 ports (kernel BR_MAX_PORTS), so an N-entry list raises the node’s egress ceiling to N×1024 — VMs spread over the list by a stable hash of the VM name, so size it with headroom (the spread is statistical, not exact). bridges keeps the raw TAP-on-bridge attachment (taps in the root netns, no per-VM network namespace or CNI plugin execution); networks runs the CNI chain per VM. Guarded egress on the egress lane (an egress-lane pool policy or any tenant policy) needs bridges and rejects a CNI network at load; none-lane pool policies ride the proxy on either
volumes unset node-local catalog of operator-managed dataset images: [ {"name":"imagenet","path":"/srv/datasets/imagenet.img","directio":"off","tenants":["acme"]}, {"name":"scratch-db","path":"/srv/datasets/scratch.img","writable":true} ]. Names match ^[a-z][a-z0-9_-]{0,19}$ and cannot start with cocoon-; paths are absolute; directio is on, off, or auto and defaults to off for both read-only and writable entries. tenants is an optional access list: empty means every authenticated scope, and a listed name the tenant set does not hold grants nothing and logs a warning at boot (the tenant API may add it later); root always has access. writable (default false) lets a claim request mode: "rw" on that entry — see Dataset volumes. The catalog is intentionally not part of the cluster digest
secrets unset credentials the egress proxy injects by name: [{"name": "gh", "header": "Authorization", "value_env": "GH_TOKEN"}]. A pool or egress-class rule references the name (on a claim where both apply, only the pool rule’s secret is injected); the value comes from the claim’s guest: false env, then the node’s environment, never this file. Without value_env the secret’s name is the env to read, and only claims supply it. See egress
egress_internal_allow unset CIDR prefixes re-admitted through the egress proxy’s SSRF guard, node-wide (every pool and tenant); prefix:port,port scopes an entry to those ports, a bare prefix admits every port. Prefer the port form. Prefixes, not a permit-private switch — the guest bridges are themselves ULA/RFC1918. See egress
egress_upstream unset per-claim upstream proxies: {"claim_env": "EGRESS_UPSTREAM", "allow": ["res.example.com", "10.1.0.0/16"]}. A claim’s guest: false entry under claim_env (http:// or socks5://, optional user:pass@, or direct) routes its egress through that upstream, which must be on allow (host names, IPs, CIDR prefixes). See upstream proxies
egress_usage_bytes false append an egress_bytes usage event with the payload bytes of every allowed egress tunnel or request when it ends; one extra usage-journal write per connection
egress_ca unset HTTPS-interception PKI: root_cert (the cluster root baked into intercepted guests; may bundle old+new roots during rotation) plus this node’s intermediate_cert/intermediate_key from sandboxd ca issue-intermediate. Required when any pool rule sets intercept
api_token unset the operator (root) credential: when set, guards the node-level endpoints (Bearer) with full access, including release-by-id cleanup. Per-sandbox tokens guard ordinary sandbox-scoped calls
meta_store unset {"kind": "pg", "dsn_env": "SANDBOXD_PG_DSN"}: keep the tenant set in one shared PostgreSQL instead of each node’s tenants.json (see shared tenant database). dsn_env names the node env holding the connection string. An optional cell also keeps the API-applied pool set there, one per cell, instead of in each node’s pools.json (see shared pool set); without it pool targets stay per node. Needs api_token. Restart only
max_fork_count 16 children a single fork may create; each is a full-RAM VM, so this bounds one request’s memory blast radius to the node’s capacity
refill_concurrency 0 (auto) concurrent VM provisioning budget, shared by warm-pool refills, fork clones, and the reap/hibernate/reconcile engine batches; the batches never take the last quarter of it (at least one slot), so a mass expiry or idle sweep cannot stop refills. 0 sizes it from the node: NumCPU*2/3 clamped to [4, 256] — a 384-core node gets 256; small nodes keep a floor of 4
release_delay_seconds 0 seconds a released VM waits in the removal queue before cocoon vm rm. The claim is dropped and journaled and its egress listener closed at release; the VM removal, the volume hold release and the egress tap unlock run on the first reap tick (5 s) after the delay, on the refill_concurrency budget like every other batch teardown, so a burst of releases does not compete with the claims still running. Until then the VM holds its memory and its volume reservations without counting as a claim. 0 removes inline
preview_listen (off) address for a preview HTTP server that serves guest ports under signed URLs; needs preview_secret
preview_secret — cluster-shared HMAC secret signing preview tokens (all nodes share one)
preview_advertise = preview_listen the browser-facing preview base URL, minted into every preview URL, so it must name a routable host (a wildcard preview_listen needs it set explicitly); nodes behind one TLS proxy may share it, while signed tokens route internally through each owner’s advertise_addr
checkpoint_dir <data_dir>/checkpoints where checkpoints and promoted templates live. Point it at a shared FUSE mount (JuiceFS over object storage, NFS) and every node sharing the mount can branch every checkpoint — records are generation-addressed with meta.json as the atomic commit pointer, so no cross-node locking is required of the filesystem. A re-publish retains its superseded export generation for at least ~1h and until a following hourly sweep so an in-flight clone that resolved the old metadata can finish; budget the current generation plus every generation retained across that grace-and-sweep window. An explicit delete can make a concurrent clone fail visibly. One contract on any shared root (mount or bucket): a template key has a single writer — promotes go to the sandbox’s owner node, and operators must not race promotes of one name from different nodes (checkpoint ids are node-generated and never collide)
checkpoint_store dir checkpoint AND promoted-template backend (both live in one store root, id-namespaced ck_/tp_): {"kind": "s3", "s3": {"bucket": "…", "prefix": "ck/", "endpoint": "…", "region": "…", "force_path_style": true}} stores checkpoints in object storage (any node claims any checkpoint, no shared mount needed). Credentials come from the standard AWS chain (env/IAM role), never this file. Re-publish retains prior export generations until Delete so an in-flight fetch that selected old metadata can finish; budget storage for those generations. When caching a new generation, each node prunes other cached generations of that record unused for at least an hour. An explicit S3 Delete can still make a concurrent fetch that has not finished materializing fail visibly. A crash between upload and the meta.json commit marker leaves orphan objects invisible to listings — add an S3 lifecycle rule to reclaim them. Absent = the dir backend at checkpoint_dir
checkpoint_ttl_hours 0 (keep forever) ages out checkpoints older than this; the sweep runs hourly and at startup. Explicit deletes never wait for it. Must be nonzero and match fleet-wide when checkpoint_peer_heal is on — it is the expiry eligibility point for a healed replica a delete broadcast missed, after which its next successful hourly sweep removes it; persistent sweep failure extends retention until one succeeds, so it is not a hard ceiling
checkpoint_peer_heal false on a cluster, lets a node pull a checkpoint it lacks from a peer — found via a live probe, not gossip — rather than failing the branch; see placement lifecycle. Three requirements, all enforced at config load: a nonempty api_token (the blob transfer between peers authenticates with it; without one the raw record stream would be open), mesh.cluster_key set (the pull presents the fleet api_token to an address learned from the peer probe, so the gossip layer carrying that address must itself be authenticated), and checkpoint_ttl_hours nonzero (a replica a delete broadcast missed becomes eligible for expiry after it, and its next successful hourly sweep removes it — so it is the finite eligibility point, not an exact ceiling). A shared checkpoint store (checkpoint_store kind s3) ignores this setting — every node already resolves every checkpoint directly, so there is nothing to heal
warm_max (pool entry) 0 (static) turns on the demand-adaptive watermark for that pool: the warm target rises from warm toward warm_max while claims arrive faster than the measured provision lead covers, and decays back over ~a minute of silence. The pool follows the target down: warm VMs above it are destroyed once the count has stayed above it for a full decay period, so a burst does not leave the pool parked at its high-water mark
warmup (pool entry) unset argv run in the golden VM after readiness and before its snapshot, so the files it touches are page-cache-resident in every clone, and again in every clone before it joins the warm pool, so those pages are already faulted into the restored VM when the first command runs — e.g. ["node", "-e", "0"] on a Node flavor. It runs under the engine’s 2-minute command timeout in silkd’s base environment (PATH, TERM, and the guest image’s proxy variables wherever nothing routes directly — the none lane and the locked bridge egress lane — with the proxy not yet serving, since a door pre-bound at refill only starts serving at claim); a non-zero exit or a timeout fails the golden build, so the pool stays unfilled until the config is fixed. Config-owned like egress: PUT /v1/pools rejects it, and a golden built with a different warmup is rebuilt
capture_trim (pool entry) false Trim the guest’s copy-on-write disk before a promote or checkpoint of this pool’s sandboxes (and of the clones of a template promoted from it), so the record carries only live data: blocks a build deleted — a COPY step’s archive, a RUN step’s caches — are no longer captured. sandboxd mounts the disk init resolved from cocoon.cow a second time at /run/sandboxd-trim through silkd as root, runs fstrim and unmounts, under the capture’s transition lock (about 15 ms on a warm clone). A failed trim is logged and the capture proceeds, larger. A hibernated or archived sandbox is captured untrimmed. Config-owned like egress: PUT /v1/pools rejects it
egress_classes unset named tenant egress layers: [{"name": "desk", "egress": {…}, "egress_upstream_env": "DESK_UPSTREAM"}]. A tenant names one with egress_class through the tenant API and its claims take that policy intersected with the pool’s. egress is required; intercept stays a pool-rule flag. Reloadable: a class’s policy reaches live claims on their next request. A class removed while tenants still name it leaves them with no egress and is logged at load and at reload
egress_upstream_env (pool or egress class entry) unset names a node environment variable holding the default upstream proxy URL for the pool’s or class’s claims (class before pool, a claim’s own entry before both); needs egress_upstream, and the URL must be on its allow. Read and checked at startup. Config-owned like egress: PUT /v1/pools rejects it
storage (pool entry) unset (cocoon’s default, 10G) size of the copy-on-write disk the pool’s golden is cold-booted with, in cocoon’s own --storage spelling ("40G", "40GiB", "40Gi": Docker/Kubernetes units, all binary) and passed to it verbatim; every clone, fork, checkpoint branch and promoted template of the pool inherits it. At least cocoon’s default of 10G. Equal sizes spelled differently build the same golden; a changed size rebuilds the golden on the next build, and claims already out keep their disk. Config-owned like egress: PUT /v1/pools rejects it
max_claims 0 (unlimited) node-wide cap on live claims; claim/fork/branch requests beyond it answer 429 with the pool state unharmed (on a cluster a non-volume claim tries a warm-peer redirect first; a volume claim answers the 429 with no redirect)
audit_log false append every relayed frame that opens an RPC — its op + addressing fields, never payloads — to <data_dir>/audit.jsonl; continuation data, data_end, stdin, and stdin_close frames pass through unrecorded. The file is size-rotated with one .1 backup. Records are {t, id, op} plus whichever addressing fields the op carries (argv, path, dest, from, to, url, session, port), plus method (GET, CONNECT, SOCKS5, …), decision, secret (the ref name, never its value) and upstream (the upstream proxy’s host:port, when one carried it) on egress records; preview accesses record as op preview, one per request; a tenant change records as op tenants with the added, changed and removed names, never a token. An opening frame whose first line exceeds 4 KiB records as op oversized with no addressing fields
idle_hibernate_seconds 0 (off) node-wide idle policy for unpooled claims (template/checkpoint claims): a none-lane claim is hibernated once it has had no open data-plane connection (relay, buffered exec, preview) and no egress request in flight for this long; the clock restarts when the last connection closes, so a long command is never cut short. The next call that reaches the guest wakes it transparently. Per-pool idle_hibernate_seconds does the same for that pool’s claims; pooled keys ignore the node-wide value, and egress pools reject it because they cannot resume safely. Opt in deliberately: a wake costs latency and the snapshot, so callers with their own idle logic must not pay twice
archive_after_seconds 0 (off) tier below hibernation: a hibernated claim idle this long is checkpointed to the store and its local VM dropped, freeing the node entirely; the next call that reaches the guest restores it transparently (a checkpoint restore’s latency) with a fresh lease of the length the claim was granted (the server default of 5m when it asked for none); a lease moves only when a client calls /renew. Keep idle_hibernate_seconds under the leases your claims use, or a woken sandbox that idles again is destroyed by the reap before it can hibernate, its archive already consumed by the wake; one still in use when its lease ends is destroyed the same way, unless it claimed on_expire: archive, which archives at the lease end whatever this setting. Requires idle_hibernate_seconds > 0 and must exceed it. Node-wide for unpooled keys; per-pool overrides for that pool
archive_delete_after_seconds 0 (keep) purge an archived claim’s store checkpoint this long after it was archived, reclaiming storage; the claim is then gone for good. On archive this retention window replaces the live claim deadline; 0 clears the deadline so the archive is kept forever. Same node-wide/per-pool split
mesh unset join a cluster (Clusters); unset = single node
pools[] — warm pools, keyed by (template, net, size). warm defaults to 4 only when both warm and warm_max are omitted; an explicit warm: 0 is kept, so warm: 0 under a warm_max starts the pool empty and grows it on demand, and warm: 0 alone registers the key with no warm VMs (PUT /v1/pools never defaults warm); net is none or egress; size is a tier, below. Retune online without a restart via PUT /v1/pools — omitted pools drain. This is the first-boot seed: once a node takes a PUT /v1/pools, the applied set persists to <data_dir>/pools.json and overrides this section on every later boot (a startup log notes it); delete pools.json to return to config-owned pools. With a meta_store cell, the cell’s set overrides it instead (see shared pool set). Egress stays config-owned either way. A template re-pulled or re-imported under the same ref is picked up within a minute: the pool rebuilds its golden and replaces its warm VMs. See state ownership

Size tiers (free-form CPU/memory is deliberately not accepted — it would fragment the warm pools):

size CPU memory
small 1 512M
medium 2 1G
large 4 4G
xlarge 4 8G
2xlarge 8 16G

Dataset volumes

Each catalog path must name a disk image containing a mountable whole-device filesystem. A read-only entry’s image must stay immutable: do not replace, truncate, or delete it while attached; publish a new catalog name or path instead. A missing path produces a startup warning and fails only claims that request it, allowing images to be distributed after sandboxd starts.

For example, build a whole-device ext4 image directly from a prepared tree, then make the published file host-read-only:

truncate -s 200G /srv/datasets/imagenet.img
mkfs.ext4 -F -d /srv/datasets/imagenet-root /srv/datasets/imagenet.img
chmod 0444 /srv/datasets/imagenet.img

Use the same dataset identity and access list for a volume name on every node, then distribute its immutable image to each node that should advertise it. sandboxd gossips only catalog names; it never copies content or gossips paths or access lists.

A claim requests up to eight unique names and may set an absolute, clean custom mount for each; the default is /volumes/<name>. Mounts must stay outside the guest OS directories and cannot duplicate or nest within one claim. Volume mounts may shadow an existing populated guest directory for that claim’s life.

A volume claim may consume an ordinary warm Cloud Hypervisor VM. sandboxd attaches after the warm pop or provision, waits for the device settle — a serial match under /sys/block, then the /dev/<name> node itself, since the kernel publishes sysfs before devtmpfs creates it — up to 2 seconds, then mounts the device before finalizing the claim: mode: "ro" (the default) attaches and mounts read-only; mode: "rw" requires the catalog entry’s writable: true and attaches and mounts read-write. Setup failure destroys the VM without quiescing — the claim was never handed out, so no workload write happened — and a popped warm VM is refilled normally.

A multi-volume claim brings every volume up concurrently: each volume’s own marker→attach→mount order is preserved, but cocoon serializes the hypervisor attach per VM, so what actually overlaps across volumes is the CLI spawns, device settles, and guest mounts. A partial failure still fails the whole claim and destroys the VM without quiescing, so any rw volume already attached keeps its dirty marker — including one the caller never got a handle to; recovery is the normal rw cycle (see Writable dataset volumes).

Warm candidates retain their normal ranking, but a volume claim’s candidate set, promoted-template intersection, and redirect follow the fleet-wide rule in cluster.

An empty catalog tenants list allows every authenticated scope; a nonempty list limits the image to those tenants, while root always bypasses it. A listed tenant the set no longer holds grants nothing and logs a warning at boot. GET /v1/volumes reports the caller-visible fleet union and holder count, plus the answering node’s current local availability and whether the entry is writable, without exposing host paths or node addresses. Applied names, effective mounts, and (for rw entries) mode are persisted with the claim. Such a claim cannot hibernate, fork, checkpoint, or promote; the idle hibernate sweep leaves it running. Release removes the VM but never deletes the operator-owned backing image — an rw release quiesces (unmounts) first, below. With directio=off, readers share the host page cache; use directio=on when cache interference matters. A dataset mounted into an egress-lane sandbox can be uploaded wherever that tenant’s egress policy permits, so treat the ACL and egress policy as one access decision. Image replication, detach, and refcounting remain out of scope.

Writable dataset volumes

A catalog entry with writable: true may be claimed with mode: "rw" (omitted, or "ro", always mounts read-only, even against a writable entry). directio still defaults to off for a writable entry — the durability nuance is that the host page cache sits between the guest’s flush and the device, so directio=on is the knob when that gap matters, not just cache interference.

Single holder, by operator contract. A writable name must be attached from exactly one node — the same weight as “replacing an attached image is operator error.” Nodes advertise only the catalog paths they actually hold, so every claim for that name, ro or rw, already funnels to that one node; running the same writable name from two nodes at once is dataset divergence that nothing in the protocol detects for you.

The <path>.dirty marker. Before the first rw attach of an image, sandboxd durably creates a <path>.dirty sidecar beside it; a clean rw release removes it. It is a write-ahead record, not a lock: it survives a killed sandboxd or a dead VMM, so anything short of a clean unmount leaves it in place. It is operator-visible — ls next to the image shows whether it’s mid-write — and travels with the image on shared storage. A leftover marker is not corruption (a journaling filesystem already makes a hard VM kill crash-consistent); it means the image needs one rw claim, which replays the journal as a side effect of mounting, followed by a clean release, before it can serve ro again. A ro claim against a dirty image is refused rather than auto-healed: replaying the journal takes the write lock, which conflicts with concurrent readers.

Concurrency, for one image name:

Attach-only claims opt out of the marker, not out of admission. A claim sent with volumes_attach_only gets the device attached and nothing else, so sandboxd neither writes nor clears <path>.dirty for it — it cannot verify a mount it did not perform. Everything in this section still applies to the default, mounting claims, and the exclusion rules above apply to attach-only claims exactly the same way, which is what keeps other tenants safe. The operator-visible difference: an image whose attach-only writer released without unmounting cleanly carries no marker, so the next ro claim is admitted and fails at mount time (500) instead of being refused early (409). Whoever hands out attach-only rw access owns that trade; see sandboxd-api.

A fuller config

The block above is the minimum. A production node with guarded egress + HTTPS interception, previews, an object-store checkpoint backend, the idle→hibernate→archive tiers, and a mesh looks like this — every field here validates on load:

{
  "listen": ":7777",
  "data_dir": "/var/lib/sandboxd",
  "advertise_addr": "10.0.0.5:7777",
  "bridges": ["br0"],
  "restore_mode": "mmap",
  "no_direct_io": true,

  "api_token": "op-root-token",

  "secrets": [
    {"name": "gh", "header": "Authorization", "value_env": "GH_TOKEN"}
  ],
  "egress_ca": {
    "root_cert": "/etc/sandboxd/egress-ca/root.crt",
    "intermediate_cert": "/etc/sandboxd/egress-ca/node-a.crt",
    "intermediate_key": "/etc/sandboxd/egress-ca/node-a.key"
  },

  "preview_listen": ":8443",
  "preview_secret": "cluster-shared-hmac-secret",
  "preview_advertise": "https://preview.example.com",

  "checkpoint_store": {
    "kind": "s3",
    "s3": {"bucket": "sandbox-ckpt", "prefix": "ck/", "region": "us-east-1"}
  },
  "checkpoint_ttl_hours": 168,

  "max_claims": 200,
  "audit_log": true,
  "idle_hibernate_seconds": 120,
  "archive_after_seconds": 3600,
  "archive_delete_after_seconds": 604800,

  "mesh": {
    "node_id": "node-a",
    "bind": "10.0.0.5:7946",
    "join": ["10.0.0.6:7946"],
    "cluster_key": "MDEyMzQ1Njc4OWFiY2RlZg=="
  },

  "pools": [
    {"template": "rt:24.04", "net": "none", "size": "small", "warm": 4, "warm_max": 12},
    {"template": "rt:24.04", "net": "egress", "size": "medium", "warm": 2,
     "egress": {"socks5": true, "allow": [
       {"host": "api.github.com", "methods": ["GET", "POST"], "secret": "gh", "intercept": true},
       {"host": "*.googleapis.com"},
       {"host": "imap.example.com", "ports": [993]}
     ]}}
  ]
}

Secret values come from the environment named by value_env (here GH_TOKEN), never the file. The egress_ca files are provisioned with sandboxd ca — see Guarded egress. On a cluster, api_token, preview_secret, and cluster_key must match on every node, and so must the tenant set.

High-density no-network pools

Both lanes use Cloud Hypervisor; net: "none" launches it with no NIC. For dense ephemeral pools, set restore_mode to mmap, enable no_direct_io, place Cocoon’s run_dir on a capacity-sized tmpfs, and set Cocoon’s cni_conf_dir empty on none-only nodes to avoid per-VM network namespaces. The last setting also removes the VMM’s network-namespace quarantine. On a 384-core host, start with refill_concurrency: 64; higher values need measurement. Pause cocoon daemon or lengthen its reconcile interval during a mass fill.

Host tuning for dense egress-lane nodes

The no-network lane is CPU-bound and needs none of this. The egress lane serializes on the kernel’s global rtnl lock — tap create, bridge attach, and the VMM’s own tap open all take it — so at hundreds of VMs host configuration dominates fill throughput:

A caching resolver for egress

The egress proxy resolves the destination of every new upstream connection (a CONNECT tunnel, a direct dial, and the proxy host on a routed upstream) with Go’s resolver, which reads /etc/resolv.conf and caches nothing. Against a network resolver one lookup measured 3.2–3.6 ms, about a fifth of a new HTTPS tunnel’s setup. Run a caching resolver on the node — the systemd-resolved stub, dnsmasq or unbound; node-local-dns on Kubernetes — and point /etc/resolv.conf at it. It honors record TTLs, and the internal-address guard still checks every address it returns.

Reserving CPU for the control plane

A saturated node can leave clone/wake execution and sandboxd itself nothing to run on — under full-node load CH clone p95 has measured in the tens of seconds. cocoon newer than v0.5.8 runs every VMM in its own cgroup v2 CPU scope (Guaranteed-at-N: a VM’s CPU count is also its hard host-CPU cap) and adds cgroup_cpus, a one-time cpuset fence keeping the whole VM population — including the virtio and io-wq worker threads vCPU affinity cannot reach — off reserved host cores. On dense nodes set the fence in cocoon’s config (e.g. cgroup_cpus: "0-379" on a 384-core host) so the reserved cores stay free for sandboxd, its cocoon invocations, and the OS; cocoon’s CPU-isolation docs cover validation semantics and the per-VM knobs.

Auth model

Three token kinds. The root api_token has full access — operators and single-tenant deployments need nothing else. Tenant tokens, added through the tenant API and kept in <data_dir>/tenants.json (or the shared tenant database), create and manage their own resources: claims, forks, checkpoints, promoted templates, and preview URLs are stamped with the tenant name; sandbox and checkpoint listings filter to the caller’s tenant, and a tenant can delete only its own checkpoints and templates (root sees everything). Operator surfaces stay root-only — per-id sandbox reads and stats, GET /v1/checkpoints/{id}/blob, GET /v1/info, PUT /v1/pools, /v1/tenants, /v1/config/reload, POST/DELETE /v1/drain, /metrics; a tenant token there is authenticated but not authorized, so it answers 403 (a wrong token stays 401). Per-sandbox tokens are unchanged: whoever holds a sandbox’s token drives that sandbox. Fork children inherit the parent’s tenant and count against its max_claims (0 = unlimited, a per-node cap, so a tenant’s cluster limit is max_claims × nodes); a tenant at its cap gets 429 exactly like a node at max_claims, and the usage journal’s claim events carry the tenant for per-tenant billing. config.json has no tenants field: a node starts with no tenants until the API adds them.

Running

sandboxd -config /etc/sandboxd/config.json

On start the node reconciles: persisted claims whose VMs cocoon still holds are re-adopted, everything else sbx--prefixed is removed. A record still in cocoon’s creating state is reclaimed through vm reconcile-stale-create (cocoon ≥ v0.5.8), which refuses while a clone is in flight instead of deleting the VM out from under it; older cocoons fall back to forced removal. Then the refill loop builds one golden snapshot per pool (a one-time cold boot + snapshot export, tens of seconds) and keeps each pool topped up with claim-ready clones. GET /v1/info shows "golden": true and warm at target when the node is ready to serve warm claims.

A minimal systemd unit (shipped as packaging/sandboxd.service):

[Unit]
Description=sandboxd
Wants=network-online.target
After=network-online.target

[Service]
ExecStart=/usr/local/bin/sandboxd -config /etc/sandboxd/config.json
Restart=on-failure
KillMode=process
Environment=SANDBOXD_LOG_LEVEL=info

[Install]
WantedBy=multi-user.target

Each record carries its syslog level, so alerting keys on systemd priority: journalctl -u sandboxd -p err lists the errors and -p warning the degradations. The one line an unparsable SANDBOXD_LOG_LEVEL prints before the logger is up carries no level, so read it with journalctl -u sandboxd or systemctl status sandboxd. With stderr off a terminal, as under this unit, the record itself is one JSON object; a terminal gets the console format. journalctl answers an empty range with -- No entries --, so a check that counts lines reads one problem where there are none — count records (-o json | wc -l) or test the output for emptiness.

Stopping sandboxd leaves VMs alive; the next start reconciles them. The unit’s KillMode=process is what guarantees it: systemd signals sandboxd alone, so a VMM that cocoon could not move into its own scope, and any cocoon call still in flight, outlive the stop instead of dying with the service’s cgroup. Claimed sandboxes are reaped when their TTL expires (default 5m, capped at 24h).

Dead VMMs and host restarts

With cocoond_socket set, the node follows cocoon daemon’s event stream (GET /v1/events on that socket): the daemon watches every VMM through a pidfd, so an exit reaches sandboxd within a daemon pass and a recycled PID never reads as a live VMM. The node then lists cocoon’s VMs only every 60 s as a backstop for a dropped event, and every 10 s while the stream is down (it reconnects with backoff) or when no socket is configured. Each list runs only while the node holds a running claim; each suspect is then checked with a single-VM cocoon vm inspect under its lock. A claim whose VMM is gone — killed, crashed, or lost with the host — is handled per vmm_restart. A VMM that exits because the guest powered off is treated the same way. A claim whose lease has already lapsed is left to the reap, unless it claimed on_expire: archive, which needs a running VM to archive:

outcome when client sees
cold restart cold (default), none-lane claim without volumes, within 3 restarts in 10 min restarts +1 and restarted_at on the sandbox row, vmm_exit then vmm_restart usage events; a call before the next pass gets 502, one during the boot waits for it, open relays see EOF and redial
failed none, an egress-lane or volume claim, the VM record gone, the restart budget spent, or 3 failed boots failed: "<reason>" on the row, vmm_exit then vmm_failed usage events, 409 from exec, relays, ports, hibernate, fork, checkpoint and promote

A cold restart boots the claim’s own copy-on-write disk: files the guest synced survive, memory and processes do not. The claim’s guest env entries are written again and its egress proxy is re-armed; the instance-metadata document is not kept, so re-send it when restarts moves. A failed claim keeps its VM and disk until release or lease end (it is destroyed at expiry even with on_expire: archive); POST /v1/sandboxes/{id}/wake cold-boots it and resets the restart budget. After a host restart the same rule applies to every running claim whose VM record survived, so hibernated, archived and running claims all come back; with sync_claims off a power loss can still lose the journal itself. Cocoon keeps each VM’s copy-on-write disk under its run_dir, so a run_dir on tmpfs (high-density pools) loses every local disk with the host: running claims then end failed, and hibernated ones cannot wake. Cold restart is verified on the direct-boot (OCI) templates this project ships; a cloudimg (UEFI) clone boots again without its per-clone cidata, and no shipped pool uses one. Without the daemon stream the poll reads cocoon vm list, whose liveness check needs a cocoon that verifies the VMM’s PID (cocoonstack/cocoon#274) to never read a recycled PID as a live VMM.

Maintenance

To empty a node for maintenance, cordon it first: POST /v1/drain (root) stops new claims and drains the warm pools; poll GET /v1/info until claimed reaches zero (or let the leases expire), then stop sandboxd. DELETE /v1/drain uncordons. The drain is not persisted — a restarted node serves again.

Shared tenant database

With meta_store set, every node reads tenants from one PostgreSQL table instead of its own tenants.json. Tenants are added, changed and removed through /v1/tenants on any node.

Shared pool set

With meta_store.cell set, the pool set applied through PUT /v1/pools belongs to the node’s cell rather than to the node: a PUT on any node of the cell replaces the warm targets of every node in it. Each pool’s warm, warm_max and other targets apply per node, as they do with pools.json, so every node of a cell runs the same targets. Nodes that differ in capacity or attachment go in different cells. A caller that gives nodes different targets — Client.SetPoolsCluster with per-node sets, or a controller that spreads a replica count across nodes — needs nodes without a cell: in one cell its per-node writes overwrite each other. Without cell, meta_store shares only the tenant set.

Reloading the config

kill -HUP <sandboxd pid> or a root POST /v1/config/reload (API) re-reads config.json, validates it with the same code as boot, and applies the reloadable settings all at once or not at all. Every field falls into one of three classes:

Effects on what already runs:

A successful reload writes an op:"config_reload" audit record listing what changed. Each node reloads its own file; on a cluster, distribute the file and reload every node.

Verifying a node

curl -s -H "Authorization: Bearer $TOKEN" http://127.0.0.1:7777/v1/info | jq .

The repository’s scripts/sandboxd-e2e.sh runs the full loop on a real node (golden build → warm pool → claim tiers → the complete verb smoke → reap → restart reconcile); set BRIDGE=<dev> to include the egress lane.

To include the read-only volume proof, put a nonempty volume-e2e.txt in the filesystem image and run VOLUME_IMAGE=/srv/datasets/imagenet.img scripts/sandboxd-e2e.sh. The script verifies two concurrent mounts, read-only enforcement, warm-pool consumption, and an unchanged source checksum. On a node using prebuilt binaries, also set VOLUME_SMOKE_BIN beside SANDBOXD_BIN, DEMO_BIN, and SMOKE_BIN.

Add VOLUME_RW_IMAGE=/srv/datasets/scratch.img (a second, writable image) to also run the writable leg: a durable write across release, second-writer exclusion, and a clean read-only claim afterward.

TLS for SDK clients

sandboxd serves plain HTTP behind a TLS-terminating proxy. Configure a stable client origin on every node reachable through that proxy:

{
  "listen": ":7777",
  "advertise_addr": "node-a.internal:7777",
  "client_advertise": "https://node-a.sandbox.example.com"
}

The mesh gossips both addresses. Client-facing owner replies, claim/template/ checkpoint redirects, and peer discovery use client_advertise. Checkpoint probing, healing, deletion broadcasts, and preview forwarding continue to use advertise_addr over internal HTTP. When client_advertise is unset, direct HTTP deployments keep their existing address behavior. Configure all cluster members before using external clients; a member without a client origin still advertises its internal address to clients. Internal SDK users must also be able to reach the configured client origins.

One proxy can serve the whole cluster, but each owner origin must route to one particular node. An entry load balancer may choose any node for the initial claim; a shared random-balancing owner origin cannot route later agent and release requests to the owning node. Unlike Preview URLs, the SDK API does not forward arbitrary sandbox requests between nodes.

Caddy reference configuration (node B uses the corresponding hostname and internal upstream):

node-a.sandbox.example.com {
    reverse_proxy node-a.internal:7777 {
        transport http {
            versions 1.1
        }
    }
}

node-b.sandbox.example.com {
    reverse_proxy node-b.internal:7777 {
        transport http {
            versions 1.1
        }
    }
}

For a private development CA, add tls internal to each site and provide Caddy’s root certificate to both SDKs using the TLS client settings. For public DNS names, Caddy can manage the certificates. Keep the upstream listeners and mesh private.

The edge must pass HTTP/1.1 Connection: Upgrade, Upgrade: silkd, and the 101 response, then relay both byte streams without response buffering. A WebSocket-only upgrade allowlist is insufficient. Agent connections must not negotiate HTTP/2. Configure stream/idle timeouts to exceed the longest relay; Caddy’s default stream timeout is unlimited. Configuration reloads may close active streams; set stream_close_delay when reloads need a drain window. The Upgrade tunnel does not need flush_interval -1.

The pinned Caddy integration runs both SDKs against two real sandboxd HTTP handlers and relays (with a fake VM/guest), with an unreachable internal owner address. It checks redirects, lookup, exec, port forwarding, half-close, release, and certificate rejection:

cd e2e
CADDY_BIN=/path/to/caddy GOWORK=off go test -race -run TestCaddyTLSCluster -v .

For hardware acceptance, run the same SDK sequence from outside the node network, including a guest HTTP server through proxy_port, and confirm an A-to-B claim never dials B’s internal address. A client-side TLS bridge alone does not translate returned owners or redirects; use it only when every returned endpoint is deliberately mapped through a local bridge.

Preview URLs

preview_listen starts a second HTTP server that serves a sandbox’s guest HTTP port under a signed, expiring shareable URL. The whole mechanism is in sandboxd: