What the stack defends, where the trust boundaries sit, what the operator must provide, and what it deliberately does not defend. Read this before exposing any part of a deployment beyond a single trusted host.
agent workload (untrusted code)
│ syscalls
guest kernel + virtio drivers
│
══ VM boundary (KVM) ═══════════════ the isolation unit
│ vsock (none lane) / nft-locked NIC (egress lane)
host: sandboxd + cocoon trusted
│ HTTP control/data plane, bearer tokens
clients: SDK / MCP / operators trusted per token scope
socks5, a SOCKS5 door to the same
allow-list that carries arbitrary TCP rather than HTTP only.
On the egress lane: the same vsock paths plus a NIC whose every
guest-initiated packet except IPv4 broadcast DHCP is dropped by an nftables lock in the host
root netns (egress); the lock is fail-closed and applied
before the claim is handed out. On both lanes the proxy refuses
loopback, private, link-local (cloud metadata), CGN, and the
IPv4-embedding IPv6 ranges, so an allow-listed name that resolves or
rebinds to an internal address cannot reach the host or a sibling.guest: false env, never the config
file; that env stays in the node-local claim record.These are hard constraints, not suggestions; the guarantees above assume all of them.
api_token, tenant tokens, and per-sandbox tokens
are cleartext bearers on the wire, and on a cluster the SDK dials every
node it is redirected to. Keep nodes and clients inside a VPC, WireGuard
mesh, or equivalent. The only surface designed to face a browser is the
preview listener, and it belongs behind a TLS-terminating proxy
(deploy). Never expose listen publicly.sandbox_egress_* nftables namespace; a second daemon would clear the
first’s locks.mesh.cluster_key unless the gossip network is itself trusted,
and open the memberlist port node-to-node only.Three credentials, in descending scope (API reference):
the root api_token (operator surfaces, full access), tenant tokens
(resource-creating verbs, everything stamped and quota’d per tenant), and
per-sandbox tokens (that sandbox only — holding a handle amplifies to
nothing node-level). Two capability tokens ride on top: preview URLs are
HMAC-signed, expire with the claim’s lease (an archived claim kept forever
has none, so its URL runs for the TTL it was minted with), and die with the
sandbox (no revocation list to leak); a checkpoint id is the unguessable capability to
branch it. On a cluster, deleting a checkpoint does not revoke that
capability fleet-wide the instant it runs: the delete is best-effort —
broadcast to every peer the node currently sees — so a peer that is offline
or partitioned at that moment keeps its own replica branchable until
checkpoint_ttl_hours ages it out (placement lifecycle).
Tenants are isolated at the API layer — listings filter, deletes answer 404
rather than confirming existence, and operator surfaces answer tenants 403.
The volume catalog is an operator-owned data boundary. A volume name and its
access list must mean the same thing fleet-wide, although membership is
node-local. An empty entry tenants list permits every authenticated scope; a
nonempty list permits only those named tenants, while the root token always
has access. Config load rejects an access-list name that is not a configured
tenant. Claim lookup returns byte-identical errors for an unknown and a
forbidden name, and catalog discovery filters before replying, so a tenant
cannot enumerate restricted entries by probing. Gossip carries names only.
Neither gossip, discovery, persisted claims, usage events, nor the sandbox index
exposes host image paths or access lists.
Read-only is integrity protection for the shared image, not confidentiality.
Mounting a dataset into an egress-lane sandbox gives that sandbox an export path
to every destination its tenant egress policy permits; review the volume access
list and egress policy together. With directio=off, concurrent readers share
the host page cache, which improves reuse but lets one tenant’s large scan evict
another’s cached pages — the same holds for a writer: a large rw write can
evict cached pages backing another volume’s readers exactly like a large scan
would. directio=on is the per-volume mitigation; v1 has no per-tenant cache
quota or accounting. Operators must keep a read-only entry’s attached image
immutable and publish a new name/path for new content.
A writable entry (writable: true) adds a channel the read-only model does
not have: whichever tenant is permitted to claim it rw changes what every
other permitted tenant reads next, and any of them can be the writer in
turn — a multi-tenant access list on a writable entry is a bidirectional
channel between those tenants, not just a shared read. Recommend a writable
entry’s tenants name exactly one tenant — an empty list permits every
authenticated scope, which for a writable entry means every tenant can write
to every other tenant’s next read. A dataset that genuinely needs multiple
writers needs an out-of-band process for who writes when, which the catalog
ACL does not provide.
Facts to plan around, stated so the boundary is honest:
intercept rules to hosts you control, or use
"inject" so only claims holding a credential for a host are intercepted
(egress).max_claims, per-tenant
caps answer 429). The API assumes callers inside the trust boundary;
front it with your own limiter if semi-trusted automation can reach it.