cocoon sandbox

Go SDK

import sandbox "github.com/cocoonstack/sandbox/sdk/go"

The SDK has no third-party dependencies (only the sibling protocol/wire module). One Client talks to one entry node; sandbox handles dial their owning node directly, so a client works unchanged against a single node or a cluster.

A complete example

Claim a sandbox, push a project into it, run a build, then freeze the built state and fan out two independent workers from that exact moment:

package main

import (
	"archive/tar"
	"bytes"
	"context"
	"errors"
	"fmt"
	"log"
	"os"
	"time"

	sandbox "github.com/cocoonstack/sandbox/sdk/go"
)

func main() {
	ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
	defer cancel()

	client, err := sandbox.Connect(os.Getenv("SANDBOXD_ADDR"),
		sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))
	if err != nil {
		log.Fatal(err)
	}

	sb, err := client.New(ctx, "rt:24.04",
		sandbox.WithSize(sandbox.Medium),
		sandbox.WithTimeout(10*time.Minute))
	if err != nil {
		log.Fatal(err)
	}
	defer sb.Close()
	fmt.Printf("claimed %s on %s\n", sb.ID, sb.Owner())

	// Push is the only ingestion path on the no-network lane.
	if err := sb.Push(ctx, "/work", projectTar()); err != nil {
		log.Fatal(err)
	}
	out, err := sb.Exec(ctx, "sh", "-c", "cd /work && make build 2>&1")
	var exit *sandbox.ExitError
	switch {
	case errors.As(err, &exit):
		log.Fatalf("build failed (rc=%d): %s", exit.Code, exit.Stderr)
	case err != nil:
		log.Fatal(err)
	}
	fmt.Print(out)

	// Freeze the built state; each branch is a fully independent sandbox.
	ckpt, err := sb.Checkpoint(ctx, "built")
	if err != nil {
		log.Fatal(err)
	}
	for i := range 2 {
		worker, err := ckpt.New(ctx)
		if err != nil {
			log.Fatal(err)
		}
		got, err := worker.Exec(ctx, "sh", "-c", fmt.Sprintf("echo worker %d && ls /work", i))
		if err != nil {
			log.Fatal(err)
		}
		fmt.Print(got)
		_ = worker.Close()
	}
}

func projectTar() *bytes.Reader {
	var buf bytes.Buffer
	tw := tar.NewWriter(&buf)
	body := "build:\n\techo built > out.txt\n"
	_ = tw.WriteHeader(&tar.Header{Name: "Makefile", Mode: 0o644, Size: int64(len(body))})
	_, _ = tw.Write([]byte(body))
	_ = tw.Close()
	return bytes.NewReader(buf.Bytes())
}

The rest of this guide is the per-method reference.

Connecting

client, err := sandbox.Connect("10.0.0.5:7777",
    sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))

HTTPS endpoints

Both SDKs accept host:port (plain HTTP), http://host[:port], or https://host[:port]. The default ports are 80 and 443. IPv6 hosts use brackets. Endpoints are origins: no userinfo, path prefix, query, or fragment.

For a public certificate, only the address changes:

client, err := sandbox.Connect("https://node-a.sandbox.example.com",
    sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))

For a private CA, load it into an x509.CertPool and pass sandbox.WithTLSConfig(&tls.Config{RootCAs: roots, MinVersion: tls.VersionTLS12}). The same trust configuration covers control requests and every agent relay. Certificate and hostname verification are enabled by default.

Python uses a standard ssl.SSLContext:

import os
import ssl
from cocoonsandbox import Client

client = Client(
    "https://node-a.sandbox.example.com",
    api_token=os.environ["SANDBOXD_TOKEN"],
    ssl_context=ssl.create_default_context(cafile="edge-ca.pem"),
)
with client.new("rt:24.04") as sb:
    assert sb.run(["true"]) == 0

Omit ssl_context for system trust. The SDK configures that context for HTTP/1.1; use a dedicated context if another caller requires a different ALPN. TLS handshakes share the existing dial timeout/cancellation budget. The agent connection then uses HTTP/1.1 Upgrade: silkd and remains a bidirectional stream; the guest protocol and port-forwarding frames do not change.

Data-plane calls share a handle’s relay connection: after a call the SDK keeps the connection for 30 seconds (WithKeepAlive tunes the window; 0 dials per call) and the next call on that handle sends its request on it, so a busy handle pays the dial, upgrade and TLS handshake once. A handle parks at most 8 idle connections. A kept connection counts as live for idle_hibernate_seconds until it closes, so keep the window below that setting; Close and Hibernate drop it at once. Long-lived streams (Watch, OpenPty, DialPort, an LSP session) take a connection of their own. Reuse needs a Unix client (the liveness peek is a MSG_PEEK); on Windows, and against a guest whose silkd predates the back-to-back protocol, every call dials as before.

An explicit scheme in an owner, redirect, peer, or Attach address wins; a bare address inherits the entry client’s scheme. Trust settings are shared across these connections. The SDK does not translate private addresses to public names: configure each node’s client_advertise as described in TLS deployment. Persist the complete owner URL with the sandbox ID and token when using Attach in another process.

Connecting to clusters

Nothing extra: dial any node. On a warm miss the entry node answers with a redirect and New follows it transparently (trying every candidate, so one dead peer never fails a claim); the returned handle is bound to the owning node’s owner_addr and all further calls go there directly.

To recover a handle when only id + token survived (say, across a process restart):

sb, err := client.Lookup(ctx, id, token)

Lookup asks the entry node, then queries all mesh peers concurrently and binds to whichever confirms ownership first.

Claiming

sb, err := client.New(ctx, "base:24.04",
    sandbox.WithNetwork(sandbox.NetEgress),
    sandbox.WithSize(sandbox.Medium),
    sandbox.WithVolumes(
        sandbox.Volume{Name: "imagenet"},
        sandbox.Volume{Name: "weights", Mount: "/models"},
        sandbox.Volume{Name: "scratch-db", Mode: "rw"}),
    sandbox.WithTimeout(10*time.Minute))
defer sb.Close()
option values default meaning
WithNetwork(n) NetNone, NetEgress NetNone Cloud Hypervisor network shape: NetNone disables the NIC and uses vsock-only I/O; NetEgress attaches a bridge/CNI NIC
WithSize(s) Small, Medium, Large, XLarge, XXLarge Small resource tier: 1cpu/512M, 2cpu/1G, 4cpu/4G, 4cpu/8G, 8cpu/16G
WithVolumes(volumes...) Volume{Name, Mount?, Mode?} entries none attach and mount up to eight unique catalog dataset disks; Mount defaults to /volumes/<name>; Mode is "ro" (default) or "rw" — "rw" requires the catalog entry’s writable: true; supported by Client.New and Template.New
WithVolumesAttachOnly() — mount attach the requested volumes without mounting them; the workload finds each device and owns the mount. Rejects a Volume.Mount locally
WithTimeout(d) duration server default 5m sandbox TTL, rounded up to seconds, server-capped at 24h. The node reaps the sandbox after the TTL even if the client vanishes

New returns when the sandbox’s silkd answers: a warm hit is milliseconds, a cold key can take the full boot. A volume claim may consume an ordinary warm VM and returns only after every requested disk is mounted; the finalized name, effective mount, and (for rw entries) mode are available in Sandbox.Volumes. Custom mounts must be absolute and clean, stay outside the guest OS tree, and cannot duplicate or nest. Sandbox.ID, Sandbox.Deadline, and Sandbox.FromCheckpoint (the lineage edge when branched) are exported. Sandbox.TemplateDigest is the exact content identity when the claim cloned a promoted template; it is empty for other sources. Owner() names the owning node, and Token() returns the per-sandbox bearer to persist with ID for a later Lookup; Close() releases the sandbox (releasing one already gone is not an error, and Close is bounded internally so it stays defer-friendly). Volume sandboxes cannot hibernate, fork, checkpoint, or promote. Passing WithVolumes or WithClaimRef to Checkpoint.New returns a local error: checkpoint branches support neither volumes nor a claim reference in this version. WithMetadata, WithArchiveOnExpire and WithEnv apply to New, Template.New, and Checkpoint.New: the labels follow the claim bounds, and a branch never inherits its source’s labels or expiry action.

WithVolumesAttachOnly() claims the same volumes without mounting them: the entries in Sandbox.Volumes carry an empty Mount, and the workload finds each device by polling /sys/block/*/serial for the catalog name — not guaranteed present when the claim returns, typically within ~100ms — then confirms the /dev/<blk> node itself exists before mounting. Everything above describes the default and is unchanged by this option. What changes is that the mount and its consistency are entirely yours: sandboxd writes and clears no dirty marker for an attach-only rw claim, because it cannot verify your unmount, so releasing without unmounting cleanly leaves the image as a crash would — see sandboxd-api for the full contract.

The caller-visible constraints are deliberate: volume claims may consume a warm VM, remain non-capturable, mount read-only by default, and require Cloud Hypervisor.

Discover the fleet entries this token may use before planning a claim:

catalog, err := client.Volumes(ctx) // []sandbox.VolumeInfo
for _, volume := range catalog {
    fmt.Println(volume.Name, volume.DefaultMount, volume.SizeBytes, volume.Available, volume.Nodes, volume.Writable)
}

Discovery returns the gossiped union and holder count; availability and size describe the connected node. Warm candidates retain normal ranking, filtered to nodes advertising every requested name. A promoted-template claim prefers a peer advertising both the template and every volume; when that intersection is empty, one volume holder may self-verify access to a shared template store before provisioning.

Claim env

sb, err := client.New(ctx, "rt:24.04", sandbox.WithEnv(map[string]sandbox.EnvVar{
    "MODE":   {Value: "prod"},
    "GW_KEY": {Value: "Bearer …", Guest: new(false)},
}))
err = sb.PatchEnv(ctx, map[string]*sandbox.EnvVar{"STAGE": {Value: "2"}, "MODE": nil})
env, err := sb.Env(ctx)     // host-only values come back empty
err = sb.SetEnv(ctx, nil)   // replaces the whole env; nil clears it

WithEnv sets the claim’s own env under the claim rules. An entry with a nil Guest is delivered into the guest; Guest: new(false) keeps it host-side, where only the node’s egress reads it (see secrets). Env, SetEnv and PatchEnv call the env verbs on the claim’s owner with the client’s API token, not the sandbox token. PatchEnv sets each entry, removes each nil one and keeps the rest as stored, so a host-only value is never resent. A host-only change applies to the next egress request at any time; a guest change needs a running guest and answers a 409 *APIError on a hibernated or archived sandbox.

Renewing

deadline, err := sb.Renew(ctx, time.Hour)   // the lease now ends an hour from now

Renew resets the lease to the given TTL from now and returns the granted deadline, which it also stores in Sandbox.Deadline. Zero asks for the server default of 5m; the TTL is rounded up to seconds and capped at 24h, as with WithTimeout. The grant is authoritative, and a renew can shorten a lease as well as extend it. Renew authenticates with the sandbox’s own token, so a handle from Lookup can renew. An archived sandbox is refused with a 409 *APIError: any call that reaches the guest wakes it, and then it can renew.

Hibernating

err := sb.Hibernate(ctx)   // snapshot + stop atomically; memory freed
// ... the next call that reaches the guest wakes it transparently:
out, err := sb.Exec(ctx, "cat", "/tmp/state")   // sessions & memory intact

Hibernate snapshots the VM and stops it in one atomic step — nothing the guest does can fall between the snapshot point and the stop. The handle stays valid: the first call that reaches the guest restores the VM (adding roughly a restore’s latency, tens of milliseconds on bare metal). The TTL keeps running — a hibernated sandbox is still reaped at its deadline, so claim with a WithTimeout that covers the idle period or Renew before it ends. When to hibernate is your policy; the node only provides the transition — unless the deployment opts into idle_hibernate_seconds (see deploy), which hibernates idle claims automatically with the same transparent wake. A claim with a connection live when the sweep checks it (a relay stream, a buffered exec, a preview dial, an egress request) is not swept; the idle clock restarts when that connection ends.

WithArchiveOnExpire() makes the lease’s end archive the sandbox instead of destroying it, whatever the deployment’s archive settings: the node hibernates it if it is running and moves it to the checkpoint store, and the next call that reaches the guest restores it. Fork children inherit the setting.

If that deployment also enables archive_after_seconds, archiving replaces the original claim deadline with the archive-retention deadline (or no deadline when archives are kept forever). Waking an archive starts a fresh lease of the length the claim was granted (the server default when it asked for none). The existing handle’s Sandbox.Deadline is the value returned when that handle was created or last renewed; call Client.Sandboxes to read the current server deadline after an archive/wake transition.

Forking

children, err := sb.Fork(ctx, 2, 10*time.Minute)   // []*Sandbox, own leases

Fork clones the sandbox into fresh, fully independent claims: memory, disk, and guest state (sessions, processes, tmpfs) duplicate at the fork point, and each child gets a distinct machine identity. The ttl bounds every child’s lifetime (zero = server default) — children never inherit the parent’s remaining lease. A running parent pauses briefly for the snapshot; a hibernated parent forks from its memory image without waking. All-or-nothing: on error no child survived. Count is capped at the node’s max_fork_count (default 16). Fork and Promote create node resources, so on a token-guarded node the client needs WithAPIToken — a sandbox handle alone cannot amplify.

Promoting to a template

tpl, err := sb.Promote(ctx, "myproj:v1")  // publish current state
child, err := tpl.New(ctx)                // clones the promoted state
fmt.Println(tpl.ContentDigest != "" && tpl.ContentDigest == child.TemplateDigest) // true
err = tpl.Delete(ctx)                     // caller owns the lifecycle

New asks for the promoted template only: once the template is deleted it answers 404 instead of taking a warm VM or cold-booting an image named like it.

Promote publishes the sandbox’s state as a template on its owning node, keyed by (name, the sandbox’s network lane, its size). Claims clone on demand (~a golden-clone’s latency); there is no warm pool for promoted templates unless the node’s config adds one. Re-promoting to the same name replaces the template. Template.ContentDigest identifies the published export bytes; a claim from that exact generation carries the same value in Sandbox.TemplateDigest. A caller pinning a mutable template name can compare the claim’s value with its expected digest and close/refuse a mismatch. Templates published by an older sandboxd have empty digests until they are re-promoted after the node is upgraded.

On the default local-disk backend templates live on one node, and on a cluster the parent claim may have been redirected — the returned Template handle is bound to the owning node. Its Delete and volume-less New reach that node; New(WithVolumes(...)) may follow one volume-placement redirect. The name-based calls (client.New("myproj:v1"), client.DeleteTemplate(...) with WithNetwork/WithSize when non-default) route cluster-wide via the mesh’s template gossip; they lag a promote or delete by about a gossip tick, so prefer the handle right after promoting (see Templates on a cluster).

Checkpoints — branching and time travel

ckpt, err := sb.Checkpoint(ctx, "after-setup")  // source keeps running
branch, err := ckpt.New(ctx)                     // fresh sandbox at the captured moment
err = ckpt.Delete(ctx)
ckpts, err := client.Checkpoints(ctx)            // node's checkpoints, newest first
known := client.Checkpoint("ck_…")               // known id; no listing round-trip

Checkpoint captures the sandbox’s full state — memory, disk, running processes — without stopping it (the same brief pause a fork takes), and ckpt.New branches any number of independent sandboxes from that exact moment; the checkpoint’s key axes apply and WithTimeout may set each branch’s TTL. Successive checkpoints of sources and branches form a tree. Checkpoints live in the node’s checkpoint store — a shared FUSE mount or a checkpoint_store of kind s3 lets any node branch them. client.Checkpoints lists the connected node’s records, while client.Checkpoint(id) creates an entry-node-bound handle for an already known id without listing. Checkpoint.New follows an owner probe/redirect and may heal a missing record locally; Checkpoint.Delete acts on the handle’s bound node. Checkpoint creation is resource-creating and takes the api token, like fork.

Checkpoint.Delete also asks every peer that node currently sees to drop any replica a heal pulled — best-effort eventual cleanup, not a fleet-wide revocation. A peer that misses the broadcast (offline, partitioned, or joined later) keeps serving branches from its replica until the node’s checkpoint_ttl_hours ages it out; enabling peer heal requires that TTL to be set, so every healed replica has a cleanup bound while healing stays on. A node later run with healing off and that TTL back at 0 keeps such a replica until an explicit delete.

Language servers (LSP)

lsp, err := sb.StartLsp(ctx, "python", "/workspace")  // flavor image provides the server
conn, err := lsp.Request(ctx)                          // JSON-RPC byte stream (frame it yourself)
// ... speak LSP over conn (Content-Length framed JSON-RPC) ...
err = lsp.Stop(ctx)

StartLsp spawns the language server the flavor image ships for the language (its argv in /etc/silkd/lsp.d/<language>; the python flavor bakes pylsp); the base image has none, so it returns silkd’s typed not_found. silkd is a broker — it pipes JSON-RPC bytes between your Request stream and the server’s stdio without parsing LSP semantics, so the caller frames (Content-Length) and correlates by request id. A server serves one Request stream for its lifetime: closing the stream ends the session and reaps the server (start a new one to keep working); Stop kills it early.

Reaching guest ports

conn, err := sb.DialPort(ctx, 8080)          // net.Conn to 127.0.0.1:8080 in the guest
l, err := sb.ProxyPort(ctx, "127.0.0.1:0", 8080)  // local listener piping to it

DialPort opens a TCP connection to a port inside the sandbox, relayed over the silkd protocol — it works on the no-network lane, where the vsock relay is the only way in. The returned net.Conn supports half-close (CloseWrite) but not deadlines; bound the ctx instead. A dead port fails with silkd’s not_found. ProxyPort serves it to unmodified local tools (browsers, curl) via a local listener.

Preview URLs

url, err := sb.PreviewURL(ctx, 8080, 30*time.Minute)

Mints a signed, shareable URL serving the guest HTTP port from a plain browser via the node’s preview listener. Minting is resource-creating and takes the api token, like fork and checkpoint. The TTL is clamped to the claim’s remaining lease, and the URL dies with the sandbox — release or reap revokes it with no extra state. Answers 501 when the node has no preview_listen configured; see deploy.

Node info

info, err := client.Info(ctx)   // pools, promoted templates, claims, drain, capacity, and mesh peers

NodeInfo.Templates lists the promoted templates the node holds, each with its ContentDigest and Labels.

NodeInfo.AtCapacity distinguishes a refill parked by node capacity from one still filling its target; AtCapacityReason carries the engine’s reason. Both decode to their zero values when the server omits them.

Running commands

out, err := sb.Exec(ctx, "python3", "script.py")   // stdout; *ExitError on rc != 0

code, err := sb.Run(ctx, sandbox.Cmd{
    Argv:   []string{"bash", "-c", "make test"},
    Cwd:    "/work",
    Env:    map[string]string{"CI": "1"},
    User:   "ubuntu",
    Stdin:  strings.NewReader(input),
    Stdout: os.Stdout,
    Stderr: os.Stderr,
})

Cmd fields: Argv (required), Cwd, Env, User (de-escalation inside the guest), Session (run inside a persistent session, below), Stdin (nil closes the child’s stdin immediately; do not share one blocking reader across Runs), Stdout/Stderr (nil discards). With Session set only Argv applies: the session owns the cwd, environment and user, its commands read /dev/null, and stderr arrives merged into stdout. A command’s environment is PATH, TERM and the lane’s proxy variables plus Env — not the image’s ENV; HOME and USER are set only with User. The guest kills a foreground command whose connection dropped only once it next writes output, so Run does not leave that to the drop: when its ctx ends after the guest reported the command’s pid, Run sends that pid a kill over a connection of its own (best effort, bounded to 5 s) before returning the ctx error. A client that only drops the connection leaves a silent command (sleep, a quiet build) running in the guest; find its pid with Ps and Kill it. Use Spawn for work that must outlive the connection. A command run inside a Session is the session’s: it reports no pid, so no cancel kills it and it does not appear in Ps; it runs to completion, and Session.Close is what ends the session.

Non-zero exits surface as *sandbox.ExitError{Code, Stderr} from Exec (alongside partial stdout); Run returns the code directly.

Background processes

pid, err := sb.Spawn(ctx, sandbox.Cmd{Argv: []string{"sh", "-c", "make build"}})
procs, err := sb.Ps(ctx)                          // []wire.ProcInfo{PID, Argv, Detached, State, ExitCode, ...}
code, exited, err := sb.Logs(ctx, pid, w, nil)    // replay the bounded output ring
code, exited, err = sb.Attach(ctx, pid, w, nil)   // replay, then follow live until exit
err = sb.Kill(ctx, pid, 0)                        // 0 = SIGKILL

Spawn returns as soon as the process starts; it keeps running with a bounded output ring. Logs replays the ring and reports the exit code once the process has ended (exited=false while running); Attach follows live output until exit — the replay and the live stream hand off atomically, so no chunk is lost or doubled between them. Killing an already-exited process is a no-op success (its OS pid may be recycled; silkd never signals a reaped child). An exited process leaves the table after 5 minutes; from then on its pid answers not_found to every verb.

Sessions

A session is a real persistent shell: cwd, env and shell state survive across calls.

sess, err := sb.NewSession(ctx,
    sandbox.WithSessionCwd("/work"),
    sandbox.WithSessionEnv(map[string]string{"PATH": "…"}))
out, err := sess.Exec(ctx, "export", "MARK=1")     // persists
out, err  = sess.Exec(ctx, "sh", "-c", "echo $MARK")
err = sess.Close(ctx)

ids, err := sb.Sessions(ctx)                       // live session ids

Idle sessions are reaped guest-side after 30 minutes. Close on a session that is already gone — and Lsp.Stop on a server that already exited — answers silkd’s not_found; ignore it in a deferred cleanup.

Files

err  := sb.WriteFile(ctx, "/work/a.txt", data, nil)   // atomic; *uint32 mode optional
data, err := sb.ReadFile(ctx, "/work/a.txt")
err  = sb.ReadFileTo(ctx, "/work/big.bin", w)          // streams into an io.Writer
ents, err := sb.ListDir(ctx, "/work")                  // []wire.DirEntry{Name,Kind,Size}
info, err := sb.Stat(ctx, "/work/a.txt")               // wire.FileInfo{Kind,Size,Mode,MtimeEpochSecs}
err  = sb.Mkdir(ctx, "/work/sub", true)                // parents
err  = sb.Remove(ctx, "/work/sub", true)               // recursive
err  = sb.Rename(ctx, "/a", "/b")

Writes stream any size and commit via temp-file rename: a mid-stream failure never leaves a truncated destination, and overwriting an executable keeps its exec bit.

Project trees

err = sb.Push(ctx, "/work", tarReader)   // extract a tar stream under /work
err = sb.Pull(ctx, "/work", tarWriter)   // stream /work back as a tar

Push is the only project-ingestion path on the no-network lane.

matches, err := sb.Find(ctx, "/work", `TODO|FIXME`, "*.go")
// []wire.Match{File, Line, Content}; glob is anchored *? wildcards on the file name
for m, err := range sb.FindSeq(ctx, "/work", `TODO`, "") { // streamed; a break ends the walk in the guest
	if err != nil {
		return err
	}
	if m.Line > 100 {
		break
	}
}

results, err := sb.Replace(ctx, []string{"/work/main.go"}, `foo`, "bar")
// []wire.Replaced{File, Replacements}; per-file atomic

Patterns are regular expressions evaluated in the guest — no shell quoting. Replace is atomic per file, not per list: it fails at the first missing, unreadable or non-UTF-8 path and the files before it stay rewritten, so pass it paths a Find just returned. A Find fails as a whole when one match’s frame would pass the 8 MiB cap.

Watching

w, err := sb.Watch(ctx, "/work", true)
defer w.Close()
for ev := range w.Events() {           // wire.Event{Kind, Path}
    fmt.Println(ev.Kind, ev.Path)      // created|modified|deleted|renamed
}
err = w.Err()                          // why the stream ended; nil after Close

Watch returns once the guest acknowledges the watch is armed — events caused after it returns are guaranteed captured. A bad path fails synchronously; if the consumer falls too far behind, Err reports a terminal overflow instead of the stream silently dropping events.

Git

err  = sb.GitClone(ctx, url, "/work/repo", "main", 0, token) // egress lane only; depth > 0 = shallow
st,  err := sb.GitStatus(ctx, "/work/repo")   // Branch, Ahead, Behind, Files[]; Truncated when the list was cut at ~1 MiB
err  = sb.GitAdd(ctx, "/work/repo", "a.txt")
hash, err := sb.GitCommit(ctx, "/work/repo", "message", "Dev <dev@example.com>")
err  = sb.GitPush(ctx, "/work/repo", token)   // egress lane only
err  = sb.GitPull(ctx, "/work/repo", token)   // egress lane only
br,  err := sb.GitBranches(ctx, "/work/repo") // Current + Branches
err  = sb.GitCreateBranch(ctx, "/work/repo", "feature")
err  = sb.GitCheckout(ctx, "/work/repo", "feature")
err  = sb.GitDeleteBranch(ctx, "/work/repo", "feature")

Results are structured (porcelain v2 under the hood), never scraped stdout. Auth tokens travel as an in-memory header, never touching guest disk. On the no-network lane, clone/push/pull fail fast with a typed unimplemented error pointing at Push.

Terminals

pty, err := sb.OpenPty(ctx, wire.PtyOpen{Cols: 120, Rows: 40})
defer pty.Close()
pty.Write([]byte("make test\n"))
io.Copy(os.Stdout, pty)                   // EOF when the shell exits
code, ok := pty.ExitCode()
err = pty.Resize(ctx, 200, 50)

A PTY is a tracked guest process (pty.PID); closing the handle (or the ctx) tears the shell down.

Node operations

Verbs for operating a node — Drain, Uncordon, SetPools, SetPoolsCluster and the tenant verbs need the root token — plus the reference the aggregated apiserver claims under:

sb, _ := client.New(ctx, "rt:24.04", sandbox.WithClaimRef("ns/workload"))
list, _ := client.Sandboxes(ctx)          // one SandboxSummary per live claim — never tokens
info, _ := client.Drain(ctx)              // cordon: refuse new claims, run leases out
info, _ = client.Uncordon(ctx)
info, _ = client.SetPools(ctx, pools)     // retune warm targets without a restart
res, _ := client.SetPoolsCluster(ctx, pools) // per-node results; retry the failures
sb = client.Attach(ownerAddr, id, token)  // bind a known handle, no lookup round-trip

err = client.PutTenant(ctx, sandbox.TenantSpec{Name: "u:42", Token: tok, MaxClaims: 10})
res, _ = client.PutTenantCluster(ctx, sandbox.TenantSpec{Name: "u:42", MaxClaims: 20}) // token kept
res, _ = client.DeleteTenantCluster(ctx, "u:42")
tenants, _ := client.Tenants(ctx)         // names, caps, live claims, removed tenants; never tokens

Sandboxes is scoped to the calling token, so a tenant sees only its own claims; the fields are those of GET /v1/sandboxes, including each claim’s Metadata, CPUCount, and MemTotalBytes. Drain leaves live claims alone — poll Info until Claimed is zero. The tenant verbs follow /v1/tenants: a TenantSpec without Token keeps the stored token, and the *Cluster forms return one NodeResult per node to retry (see cluster). SandboxSummary.Tenant names each claim’s tenant.

Error handling

var e *wire.ErrorResp
if errors.As(err, &e) && e.Kind == wire.KindUnimplemented {
    // no-network lane: fall back to sb.Push
}

Context cancellation is honored on every call that takes a ctx: canceling it closes the underlying connection. Data-plane calls return ctx.Err() directly; control-plane verbs return a *url.Error wrapping it, so test with errors.Is, never ==. Close is the exception: it takes no ctx.