import sandbox "github.com/cocoonstack/sandbox/sdk/go"
The SDK has no third-party dependencies (only the sibling protocol/wire
module). One Client talks to one entry node; sandbox handles dial their
owning node directly, so a client works unchanged against a single node or a
cluster.
Claim a sandbox, push a project into it, run a build, then freeze the built state and fan out two independent workers from that exact moment:
package main
import (
"archive/tar"
"bytes"
"context"
"errors"
"fmt"
"log"
"os"
"time"
sandbox "github.com/cocoonstack/sandbox/sdk/go"
)
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
client, err := sandbox.Connect(os.Getenv("SANDBOXD_ADDR"),
sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))
if err != nil {
log.Fatal(err)
}
sb, err := client.New(ctx, "rt:24.04",
sandbox.WithSize(sandbox.Medium),
sandbox.WithTimeout(10*time.Minute))
if err != nil {
log.Fatal(err)
}
defer sb.Close()
fmt.Printf("claimed %s on %s\n", sb.ID, sb.Owner())
// Push is the only ingestion path on the no-network lane.
if err := sb.Push(ctx, "/work", projectTar()); err != nil {
log.Fatal(err)
}
out, err := sb.Exec(ctx, "sh", "-c", "cd /work && make build 2>&1")
var exit *sandbox.ExitError
switch {
case errors.As(err, &exit):
log.Fatalf("build failed (rc=%d): %s", exit.Code, exit.Stderr)
case err != nil:
log.Fatal(err)
}
fmt.Print(out)
// Freeze the built state; each branch is a fully independent sandbox.
ckpt, err := sb.Checkpoint(ctx, "built")
if err != nil {
log.Fatal(err)
}
for i := range 2 {
worker, err := ckpt.New(ctx)
if err != nil {
log.Fatal(err)
}
got, err := worker.Exec(ctx, "sh", "-c", fmt.Sprintf("echo worker %d && ls /work", i))
if err != nil {
log.Fatal(err)
}
fmt.Print(got)
_ = worker.Close()
}
}
func projectTar() *bytes.Reader {
var buf bytes.Buffer
tw := tar.NewWriter(&buf)
body := "build:\n\techo built > out.txt\n"
_ = tw.WriteHeader(&tar.Header{Name: "Makefile", Mode: 0o644, Size: int64(len(body))})
_, _ = tw.Write([]byte(body))
_ = tw.Close()
return bytes.NewReader(buf.Bytes())
}
The rest of this guide is the per-method reference.
client, err := sandbox.Connect("10.0.0.5:7777",
sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))
Connect(addr, opts...) — addr accepts a comma-separated seed list for
forward compatibility; the current version uses the first entry.WithAPIToken(token) — the node token: a root api_token (full access)
or a tenant token (resource-creating verbs only; operator surfaces like
Info answer it 403). On a cluster every node shares the same root token
and the same tenants set.WithHTTPClient(client) — replace the control-plane HTTP client when the
caller needs a custom transport, proxy, or timeout. An http.Transport’s
TLS configuration also supplies the agent relay’s certificate settings;
HTTP proxies and custom dialers apply only to control requests.WithTLSConfig(config) — set certificate verification for both HTTPS
requests and agent relays. With WithHTTPClient, its transport must be an
*http.Transport; the SDK clones it rather than changing the caller’s client.Both SDKs accept host:port (plain HTTP), http://host[:port], or
https://host[:port]. The default ports are 80 and 443. IPv6 hosts use brackets.
Endpoints are origins: no userinfo, path prefix, query, or fragment.
For a public certificate, only the address changes:
client, err := sandbox.Connect("https://node-a.sandbox.example.com",
sandbox.WithAPIToken(os.Getenv("SANDBOXD_TOKEN")))
For a private CA, load it into an x509.CertPool and pass
sandbox.WithTLSConfig(&tls.Config{RootCAs: roots, MinVersion: tls.VersionTLS12}).
The same trust configuration covers control requests and every agent relay.
Certificate and hostname verification are enabled by default.
Python uses a standard ssl.SSLContext:
import os
import ssl
from cocoonsandbox import Client
client = Client(
"https://node-a.sandbox.example.com",
api_token=os.environ["SANDBOXD_TOKEN"],
ssl_context=ssl.create_default_context(cafile="edge-ca.pem"),
)
with client.new("rt:24.04") as sb:
assert sb.run(["true"]) == 0
Omit ssl_context for system trust. The SDK configures that context for
HTTP/1.1; use a dedicated context if another caller requires a different ALPN.
TLS handshakes share the existing dial timeout/cancellation budget. The agent
connection then uses HTTP/1.1 Upgrade: silkd and remains a bidirectional
stream; the guest protocol and port-forwarding frames do not change.
Data-plane calls share a handle’s relay connection: after a call the SDK
keeps the connection for 30 seconds (WithKeepAlive tunes the window; 0
dials per call) and the next call on that handle sends its request on it, so
a busy handle pays the dial, upgrade and TLS handshake once. A handle parks
at most 8 idle connections. A kept connection counts as live for
idle_hibernate_seconds until it closes, so keep the window below that
setting; Close and Hibernate drop it at once. Long-lived streams
(Watch, OpenPty, DialPort, an LSP session) take a connection of their
own. Reuse needs a Unix client (the liveness peek is a MSG_PEEK); on
Windows, and against a guest whose silkd predates the back-to-back protocol,
every call dials as before.
An explicit scheme in an owner, redirect, peer, or Attach address wins;
a bare address inherits the entry client’s scheme. Trust settings are shared
across these connections. The SDK does not translate private addresses to
public names: configure each node’s client_advertise as described in
TLS deployment. Persist the complete owner URL
with the sandbox ID and token when using Attach in another process.
Nothing extra: dial any node. On a warm miss the entry node answers with a
redirect and New follows it transparently (trying every candidate, so one
dead peer never fails a claim); the returned handle is bound to the owning
node’s owner_addr and all further calls go there directly.
To recover a handle when only id + token survived (say, across a process
restart):
sb, err := client.Lookup(ctx, id, token)
Lookup asks the entry node, then queries all mesh peers concurrently and
binds to whichever confirms ownership first.
sb, err := client.New(ctx, "base:24.04",
sandbox.WithNetwork(sandbox.NetEgress),
sandbox.WithSize(sandbox.Medium),
sandbox.WithVolumes(
sandbox.Volume{Name: "imagenet"},
sandbox.Volume{Name: "weights", Mount: "/models"},
sandbox.Volume{Name: "scratch-db", Mode: "rw"}),
sandbox.WithTimeout(10*time.Minute))
defer sb.Close()
| option | values | default | meaning |
|---|---|---|---|
WithNetwork(n) |
NetNone, NetEgress |
NetNone |
Cloud Hypervisor network shape: NetNone disables the NIC and uses vsock-only I/O; NetEgress attaches a bridge/CNI NIC |
WithSize(s) |
Small, Medium, Large, XLarge, XXLarge |
Small |
resource tier: 1cpu/512M, 2cpu/1G, 4cpu/4G, 4cpu/8G, 8cpu/16G |
WithVolumes(volumes...) |
Volume{Name, Mount?, Mode?} entries |
none | attach and mount up to eight unique catalog dataset disks; Mount defaults to /volumes/<name>; Mode is "ro" (default) or "rw" — "rw" requires the catalog entry’s writable: true; supported by Client.New and Template.New |
WithVolumesAttachOnly() |
— | mount | attach the requested volumes without mounting them; the workload finds each device and owns the mount. Rejects a Volume.Mount locally |
WithTimeout(d) |
duration | server default 5m | sandbox TTL, rounded up to seconds, server-capped at 24h. The node reaps the sandbox after the TTL even if the client vanishes |
New returns when the sandbox’s silkd answers: a warm hit is milliseconds,
a cold key can take the full boot. A volume claim may consume an ordinary warm
VM and returns only after every requested disk is mounted; the finalized
name, effective mount, and (for rw entries) mode are available in
Sandbox.Volumes. Custom mounts must be absolute and clean, stay outside the
guest OS tree, and cannot duplicate or nest.
Sandbox.ID, Sandbox.Deadline, and
Sandbox.FromCheckpoint (the lineage edge when branched) are exported.
Sandbox.TemplateDigest is the exact content identity when the claim cloned a
promoted template; it is empty for other sources. Owner() names the owning
node, and Token() returns the per-sandbox bearer to persist with ID for a
later Lookup; Close() releases the sandbox (releasing one already gone is
not an error, and Close is bounded internally so it stays defer-friendly).
Volume sandboxes cannot hibernate, fork, checkpoint, or promote. Passing
WithVolumes or WithClaimRef to Checkpoint.New returns a local error:
checkpoint branches support neither volumes nor a claim reference in this version.
WithMetadata, WithArchiveOnExpire and WithEnv apply to New, Template.New, and
Checkpoint.New: the labels follow the claim bounds,
and a branch never inherits its source’s labels or expiry action.
WithVolumesAttachOnly() claims the same volumes without mounting them: the
entries in Sandbox.Volumes carry an empty Mount, and the workload finds
each device by polling /sys/block/*/serial for the catalog name — not
guaranteed present when the claim returns, typically within ~100ms — then
confirms the /dev/<blk> node itself exists before mounting. Everything
above describes the default and is unchanged by this option. What changes is
that the mount and its consistency are entirely yours: sandboxd writes and
clears no dirty marker for an attach-only rw claim, because it cannot verify
your unmount, so releasing without unmounting cleanly leaves the image as a
crash would — see
sandboxd-api for the full contract.
The caller-visible constraints are deliberate: volume claims may consume a warm VM, remain non-capturable, mount read-only by default, and require Cloud Hypervisor.
Discover the fleet entries this token may use before planning a claim:
catalog, err := client.Volumes(ctx) // []sandbox.VolumeInfo
for _, volume := range catalog {
fmt.Println(volume.Name, volume.DefaultMount, volume.SizeBytes, volume.Available, volume.Nodes, volume.Writable)
}
Discovery returns the gossiped union and holder count; availability and size describe the connected node. Warm candidates retain normal ranking, filtered to nodes advertising every requested name. A promoted-template claim prefers a peer advertising both the template and every volume; when that intersection is empty, one volume holder may self-verify access to a shared template store before provisioning.
sb, err := client.New(ctx, "rt:24.04", sandbox.WithEnv(map[string]sandbox.EnvVar{
"MODE": {Value: "prod"},
"GW_KEY": {Value: "Bearer …", Guest: new(false)},
}))
err = sb.PatchEnv(ctx, map[string]*sandbox.EnvVar{"STAGE": {Value: "2"}, "MODE": nil})
env, err := sb.Env(ctx) // host-only values come back empty
err = sb.SetEnv(ctx, nil) // replaces the whole env; nil clears it
WithEnv sets the claim’s own env under the claim rules.
An entry with a nil Guest is delivered into the guest; Guest: new(false)
keeps it host-side, where only the node’s egress reads it (see
secrets). Env, SetEnv and PatchEnv
call the env verbs on
the claim’s owner with the client’s API token, not the sandbox token.
PatchEnv sets each entry, removes each nil one and keeps the rest as stored,
so a host-only value is never resent. A host-only change applies to the next
egress request at any time; a guest change needs a running guest and answers
a 409 *APIError on a hibernated or archived sandbox.
deadline, err := sb.Renew(ctx, time.Hour) // the lease now ends an hour from now
Renew resets the lease to the given TTL from now and returns the granted
deadline, which it also stores in Sandbox.Deadline. Zero asks for the server
default of 5m; the TTL is rounded up to seconds and capped at 24h, as with
WithTimeout. The grant is authoritative, and a renew can shorten a lease as
well as extend it. Renew authenticates with the sandbox’s own token, so a
handle from Lookup can renew. An archived sandbox is refused with a 409
*APIError: any call that reaches the guest wakes it, and then it can renew.
err := sb.Hibernate(ctx) // snapshot + stop atomically; memory freed
// ... the next call that reaches the guest wakes it transparently:
out, err := sb.Exec(ctx, "cat", "/tmp/state") // sessions & memory intact
Hibernate snapshots the VM and stops it in one atomic step — nothing the
guest does can fall between the snapshot point and the stop. The handle
stays valid: the first call that reaches the guest restores the VM (adding
roughly a restore’s latency, tens of milliseconds on bare metal). The TTL
keeps running — a hibernated sandbox is still reaped at its deadline, so
claim with a WithTimeout that covers the idle period or Renew before it
ends. When to hibernate is your policy; the node only provides the
transition — unless the deployment opts into idle_hibernate_seconds (see
deploy), which hibernates idle claims
automatically with the same transparent wake. A claim with a connection live
when the sweep checks it (a relay stream, a buffered exec, a preview dial, an
egress request) is not swept; the idle clock restarts when that connection ends.
WithArchiveOnExpire() makes the lease’s end archive the sandbox instead of
destroying it, whatever the deployment’s archive settings: the node hibernates
it if it is running and moves it to the checkpoint store, and the next call
that reaches the guest restores it. Fork children inherit the setting.
If that deployment also enables archive_after_seconds, archiving replaces
the original claim deadline with the archive-retention deadline (or no
deadline when archives are kept forever). Waking an archive starts a fresh
lease of the length the claim was granted (the server default when it asked
for none). The existing handle’s Sandbox.Deadline is the value
returned when that handle was created or last renewed; call
Client.Sandboxes to read the current server deadline after an
archive/wake transition.
children, err := sb.Fork(ctx, 2, 10*time.Minute) // []*Sandbox, own leases
Fork clones the sandbox into fresh, fully independent claims: memory,
disk, and guest state (sessions, processes, tmpfs) duplicate at the fork
point, and each child gets a distinct machine identity. The ttl bounds every
child’s lifetime (zero = server default) — children never inherit the
parent’s remaining lease. A running parent pauses briefly for the snapshot;
a hibernated parent forks from its memory image without waking.
All-or-nothing: on error no child survived. Count is capped at the node’s
max_fork_count (default 16).
Fork and Promote create node resources, so on a token-guarded node the
client needs WithAPIToken — a sandbox handle alone cannot amplify.
tpl, err := sb.Promote(ctx, "myproj:v1") // publish current state
child, err := tpl.New(ctx) // clones the promoted state
fmt.Println(tpl.ContentDigest != "" && tpl.ContentDigest == child.TemplateDigest) // true
err = tpl.Delete(ctx) // caller owns the lifecycle
New asks for the promoted template only: once the template is deleted it
answers 404 instead of taking a warm VM or cold-booting an image named like it.
Promote publishes the sandbox’s state as a template on its owning node,
keyed by (name, the sandbox’s network lane, its size). Claims clone on
demand (~a golden-clone’s latency); there is no warm pool for promoted
templates unless the node’s config adds one. Re-promoting to the same name
replaces the template. Template.ContentDigest identifies the published
export bytes; a claim from that exact generation carries the same value in
Sandbox.TemplateDigest. A caller pinning a mutable template name can compare
the claim’s value with its expected digest and close/refuse a mismatch.
Templates published by an older sandboxd have empty digests until they are
re-promoted after the node is upgraded.
On the default local-disk backend templates live on one node, and on a
cluster the parent claim may have been redirected — the returned Template
handle is bound to the owning node. Its Delete and volume-less New reach
that node; New(WithVolumes(...)) may follow one volume-placement redirect.
The name-based calls
(client.New("myproj:v1"), client.DeleteTemplate(...) with
WithNetwork/WithSize when non-default) route cluster-wide via the
mesh’s template gossip; they lag a promote or delete by about a gossip
tick, so prefer the handle right after promoting (see
Templates on a cluster).
ckpt, err := sb.Checkpoint(ctx, "after-setup") // source keeps running
branch, err := ckpt.New(ctx) // fresh sandbox at the captured moment
err = ckpt.Delete(ctx)
ckpts, err := client.Checkpoints(ctx) // node's checkpoints, newest first
known := client.Checkpoint("ck_…") // known id; no listing round-trip
Checkpoint captures the sandbox’s full state — memory, disk, running
processes — without stopping it (the same brief pause a fork takes), and
ckpt.New branches any number of independent sandboxes from that exact
moment; the checkpoint’s key axes apply and WithTimeout may set each
branch’s TTL. Successive checkpoints of sources and branches form a tree.
Checkpoints live in the node’s checkpoint store — a shared FUSE mount or
a checkpoint_store of kind s3 lets any node branch them.
client.Checkpoints lists the connected node’s records, while
client.Checkpoint(id) creates an entry-node-bound handle for an already known
id without listing. Checkpoint.New follows an owner probe/redirect and may
heal a missing record locally; Checkpoint.Delete acts on the handle’s bound
node. Checkpoint creation is resource-creating and takes the api token, like
fork.
Checkpoint.Delete also asks every peer that node currently sees to drop any
replica a heal pulled — best-effort eventual cleanup, not a fleet-wide
revocation. A peer that misses the broadcast (offline, partitioned, or joined
later) keeps serving branches from its replica until the node’s
checkpoint_ttl_hours ages it out; enabling peer heal requires that TTL to be
set, so every healed replica has a cleanup bound while healing stays on. A node
later run with healing off and that TTL back at 0 keeps such a replica until an
explicit delete.
lsp, err := sb.StartLsp(ctx, "python", "/workspace") // flavor image provides the server
conn, err := lsp.Request(ctx) // JSON-RPC byte stream (frame it yourself)
// ... speak LSP over conn (Content-Length framed JSON-RPC) ...
err = lsp.Stop(ctx)
StartLsp spawns the language server the flavor image ships for the
language (its argv in /etc/silkd/lsp.d/<language>; the python flavor bakes
pylsp); the base image has none, so it returns silkd’s typed not_found.
silkd is a broker — it pipes JSON-RPC bytes between your Request stream
and the server’s stdio without parsing LSP semantics, so the caller frames
(Content-Length) and correlates by request id. A server serves one
Request stream for its lifetime: closing the stream ends the session and
reaps the server (start a new one to keep working); Stop kills it early.
conn, err := sb.DialPort(ctx, 8080) // net.Conn to 127.0.0.1:8080 in the guest
l, err := sb.ProxyPort(ctx, "127.0.0.1:0", 8080) // local listener piping to it
DialPort opens a TCP connection to a port inside the sandbox, relayed over
the silkd protocol — it works on the no-network lane, where the vsock relay
is the only way in. The returned net.Conn supports half-close
(CloseWrite) but not deadlines; bound the ctx instead. A dead port fails
with silkd’s not_found. ProxyPort serves it to unmodified local tools
(browsers, curl) via a local listener.
url, err := sb.PreviewURL(ctx, 8080, 30*time.Minute)
Mints a signed, shareable URL serving the guest HTTP port from a plain
browser via the node’s preview listener. Minting is resource-creating and
takes the api token, like fork and checkpoint. The TTL is clamped to the
claim’s remaining lease, and the URL dies with the sandbox — release or reap
revokes it with no extra state. Answers 501 when the node has no
preview_listen configured; see deploy.
info, err := client.Info(ctx) // pools, promoted templates, claims, drain, capacity, and mesh peers
NodeInfo.Templates lists the promoted templates the node holds, each with its
ContentDigest and Labels.
NodeInfo.AtCapacity distinguishes a refill parked by node capacity from one
still filling its target; AtCapacityReason carries the engine’s reason. Both
decode to their zero values when the server omits them.
out, err := sb.Exec(ctx, "python3", "script.py") // stdout; *ExitError on rc != 0
code, err := sb.Run(ctx, sandbox.Cmd{
Argv: []string{"bash", "-c", "make test"},
Cwd: "/work",
Env: map[string]string{"CI": "1"},
User: "ubuntu",
Stdin: strings.NewReader(input),
Stdout: os.Stdout,
Stderr: os.Stderr,
})
Cmd fields: Argv (required), Cwd, Env, User (de-escalation inside
the guest), Session (run inside a persistent session, below), Stdin
(nil closes the child’s stdin immediately; do not share one blocking reader
across Runs), Stdout/Stderr (nil discards). With Session set only
Argv applies: the session owns the cwd, environment and user, its commands
read /dev/null, and stderr arrives merged into stdout. A command’s
environment is PATH, TERM and the lane’s proxy variables plus Env — not
the image’s ENV; HOME and USER are set only with User. The guest kills
a foreground command whose connection dropped only once it next writes output,
so Run does not leave that to the drop: when its ctx ends after the guest
reported the command’s pid, Run sends that pid a kill over a connection of
its own (best effort, bounded to 5 s) before returning the ctx error. A client
that only drops the connection leaves a silent command (sleep, a quiet build)
running in the guest; find its pid with Ps and Kill it. Use Spawn for
work that must outlive the connection. A command run inside a Session is the
session’s: it reports no pid, so no cancel kills it and it does not appear in
Ps; it runs to completion, and Session.Close is what ends the session.
Non-zero exits surface as *sandbox.ExitError{Code, Stderr} from Exec
(alongside partial stdout); Run returns the code directly.
pid, err := sb.Spawn(ctx, sandbox.Cmd{Argv: []string{"sh", "-c", "make build"}})
procs, err := sb.Ps(ctx) // []wire.ProcInfo{PID, Argv, Detached, State, ExitCode, ...}
code, exited, err := sb.Logs(ctx, pid, w, nil) // replay the bounded output ring
code, exited, err = sb.Attach(ctx, pid, w, nil) // replay, then follow live until exit
err = sb.Kill(ctx, pid, 0) // 0 = SIGKILL
Spawn returns as soon as the process starts; it keeps running with a
bounded output ring. Logs replays the ring and reports the exit code once
the process has ended (exited=false while running); Attach follows live
output until exit — the replay and the live stream hand off atomically, so
no chunk is lost or doubled between them. Killing an already-exited process
is a no-op success (its OS pid may be recycled; silkd never signals a
reaped child). An exited process leaves the table after 5 minutes; from then
on its pid answers not_found to every verb.
A session is a real persistent shell: cwd, env and shell state survive across calls.
sess, err := sb.NewSession(ctx,
sandbox.WithSessionCwd("/work"),
sandbox.WithSessionEnv(map[string]string{"PATH": "…"}))
out, err := sess.Exec(ctx, "export", "MARK=1") // persists
out, err = sess.Exec(ctx, "sh", "-c", "echo $MARK")
err = sess.Close(ctx)
ids, err := sb.Sessions(ctx) // live session ids
Idle sessions are reaped guest-side after 30 minutes. Close on a session
that is already gone — and Lsp.Stop on a server that already exited — answers
silkd’s not_found; ignore it in a deferred cleanup.
err := sb.WriteFile(ctx, "/work/a.txt", data, nil) // atomic; *uint32 mode optional
data, err := sb.ReadFile(ctx, "/work/a.txt")
err = sb.ReadFileTo(ctx, "/work/big.bin", w) // streams into an io.Writer
ents, err := sb.ListDir(ctx, "/work") // []wire.DirEntry{Name,Kind,Size}
info, err := sb.Stat(ctx, "/work/a.txt") // wire.FileInfo{Kind,Size,Mode,MtimeEpochSecs}
err = sb.Mkdir(ctx, "/work/sub", true) // parents
err = sb.Remove(ctx, "/work/sub", true) // recursive
err = sb.Rename(ctx, "/a", "/b")
Writes stream any size and commit via temp-file rename: a mid-stream failure never leaves a truncated destination, and overwriting an executable keeps its exec bit.
err = sb.Push(ctx, "/work", tarReader) // extract a tar stream under /work
err = sb.Pull(ctx, "/work", tarWriter) // stream /work back as a tar
Push is the only project-ingestion path on the no-network lane.
matches, err := sb.Find(ctx, "/work", `TODO|FIXME`, "*.go")
// []wire.Match{File, Line, Content}; glob is anchored *? wildcards on the file name
for m, err := range sb.FindSeq(ctx, "/work", `TODO`, "") { // streamed; a break ends the walk in the guest
if err != nil {
return err
}
if m.Line > 100 {
break
}
}
results, err := sb.Replace(ctx, []string{"/work/main.go"}, `foo`, "bar")
// []wire.Replaced{File, Replacements}; per-file atomic
Patterns are regular expressions evaluated in the guest — no shell quoting.
Replace is atomic per file, not per list: it fails at the first missing,
unreadable or non-UTF-8 path and the files before it stay rewritten, so pass it
paths a Find just returned. A Find fails as a whole when one match’s frame
would pass the 8 MiB cap.
w, err := sb.Watch(ctx, "/work", true)
defer w.Close()
for ev := range w.Events() { // wire.Event{Kind, Path}
fmt.Println(ev.Kind, ev.Path) // created|modified|deleted|renamed
}
err = w.Err() // why the stream ended; nil after Close
Watch returns once the guest acknowledges the watch is armed — events
caused after it returns are guaranteed captured. A bad path fails
synchronously; if the consumer falls too far behind, Err reports a terminal
overflow instead of the stream silently dropping events.
err = sb.GitClone(ctx, url, "/work/repo", "main", 0, token) // egress lane only; depth > 0 = shallow
st, err := sb.GitStatus(ctx, "/work/repo") // Branch, Ahead, Behind, Files[]; Truncated when the list was cut at ~1 MiB
err = sb.GitAdd(ctx, "/work/repo", "a.txt")
hash, err := sb.GitCommit(ctx, "/work/repo", "message", "Dev <dev@example.com>")
err = sb.GitPush(ctx, "/work/repo", token) // egress lane only
err = sb.GitPull(ctx, "/work/repo", token) // egress lane only
br, err := sb.GitBranches(ctx, "/work/repo") // Current + Branches
err = sb.GitCreateBranch(ctx, "/work/repo", "feature")
err = sb.GitCheckout(ctx, "/work/repo", "feature")
err = sb.GitDeleteBranch(ctx, "/work/repo", "feature")
Results are structured (porcelain v2 under the hood), never scraped stdout.
Auth tokens travel as an in-memory header, never touching guest disk. On the
no-network lane, clone/push/pull fail fast with a typed unimplemented
error pointing at Push.
pty, err := sb.OpenPty(ctx, wire.PtyOpen{Cols: 120, Rows: 40})
defer pty.Close()
pty.Write([]byte("make test\n"))
io.Copy(os.Stdout, pty) // EOF when the shell exits
code, ok := pty.ExitCode()
err = pty.Resize(ctx, 200, 50)
A PTY is a tracked guest process (pty.PID); closing the handle (or the
ctx) tears the shell down.
Verbs for operating a node — Drain, Uncordon, SetPools,
SetPoolsCluster and the tenant verbs need the root token — plus the reference
the aggregated apiserver claims under:
sb, _ := client.New(ctx, "rt:24.04", sandbox.WithClaimRef("ns/workload"))
list, _ := client.Sandboxes(ctx) // one SandboxSummary per live claim — never tokens
info, _ := client.Drain(ctx) // cordon: refuse new claims, run leases out
info, _ = client.Uncordon(ctx)
info, _ = client.SetPools(ctx, pools) // retune warm targets without a restart
res, _ := client.SetPoolsCluster(ctx, pools) // per-node results; retry the failures
sb = client.Attach(ownerAddr, id, token) // bind a known handle, no lookup round-trip
err = client.PutTenant(ctx, sandbox.TenantSpec{Name: "u:42", Token: tok, MaxClaims: 10})
res, _ = client.PutTenantCluster(ctx, sandbox.TenantSpec{Name: "u:42", MaxClaims: 20}) // token kept
res, _ = client.DeleteTenantCluster(ctx, "u:42")
tenants, _ := client.Tenants(ctx) // names, caps, live claims, removed tenants; never tokens
Sandboxes is scoped to the calling token, so a tenant sees only its own
claims; the fields are those of GET /v1/sandboxes,
including each claim’s Metadata, CPUCount, and MemTotalBytes.
Drain leaves live claims alone — poll Info until Claimed is zero.
The tenant verbs follow /v1/tenants: a
TenantSpec without Token keeps the stored token, and the *Cluster forms
return one NodeResult per node to retry (see
cluster). SandboxSummary.Tenant names each claim’s
tenant.
*sandbox.ExitError — non-zero exit from Exec (Code, Stderr)*wire.ErrorResp — a typed guest-side failure; Kind is one of
wire.KindBadRequest (including a spawn the guest cannot start: missing
binary, a cwd that is not a directory, no exec bit, a User the guest
cannot resolve), KindNotFound,
KindUnimplemented, KindInternal (import
github.com/cocoonstack/sandbox/protocol/wire)var e *wire.ErrorResp
if errors.As(err, &e) && e.Kind == wire.KindUnimplemented {
// no-network lane: fall back to sb.Push
}
Context cancellation is honored on every call that takes a ctx: canceling it
closes the underlying connection. Data-plane calls return ctx.Err()
directly; control-plane verbs return a *url.Error wrapping it, so test with
errors.Is, never ==. Close is the exception: it takes no ctx.