States, shutdown behavior, cloud-init first boot, data disks, performance tuning, and live status.
| State | Description |
|---|---|
creating |
DB placeholder written, disks being prepared; vm start and vm stop refuse it, and an ownerless one (its create or clone died) is reclaimed by vm reconcile-stale-create, gc or the daemon |
created |
Registered, hypervisor process not yet started |
running |
Hypervisor process alive, guest is up |
stopped |
Hypervisor process exited cleanly |
error |
Start, stop, or restore failed — recover with vm restore; a failed restore also quarantines the record, so vm start is refused until a later restore succeeds or vm rm deletes it. Not sticky once the VM has been up: the daemon and the next vm start converge a dead VMM to stopped / unexpected-exit, and a VMM that is still alive is repaired to running; only a failed first start leaves error in place |
stopped (stale) |
Rendered, never persisted: the record reads running but the VMM is gone; vm list -o json, vm status --event --format json and vm inspect report "state": "stopped" with "stale": true |
stop_timeout_seconds in config or --timeout flag) → vm.shutdown → SIGTERM → 5s → SIGKILL. Cloud Hypervisor delivers the button through its GED device on both boot paths; a direct-boot guest with systemd (the sandbox images) powers off in well under a second, one whose init ignores the button waits out the timeout; the Android images handle it through cocoon-power (see OS images)ssh shutdown /s /t 0 before stopping, or --force to skip the ACPI timeout (see known issues)SendCtrlAltDel → up to stop_timeout_seconds (30s) for the guest to halt → SIGTERM → 5s → SIGKILL--force): skip the ACPI window — Cloud Hypervisor still issues vm.shutdown to flush disks, then SIGTERM → 5s → SIGKILL; Firecracker goes straight to SIGTERMvm rm --force): same immediate path as force stop, then delete — no graceful window| Flag | Default | Description |
|---|---|---|
--force |
false |
Skip graceful ACPI shutdown, immediate kill |
--timeout |
0 (use config default) |
ACPI shutdown timeout in seconds |
Every VM’s hypervisor process is spawned directly into its own cgroup v2 scope (<cgroup_parent>/vm-<id>.scope, default parent cocoon.slice) via CLONE_INTO_CGROUP — vCPU threads, virtio queue workers, and io_uring kernel workers all land inside. The vCPU count alone does not bound host consumption (a 1-vCPU VM under I/O measures 111–113% of a core); the scope does.
Defaults are Kubernetes-style Guaranteed at N for --cpu N: quota = N cores (--cpu caps the long-run average, not just a topology hint), weight = N (proportional share under contention), burst = one period of quota. The burst credit refills from idle periods and never raises the long-run cap; it is there because the VMM’s virtio and io_uring workers bill to the same budget as the vCPUs, and without it a guest saturating its N vCPUs under I/O drains the budget mid-period and every vCPU freezes together until the period ends. Override any raw knob: lower --cpu-weight for burstable overcommit, raise --cpu-quota-us when sustained demand exceeds N (burst only absorbs demand fluctuating around the cap), --cpu-burst-us -1 for a strict no-burst cap. Two caveats at defaults: a saturated VM doing I/O pays its virtio service out of the N-core budget (~13% floor case), and cpu.weight only arbitrates real runqueue contention — a parent bandwidth limit is consumed first-come-first-served, not by weight.
cgroup_cpus fences the whole VM population onto a host cpu subset (e.g. 0-14 on a 16-core host keeps core 15 for the OS, the API consumer, and clone/wake execution). --cpuset-cpus pins one VM to specific cores inside the fence; --cpuset-cpus auto lets launch pick the least-loaded last-level-cache domain (one L3 complex, both SMT threads of every core) so the VM’s vCPUs, virtio workers, and io_uring threads share one cache; auto works on either backend, and only Cloud Hypervisor pins queue threads, so a Firecracker VM records the cpuset alone. Both are validated by cocoon against the effective sets — the kernel silently degrades ungrantable cpuset requests rather than failing — and shrinking the fence is refused while a running VM’s cpuset conflicts; queue pins outside a new fence are widened by the kernel until that VM’s next launch. Clearing cgroup_cpus converges: the stale fence is reset on the next launch.
Writable virtio-blk queue threads are pinned on every Cloud Hypervisor launch of a multi-vCPU VM, cpuset or not: each launch (start, clone, restore) counts what the live VMs across both backends already hold, their queue pins or, for a VM without pins, its cpuset, picks the least-loaded cache domain inside the fence, and gives every queue its own hardware thread there — distinct cores first, SMT siblings after, round-robin once a domain has fewer threads than the VM has queues — so a queue worker keeps its working set and its NVMe submission queue on one core; pair it with --cpuset-cpus auto to keep the vCPUs in that cache too. An explicit cpuset round-robins the queues over that cpuset instead; single-vCPU VMs and read-only disks are never pinned. cocoon vm inspect shows the result as cpuset and queue_cpus; the placement is re-derived on every launch, so a stopped or failed VM holds none.
Provisioning is exempt from the ceiling: clone/restore launch with the quota at max — weight, fence, and placement still apply — and the finite quota is armed after the memory load completes, before the guest resumes, so snapshot loading runs at control-plane speed while the guest never executes uncapped. The mmap mode’s deferred first-touch is guest-lifetime work whose cost keys on host cache state: cold cache faults from disk — I/O-bound, invisible to the quota; hot cache (sibling clones of one golden) turns guest writes into pure-CPU copy-on-write, fully exposed at wall = cpu_time / min(quota, busy threads). Time-to-first-exec is flat either way; only bulk re-touch of a large working set feels the ceiling, and only when oversold (quota under one core per busy thread). For such shapes express low priority with --cpu-weight instead of a tight quota, or use copy mode to prepay the load inside the exempt provisioning window.
cgroup knobs are host-side policy, like networking: snapshots record the source VM’s values but never apply them — a clone takes its policy from flags (defaults otherwise), restore keeps the target VM’s. Scopes are removed when the VMM dies (stop, hibernate, delete, crash convergence) and orphans are swept by cocoon gc. cocoon vm list shows per-VM throttling as THROTTLED (nr_throttled/throttled_usec from cpu.stat).
Requirements: cgroup v2 unified hierarchy with the cpu controller (kernel ≥ 5.14 for burst; on older kernels the defaulted burst degrades to none while an explicit --cpu-burst-us fails), running cocoon as root (production shape). Non-root works inside a systemd user slice with delegated controllers (systemd-run --user --scope), where user slices typically delegate cpu but not cpuset — fence/placement then fail preflight with the exact missing file named; queue pinning is thread affinity and needs no controller.
# Guaranteed at N (default): 2-core long-run cap, share 2, one period of burst credit
cocoon vm run --cpu 2 --memory 2G --name vm1 ghcr.io/cocoonstack/cocoon/ubuntu:24.04
# Strict cap, no burst: every period is hard-limited to 2 cores
cocoon vm run --cpu 2 --cpu-burst-us -1 --name strict1 ...
# Burstable overcommit: reach 2 cores when idle, shrink by weight under pressure
cocoon vm run --cpu 2 --cpu-weight 25 --name burst1 ...
# Headroom for virtio I/O service: guest keeps its full 2 cores under load
cocoon vm run --cpu 2 --cpu-quota-us 230000 --name io-heavy ...
# Metered: 0.5-core long-run average, bounded 1-core spikes
cocoon vm run --cpu 1 --cpu-quota-us 50000 --cpu-burst-us 50000 --name metered ...
# Pinning (NUMA / isolation-sensitive only — wastes idle cores)
cocoon vm run --cpu 2 --cpuset-cpus 2-3 --name pinned ...
# One cache domain, chosen at launch by load — the whole VM shares an L3
cocoon vm run --cpu 8 --cpuset-cpus auto --name llc ...
# Machine fence (config, not a flag): the fleet never touches core 15
COCOON_CGROUP_CPUS=0-14 cocoon vm run ...
# Clones never inherit snapshot policy — give it explicitly or get defaults
cocoon vm clone golden --name c1 # Guaranteed at N
cocoon vm clone golden --name c2 --cpu-weight 10 # explicit share
Rules of thumb: density → weight overcommit; single-VM performance → raised quota; billing semantics → quota + burst; pinning only when NUMA or isolation demands it. cocoon vm list’s THROTTLED column (count/total time from cpu.stat) shows who is hitting their cap.
The fence bounds the VMs; the caller’s own work (clone, restore, the API consumer) is deliberately not cocoon’s to manage — set it on the invoking service’s systemd unit, which writes the same cgroup v2 files:
# /etc/systemd/system/<your-service>.service.d/cpu.conf
[Service]
CPUWeight=1000 # management plane wins contention on the VM cores
With cgroup_cpus=0-14, the reserved core 15 has no VM competition and acts as the control plane’s fast lane without pinning; prefer that over AllowedCPUs=15, which would confine clone’s multi-core memory restore to one core. One-off commands: systemd-run --scope -p CPUWeight=1000 cocoon vm clone ....
vm create --hugepages; VM memory is backed by 2 MiB hugepages for reduced TLB pressure, and in exchange snapshots of that VM restore via eager copy only (the mmap fast path needs plain private-anon memory). Firecracker rejects --hugepages: FC cannot restore a hugetlbfs-backed snapshot, which would break hibernate/clone--mergeable at golden creation; guest memory is madvised MADV_MERGEABLE so host KSM can dedup identical pages across VMs — the flag persists through snapshot/clone/restore (it lives in the snapshot’s CH config, not the CLI), so build the golden with it or rebuild. cocoon only sets the madvise: enabling and tuning the scanner (/sys/kernel/mm/ksm/run, pages_to_scan) is the operator’s. Excludes --hugepages/--shared-memory (KSM merges only plain private pages); mmap-cloned siblings already share untouched pages via the page cache, so KSM’s gain is dirtied-but-equal and cross-golden pages — measure density on your fleet, and weigh ksmd CPU plus the cross-VM dedup timing side channel in multi-tenant setupsdirect=off), writable raw COW and data disks use O_DIRECT (direct=on) to avoid host cache buildup and guest flush storms, and qcow2 overlays stay buffered — Cloud Hypervisor applies the disk’s direct flag to the backing file too, and O_DIRECT there would give every VM its own read of the shared base instead of one page-cache copy--no-balloon opts a VM out at create, and clone/restore inherit it from the snapshot, which owns the restored device tree)--no-watchdog is an explicit compatibility opt-out for guests whose watchdog driver is unsafe during reboot. Firecracker exposes no watchdog device and rejects --fc --no-watchdog--cpu N boots N vCPUs with max set to the host’s online core count, so the guest sizes its ACPI CPU table for the host rather than for N — a cold-boot cost that grows with host sizeFirecracker attaches virtio devices over MMIO by default (and adds pci=off to the guest command line itself). --pci on vm run/vm create starts Firecracker with --enable-pci, so the guest enumerates virtio-pci devices instead; this is the prerequisite for Firecracker device hot-plug (Developer Preview upstream). The transport is fixed for the VM’s lifetime and travels with its snapshots: clones and restores relaunch on the transport the snapshot was taken with, and a snapshot cannot move between transports. The guest kernel needs CONFIG_PCI, CONFIG_VIRTIO_PCI and CONFIG_ACPI; distro kernels have both transports, while the sandbox guest kernel is built without CONFIG_VIRTIO_MMIO, so sandbox images can never use the MMIO transport (booting them on Firecracker at all is not validated yet). Cloud Hypervisor is always virtio-pci, so --pci is rejected there.
With --pci, vm net, vm disk attach/detach, clone --nics and clone --data-disk work as on Cloud Hypervisor, with one difference: Firecracker does not notify the guest, so cocoon prints the guest-side step after each change (echo 1 > /sys/bus/pci/rescan after a plug, echo 1 > /sys/.../device/../remove for the stale node after an unplug; --output json returns them as hints). Unmount or down a device inside the guest before detaching it. Hot-attached disks show up as /dev/vdX (Firecracker has no virtio-blk serial) and, as on Cloud Hypervisor, block snapshot and hibernate until detached.
Cloudimg VMs receive a NoCloud cidata disk (FAT12 with CIDATA volume label) containing:
#cloud-config with configurable user/password (--user/--password, defaults to root/cocoon)/etc/systemd/network/15-cocoon-id*.network files matching current MAC (MACAddress=), used when netplan PERM-MAC matching cannot applyThe cidata disk is automatically excluded on subsequent boots — after the first successful start, the VM record is marked as first_booted and the cidata disk is no longer attached, preventing cloud-init from re-running.
Note: --user/--password only apply to cloudimg VMs (cloud-init). OCI VM images bake credentials at build time — every official os-image/ubuntu/* image ships openssh-server enabled with PermitRootLogin yes and the default root:cocoon credentials. Host-to-guest control plane operations (kubectl exec, kubectl logs) prefer cocoon-agent over vsock; SSH stays available as the human-on-keyboard path.
--data-disk attaches additional virtio-blk disks beyond the rootfs/COW. Cocoon manages each disk’s lifecycle (sparse raw file under the VM’s runDir, optional ext4 mkfs at create time, full participation in snapshot/clone/restore), so the user only chooses size, optional fstype/mount, and DirectIO policy.
# OCI + CH: two disks, mounted manually inside the guest
cocoon vm run --data-disk size=20G,name=db --data-disk size=50G,name=cache <oci-image>
# Cloudimg + CH: cloud-init writes /etc/fstab from the spec, disks auto-mount on boot
cocoon vm run --data-disk size=20G,name=db,mount=/mnt/db <cloudimg>
# Unformatted disk, guest is responsible for mkfs and mount
cocoon vm run --data-disk size=20G,name=raw,fstype=none <oci-image>
--data-disk accepts comma-separated key=value pairs (repeatable):
| Key | Default | Notes |
|---|---|---|
size |
required | Minimum 16 MiB; goes through units.RAMInBytes so 512M, 2G, etc. all work |
name |
dataN (auto) |
1-20 chars, [a-z][a-z0-9_-]{0,19}, cocoon- prefix reserved; auto-numbered names skip any explicit one already taken |
fstype |
ext4 |
ext4 (cocoon mkfs’s it) or none (guest must format); xfs is not supported in Phase 1 |
mount |
/mnt/<name> |
Cloudimg+CH only — emitted as a cloud-init mounts: row using /dev/disk/by-id/virtio-<name>. Pass mount= (empty) to skip auto-mount even when fstype is ext4. fstype=none requires mount=/empty |
directio |
auto |
on forces direct=on, off forces page-cache, auto inherits VM-level --no-direct-io. CH only — FC has no DirectIO knob and logs a warn |
| Cloud Hypervisor | Firecracker | |
|---|---|---|
| Guest device naming | /dev/disk/by-id/virtio-<name> (stable) and /dev/vdX |
/dev/vdX only — FC has no virtio-blk serial field |
Cloud-init mounts: auto-mount |
Yes (cloudimg path) | N/A (FC has no cloudimg) |
| Per-disk DirectIO override | Yes | Ignored with warn |
| Snapshot/clone/restore | Yes — sidecar carries Role/MountPoint/FSType | Yes — sidecar in cocoon.json carries the same |
Phase 1 inherits data disks 1:1: snapshot reflinks each data-<name>.raw into the snapshot tar, clone re-creates them under the new VM’s runDir (and regenerates cidata so cloud-init re-mounts on the new identity), and restore rolls all data disks back to the snapshot timepoint along with the rootfs and memory state. Cloud Hypervisor clones can additionally CREATE fresh data disks at clone time via --data-disk (hot-added after restore — the snapshot’s device tree itself cannot grow); clone-created disks are hot-added after cidata is regenerated, so they are never auto-mounted — mount= has no effect on vm clone --data-disk; mount them inside the guest. Removing inherited disks at clone time is not supported, and Firecracker clones accept --data-disk only from a --pci snapshot (MMIO cannot hot-plug).
Restore preflight verifies sidecar integrity, file presence (vmstate, memory, COW, every data-*.raw — presence, not size; --from-dir and snapshot import separately check the COW size against the envelope), per-index Path/RO agreement between the sidecar and CH config.json, Role/Serial agreement between the sidecar and the VM record, and on Firecracker that every COW and data disk is recorded at this VM’s own path, all before killing the running VM, so a malformed or imported snapshot fails fast and leaves the live VM untouched.
cocoon vm status provides real-time VM state monitoring with two modes:
# One-shot snapshot (default)
cocoon vm status
# Refresh mode — clears and redraws like `watch`
cocoon vm status --watch
# Event stream mode — appends state changes (for scripting / vk-cocoon)
cocoon vm status --event
# Filter specific VMs, custom poll interval
cocoon vm status --event -n 2 my-vm other-vm
State changes are detected via fsnotify on the meta store’s directory (the sqlite database’s directory, or the json index file’s parent, since an atomic rename changes the file’s inode; sub-second latency), with a configurable poll interval as fallback. Event mode emits ADDED, MODIFIED, and DELETED lines suitable for machine consumption.