cocoon

Snapshots, Clone & Restore

Capture running VMs; clone them into new identities; restore or hibernate in place; move snapshots between hosts.

Overview

Cocoon supports snapshotting a running VM and cloning it into one or more new VMs.

Workflow

# 1. Snapshot a running VM
cocoon snapshot save --name my-snap my-vm

# 2. List snapshots
cocoon snapshot list

# 3. Clone a new VM from the snapshot
cocoon vm clone my-snap

# 4. Clone with a custom name
cocoon vm clone --name fresh-clone my-snap

# 5. Delete a snapshot
cocoon snapshot rm my-snap

What Gets Captured

A snapshot contains the full VM state:

snapshot save and vm hibernate are refused while a vhost-user-fs share, a VFIO device or a hot-attached disk is present; see Devices.

On Cloud Hypervisor, snapshot save and vm hibernate first let a running guest finish inflating its balloon, waiting up to a minute and moving on once inflation stalls for 2 s, so clones restore with a settled balloon instead of inflating it again at a high CPU cost.

Clone Constraints

CPU, memory, and storage are fixed at snapshot time on both backends: the guest is reconstructed from the snapshot’s binary device state, so vm clone and vm restore do not accept --cpu, --memory or --storage at all. NIC count inherits by default; Cloud Hypervisor clones can override it via --nics N (cocoon hot-swaps the snapshot’s NICs for a fresh set right after restore). Firecracker clones inherit the virtio transport with the snapshot: an MMIO clone must keep the snapshot’s NIC topology (network_overrides only retargets existing interfaces) and rejects --data-disk, while a --pci clone restores the snapshot’s NICs, then hot-plugs the --nics delta and any --data-disk after restore and prints the guest-side rescan. On Firecracker, the target network’s MTU must equal the snapshot’s, because the guest keeps the MTU the snapshot advertised (snapshots taken before nic_mtus was recorded skip this check); Cloud Hypervisor clones swap in fresh NICs after restore. Fresh data disks can be added to a Cloud Hypervisor clone via --data-disk (hot-added after restore). Create a fresh VM with cocoon vm run if a different CPU/memory/storage shape is needed.

Memory Restore Modes & Concurrency

Cloud Hypervisor clones and restores load guest memory via --restore-mode:

Firecracker restore is always memory-mapped; --restore-mode is accepted but ignored there.

Concurrent clones are first-class: sibling clones of one snapshot run fully in parallel, and each clone holds a shared lease on its snapshot for the duration of setup, so a concurrent snapshot rm fails fast with snapshot <id> is in use by an active clone/restore/export and a GC sweep skips it until its next cycle, instead of destroying work in flight; Firecracker clones also hold a shared per-VM lease on a managed source, so vm rm/vm restore of the source waits too. A source that was already deleted still clones (its drives travel inside the snapshot).

Post-Clone Guest Setup

After cloning, the guest resumes with new NICs — Cloud Hypervisor clones hot-swap in fresh MAC addresses automatically, while a Firecracker clone keeps the source VM’s MACs in the restored vmstate and cocoon vm clone prints the ip link set dev ethN down / address <MAC> / up triple to run first — but the guest OS still has the old IP configuration. You must reconfigure networking inside the guest: cocoon vm clone prints the exact steps for that VM — a --no-balloon clone has no balloon to release, so it gets no drop_caches line.

Cloudimg VMs (cloud-init re-initialization):

# Release balloon memory (the snapshot's memory pages are still cached)
echo 3 > /proc/sys/vm/drop_caches

# Clean old network configs from snapshot and reconfigure via cloud-init
rm -f /etc/systemd/network/10-*.network
cloud-init clean --logs --seed --configs network && cloud-init init --local && cloud-init init
cloud-init modules --mode=config && systemctl restart systemd-networkd

OCI VMs (MAC-based systemd-networkd reconfiguration — the new values are printed by cocoon vm clone):

# Release balloon memory
echo 3 > /proc/sys/vm/drop_caches

# Clean old network configs from snapshot, then set the hostname
rm -f /etc/systemd/network/10-*.network
hostnamectl set-hostname <VM_NAME>

# Write new MAC-based configs
# (cocoon vm clone prints a ready-to-paste loop with actual MAC/IP/GW values)
macs=('<MAC0>' '<MAC1>')
addrs=('<NEW_IP0>/<PREFIX>' '<NEW_IP1>/<PREFIX>')
gws=('<GATEWAY0>' '<GATEWAY1>')
for i in "${!macs[@]}"; do
  f="/etc/systemd/network/10-${macs[$i]//:/}.network"
  printf '[Match]\nMACAddress=%s\n\n[Network]\nAddress=%s\n' "${macs[$i]}" "${addrs[$i]}" > "$f"
  [ -n "${gws[$i]}" ] && printf 'Gateway=%s\n' "${gws[$i]}" >> "$f"
done
systemctl restart systemd-networkd

The cocoon vm clone command prints these hints with the actual values after a successful clone; the gws lines appear only when a NIC has a gateway, DHCP-only NICs (a --bridge clone) get a separate block writing DHCP=ipv4 units, and a Windows guest gets a Get-PnpDevice rebind plus netsh interface ipv4 set address instead.

Export & Import

Snapshots can be exported to portable tar archives (gzip with --gzip) for transfer between hosts or clusters, and imported back:

# Export a snapshot to a file
cocoon snapshot export my-snap --gzip -o my-snap.tar.gz

# Import on another host
cocoon snapshot import my-snap.tar.gz --name imported-snap

# Clone from the imported snapshot
cocoon vm clone imported-snap

# Or pipe directly between hosts (no intermediate file)
cocoon snapshot export my-snap -o - | ssh host2 cocoon snapshot import --name my-snap

The archive contains the snapshot config, VM config, COW disk, memory ranges, and device state — sparse files carry cocoon’s own pax records so their holes survive the round-trip — everything needed to reconstruct the snapshot on a different machine. The envelope also lists the files the archive carries, and snapshot import refuses an archive that ends before every listed file has landed — a plain tar cut at an entry boundary otherwise reads as complete. snapshot export --to-dir writes the same files as a directory that vm clone --from-dir and vm restore --from-dir consume without a tar round-trip. Do not repack an export with a third-party tar: dropping cocoon’s sparse records can rebuild a short, shifted disk. --from-dir and snapshot import reject a cow.raw whose size differs from the envelope’s recorded storage size.

Cross-Node Clone

When cloning an imported snapshot on a node that does not have the original base image, use --pull to auto-pull it:

# On the target node: import snapshot and clone with auto-pull
cocoon snapshot import my-snap.tar.gz --name my-snap
cocoon vm clone --pull my-snap

The --pull flag uses the image digest recorded at snapshot time: OCI images are pulled by digest, so a moved tag cannot be silently substituted; for cloud images a digest mismatch is logged and the clone then fails.

Note: --pull only works for registry-pulled images (OCI and cloudimg). For imported images (local qcow2/tar files), the base image must be transferred manually to the target node before cloning.

Restore

Restore reverts a VM — running or stopped — to a previous snapshot’s state in-place:

# Restore a VM to a previous snapshot
cocoon vm restore my-vm my-snap

A restore from a local snapshot store stops the VM and populates its run dir in place, quarantining the run dir if the population fails part-way so the VM cannot boot mixed-vintage files; a streamed restore stages the snapshot into a scratch directory first and (re)starts the hypervisor only after the full extraction succeeds, so a truncated or corrupt stream errors out with the VM in its prior state. Network is fully preserved — same IP, same MAC, same network namespace; restore always runs the same network self-heal as vm start, so a hibernated VM resumes even after a host reboot. No guest-side reconfiguration is needed (unlike clone).

Hibernate

Hibernate atomically snapshots a running VM and stops it, releasing its memory; resume with vm restore:

cocoon vm hibernate my-vm --name nap
cocoon vm restore my-vm nap

Pause, capture, persist, and VMM termination share one pause window: the snapshot point and the stop coincide (nothing the guest does can be lost in between), and the VMM dies only after the snapshot is durably saved — if saving fails (disk full, snapshot DB error), the VM simply resumes running and the command fails. Restore of the hibernated VM preserves machine identity (entropy-only reseed): running processes, sessions, and tmpfs contents continue where they stopped.

Restore Constraints