cocoon

Garbage Collection

Cross-module GC, snapshot LRU eviction, and scheduled cleanup.

How It Works

cocoon gc performs cross-module garbage collection:

  1. Recover the interrupted deletions of every module that has a recovery step (snapshot, both hypervisor backends, CNI), so discovery sees no half-removed resources
  2. Snapshot each module’s index (a loose read; every destructive decision is revalidated later under the module’s own entity locks and tombstone leases); a module whose index read fails aborts the whole cycle before anything is collected
  3. Resolve each module identifies unreferenced resources using the full snapshot set (e.g., image GC checks VM and snapshot records for blob references)
  4. Collect — recheck ownership and delete remaining targets; CNI checks both VM backends under the VM ops lock, bridge checks the current owner after locating each TAP, and the cgroup module checks it before removing an empty scope. The bridge and cgroup rechecks need no lock because create writes the VM record before it provisions a TAP or a scope, so a live device always has a record. A failed owner lookup preserves that resource, reports an error, and lets the sweep continue with the rest.

This ensures blobs referenced by running VMs or saved snapshots are never deleted. cocoon gc also terminates a VMM still bound to an orphaned run dir (SIGTERM, then SIGKILL) before removing the dir.

Log Output

Every collected item is logged at INFO level with a structured key=value payload in its message under func gc.<module>, and a summary line ends the cycle; the summary counts candidates identified, not deletions confirmed. When stderr is not a terminal (the systemd unit below, a pipe, a file), each record is one JSON line. Sample:

{"level":"info","func":"gc.snapshot","time":"2026-04-19T03:00:12Z","message":"collected id=XEOU... name=ubuntu-hot-testing:v1 bytes=3221225472 last_accessed=2026-04-12T10:30:00Z reason=lru-age"}
{"level":"info","func":"gc.snapshot","time":"2026-04-19T03:00:12Z","message":"collected id=2GQVEA... name= bytes=0 last_accessed=never reason=orphan"}
{"level":"info","func":"gc.cloud-hypervisor","time":"2026-04-19T03:00:12Z","message":"collected id=ABC123 reason=orphan-runDir"}
{"level":"info","func":"gc.oci","time":"2026-04-19T03:00:12Z","message":"collected blob=b40150c1c2717d... reason=unreferenced"}
{"level":"info","func":"gc.cni","time":"2026-04-19T03:00:12Z","message":"collected id=JKLMN netns=cocoon-JKLMN reason=orphan"}
{"level":"info","func":"gc.bridge","time":"2026-04-19T03:00:12Z","message":"collected id=MNOPQ iface=btMNOPQ-0 reason=orphan-tap"}
{"level":"info","func":"gc.Run","time":"2026-04-19T03:00:12Z","message":"completed: cloud-hypervisor=1 cni=1 oci=4 snapshot=3 (failures: 0, duration: 230ms)"}

Filter with awk / grep:

journalctl -u cocoon-gc.service --since today | grep '"func":"gc.snapshot".*reason=lru-'
journalctl -u cocoon-gc.service --since today | awk '/"func":"gc.Run"/'

Reasons:

Snapshot LRU Eviction

Bare cocoon gc only reclaims orphans (on-disk data with no DB record), missing-dir records (a record whose data dir is gone) and stale pending records (a save died mid-flight; every save holds its snapshot’s build lease from placeholder to finalize, so a pending record whose lease GC can acquire is provably ownerless — no age wait). To also evict healthy snapshots by access recency, pass --snapshot:

Flag Effect
--snapshot Enable LRU eviction. Bare flag = evict every non-pending snapshot.
--snapshot-keep N Keep at most N most-recently-accessed snapshots.
--snapshot-age DUR Evict snapshots last accessed before this duration (e.g. 720h for 30d).
--snapshot-size SZ Evict oldest snapshots until total size ≤ this (e.g. 100GiB; suffixes are binary, GB = GiB).
--snapshot-dry-run Log which snapshots would be LRU-evicted; act on nothing. Snapshot-only — orphans and other GC modules still execute.

Sub-flags combine as union of evictions (intersection of kept) — a snapshot is kept only if it passes every active criterion. All sub-flags require --snapshot; negative values are rejected.

LastAccessedAt is updated on Restore, vm clone (via DataDir) and snapshot export, set at finalize time by snapshot save and to the creation time by snapshot import, so a fresh snapshot is never age-evicted by the next sweep. Inspect and list do not count as access.

# Preview what 30-day eviction would remove (snapshot-only — other GC modules still run)
cocoon gc --snapshot --snapshot-age=720h --snapshot-dry-run

# Production: weekly cleanup, keep 50 newest within 7 days
cocoon gc --snapshot --snapshot-age=168h --snapshot-keep=50

# Cap storage at 100GB
cocoon gc --snapshot --snapshot-size=100GiB

# Nuke all snapshots (dev / test reset)
cocoon gc --snapshot

Scheduled Snapshot GC

cocoon gc is a one-shot, lock-safe operation — drive periodic execution from a systemd timer or cron. If you already run cocoon daemon --gc-interval, that sweep covers orphans and a timer is only needed for LRU eviction. Recommended template (systemd):

# /etc/systemd/system/cocoon-gc.service
[Unit]
Description=Cocoon snapshot GC (LRU eviction)

[Service]
Type=oneshot
ExecStart=/usr/local/bin/cocoon gc --snapshot --snapshot-age=168h --snapshot-keep=50
StandardOutput=journal
StandardError=journal
# /etc/systemd/system/cocoon-gc.timer
[Unit]
Description=Run cocoon snapshot GC daily

[Timer]
OnCalendar=daily
RandomizedDelaySec=1h
Persistent=true
Unit=cocoon-gc.service

[Install]
WantedBy=timers.target

Enable: systemctl enable --now cocoon-gc.timer.

For cron, drop a one-liner into /etc/cron.daily/cocoon-gc:

#!/bin/sh
exec /usr/local/bin/cocoon gc --snapshot --snapshot-age=168h --snapshot-keep=50