Repository navigation
v0.4+ candidate: live-fork via memfd-backed source RAM + uffd_wp (deferred from v0.3) #101
Description
Activity
Update 2026-05-19 — narrowing scope. v0.3 phase 1 shipped 143× on vanilla Firecracker (no fork required). The
firecracker-patch/directory that originally drafted aMemoryBackend::Memfdpatch has been removed (PR opened).Concretely, we're now explicitly NOT taking the memfd route, for four reasons:
- Phase 1 cleared ~85 % of the original target headroom (30 ms target → 205 ms achieved on 4 GiB SSD) without any Firecracker change.
- memfd's value-add isn't sharing —
mmap(MAP_PRIVATE)ofmemory.binacross N children already gives kernel-CoW fan-out for free. memfd's real value is letting uffd_wp track source writes cleanly, but that's only useful for the live-fork architecture indocs/design/userfaultfd.md, which has its own deferred open questions. - Fork maintenance is real cost: own musl-via-docker CI, rebase on every upstream tag (~quarterly), track CVEs, weaken the "vanilla Firecracker" trust story users rely on. CodeSandbox went this route and hasn't upstreamed.
- Cheaper alternatives close the remaining gap without touching Firecracker: phase 2 (NVMe + io_uring), phase 3 (pre-emptive background snapshot with reflink), phase 1d (per-sandbox shadow file lifting the first-BRANCH-only restriction).
The revival criteria from the original issue still hold. Specifically: if phase 2/3/1d ship and the remaining pause floor is still the user-visible bottleneck on a real workload, OR a specific external user commits to using a
deeplethe/firecrackerfork, OR the source-divergence sync mechanism gets a paper-grade end-to-end sketch, OR upstream Firecracker accepts external-memfd injection — at that point we re-derive the patch from scratch against the then-current Firecracker tag.crates/forkd-uffd/and theMemoryBackend::Userfaultenum stay in the repo as scaffolding. The page-fault handshake parser is reusable independent of the live-fork decision.Reasoning in
docs/design/userfaultfd.md§ "Why we won't fork Firecracker".- added a commit that references this issue
on May 19, 2026 Draft RFC + implementation plan now live: #156
8-week phased plan: PoC → integrate into
forkd-uffd→ pause-window bench → hardening (write-heavy / NUMA / pre-5.7 fallback) → launch.Target: BRANCH pause < 10 ms by removing the synchronous memory write entirely. Switch source RAM to memfd, arm
UFFDIO_WRITEPROTECT, async dirty-page copier on uffd handler.Comments / prior-art pointers especially welcome on the open questions in the doc — particularly the behavior of
UFFD_WPon memfd-backed VMAs underKVM_RUN.Surface RFC posted at #174 —
DESIGN-v0.4-USER-API.md.Pins the CLI / REST / SDK shape before implementation starts so we don't re-litigate
--livevs--mode livevs--wpmid-coding. Companion to the existing kernel-mechanism doc (DESIGN-v0.4.md), use-case inventory (DESIGN-v0.4-USE-CASES.md), and FC integration spike (DESIGN-v0.4-PHASE3-SPIKE.md).Five open questions tagged for review in the doc; the most consequential is fail-vs-silent-fallback when the kernel doesn't support
uffd_wpon memfd. Feedback welcome here or on the PR.Closing — v0.4 live BRANCH ships end-to-end across REST, CLI, SDKs, doctor, docs, and bench (Phase 7.1–7.5, PRs #204–#210). User-facing surface complete; release artifacts published.
What landed:
- Phase 7.1 — REST canonical
modefield on BRANCH (feat(controller): Phase 7.1 — canonicalmodefield on BRANCH #204) - Phase 7.2 — CLI
--live/--no-wait(feat(cli): Phase 7.2 —--liveand--no-waitflags onforkd snapshot#205) - Phase 7.3 — Python / TypeScript / MCP SDK
mode/wait/live_fork(feat(sdk): Phase 7.3 —mode/wait/live_forkacross Python, TS, MCP #206) - Phase 7.4 —
forkd doctoruffd_wp + memfd_create capability checks (feat(doctor,uffd): Phase 7.4 — uffd_wp + memfd_create capability checks #207) - Phase 7.5 — bench harness + RESULTS-v0.4.md on a clean source (bench(v0.4): Phase 7.5 — live BRANCH pause-window data on a clean source #210)
- docs sweep — README + API.md + CHANGELOG + DESIGN-v0.4.md (docs: align README + API.md + CHANGELOG with shipped v0.4 live-fork surface #208)
forkd fork --live-forkfor local-boot path (feat(cli):forkd fork --live-fork— memfd-backed children for local v0.4 path #211)
Bench (1.5 GiB python-numpy, Intel i7-12700, ext4 HDD):
mode pause p50 pause p90 RT p50 live 56 ms 64 ms 13.7 s (sync) / 69 ms (async) diff 202 ms 418 ms 13.5 s full 13 550 ms 14 268 ms 13.6 s 3.6× faster pause vs Diff; 200× faster RT with
wait: false. Full data + CSV inbench/live-fork-pause-window/RESULTS-v0.4.md.Outstanding (tracked separately):
- Vendored FC patch upstream proposal — firecracker-microvm/firecracker#5912
- Daemon-spawn CLI verb for
--live-fork— feat(cli): expose--live-forkonforkd fork/from-image/runfor spawn-time opt-in #209
v0.5 (diff snapshot chains, M2.1) is in flight on
main; new tracking issue to follow.- Phase 7.1 — REST canonical
- added 7 commits that reference this issue
on Aug 11, 2026
Status: Deferred from v0.3. The design and scaffolding are in the repo (see "What's already here" below) so the work can be picked up cleanly if/when the cost-benefit changes. The reason for deferral is in "Why deferred" below.
Goal
Cut BRANCH pause-window to ~30 ms regardless of source memory size, by replacing "pause source, write full memory.bin, resume" with "pause source, register WP on its memory, resume; children inherit a memfd view and the WP handler resolves source's post-fork writes lazily."
Motivating measurement:
bench/pause-window/RESULTS-v0.2.md. Today's pause is storage-bound: 163 ms (tmpfs) to 4.26 s (SATA SSD) for 513 MiB, scaling linearly with source memory. v0.2.5's prewarm (PR #100) flattens the cold/warm ratio but doesn't change the absolute floor.What's already here (scaffolding from v0.3 cycle)
docs/design/userfaultfd.mdcrates/forkd-uffd/socketpair(2); no event loop yetcrates/forkd-vmm/src/lib.rsMemoryBackend::Userfaultbail!s inrestore_many_withso no caller can rely on itfirecracker-patch/v0.3-memfd-backend.patchWhy deferred
Architecture isn't closed. memfd injection is one piece; the harder piece is what happens to source's post-fork writes (its private CoW pages diverge from the memfd, so children see boot state, not source's current state). The prior art that's most often cited — MITOSIS, NFork, CodeSandbox — each solves a different problem (cross-host RDMA fork, kernel page-table sharing, cold-snapshot startup). None of them is a drop-in for "instant BRANCH from a long-running source."
Firecracker fork is real maintenance. Forking and patching is ~1 week initial + permanent: keep a
deeplethe/firecrackerrepo, rebase the patch on every upstream tag, run our own musl-via-docker CI, publish per-version releases, weaken the "uses official firecracker" trust story.Cheaper alternatives haven't been exhausted. Things that don't require a firecracker patch and could land in v0.3 instead:
enable_diff_snapshots: true+track_dirty_pages. The second-and-after BRANCH from the same source only writes pages dirtied since the last snapshot. Typically 5-10x speedup for repeated fan-out. ETA 3-5 days. Likely the biggest single win available.Together these get us most of the perceived win (sub-second BRANCH on commodity SSD) without any firecracker work.
Revival criteria
This stops being deferred when at least two of:
What gets touched if revived
crates/forkd-uffd/— add the UFFDIO_REGISTER / UFFDIO_COPY / UFFDIO_WP event loop on top of today's handshake parser.firecracker-patch/— refresh the patch against the current firecracker tag, get it compile-tested, fork the upstream repo.crates/forkd-vmm/src/lib.rs— wireMemoryBackend::Userfaultto actually spawn the handler, create the memfd, send it across the UDS instead ofbail!.docs/ROADMAP.md,docs/design/userfaultfd.md— update from "deferred" back to "in flight."Related work shipped in the v0.3 cycle anyway
MemoryBackendenum + scaffolding (no behavior change, but the API shape is stable for future use).