Skip to content

Which benchmark results support an adoption decision?

Main conclusions

Choose execution modes using measured workloads with matching budgets and timing boundaries; use different configurations to understand their individual performance levels. Startup, tool execution, application and resource occupancy answer different questions.

Decision Required evidence
Choose local tool execution Native, Docker and pVisor host/staged/VM task comparisons
Choose an independent guest kernel Firecracker, QEMU and pVisor startup plus tool data
Set capacity or timeouts Success rates, complete-task waiting and matching resource accounting

Motivation

Faster startup does not guarantee faster compilation. Writable mounts and staged review also provide different workflows. Adoption decisions require speed, execution boundaries and how changes reach the original directory.

Experiment design

Same-host controls

B-STARTUP, B-FS-TOOLS and B-AGENT-TASK use one offline tool environment and fixed inputs across native, pVisor host/staged/VM, private rootless Docker, Firecracker PCI and QEMU q35/microvm. Images, tools and daemon are prepared beforehand. Every execution uses the same two-core affinity; tool VMs have matching configured memory. Each job gets a fresh workspace, with seeded randomized runtime order. Complete Ubuntu, macOS and distinct binaries remain separate cohorts.

Downloads, benchmark compilation, image import and input copying are excluded. Timing separates first valid output, internal tools plus validation, returned results and launch-to-process-exit. Successful samples require validated outputs, observed executor and complete staging. Staged modes also verify that original host files remain unchanged.

Distributions and failures

Cohorts are not pooled, and slow samples are not removed after measurement. Failures and invalid outputs are counted separately and never treated as zero latency. Capacity guards are not completed jobs. P95 is descriptive: it is omitted below 30 samples. Public P99 is omitted; small cohorts do not support stable tail-latency claims.

Separated clusters are reported with counts and individual medians. The descriptive split rule is in the runner manual; it establishes no cause. Engineering A/B percentage claims need a 95% bootstrap interval for the median difference; user pages do not narrate optimization work.

Isolation and resources

Docker writable bind mounts write directly to the host; pVisor staging retains changes until apply. Firecracker/QEMU use private ext4, while pVisor VM uses virtio-fs with different kernels and devices. This is not a pure VMM or production-security ranking. Standalone staging boundaries follow actual records and isolation checks.

RSS sums sampled process scopes and may double-count shared pages or miss short peaks. Docker must include actual container processes and identify its dedicated daemon. Whole-group memory must include all descendants, backing and charged cache, without double-counting overlapping cgroups. Configured RAM, RSS, macOS RAM proxies and net physical memory are distinct metrics.

Data and analysis

Prepared environments

Startup · Files and tools · Repair tasks and CLIs. Tables retain cohort identities and counts, with pinned artifacts rather than claims about current third-party releases.

Complete distributions

Ubuntu uses its distribution kernel, initrd, systemd and private disks; pVisor reuses prepared tool directories. The comparison answers deployment waiting, rather than isolating VMM cost.

QEMU complete distributions

q35 and microvm use one Ubuntu template. Their samples and percentiles remain separate from other cohorts.

Tasks and resources

Application · Network · Concurrency · Isolation · Replay · Review. Unmeasured industry alternatives are marked explicitly, without marketing numbers filling gaps.

Filesystem engineering experiments

Engineering A/B, instrumented profiles and diagnostic probes remain in technical analysis, outside cross-product main tables.

Data location and downloads

Markdown holds derived user-facing tables, with downloadable CSVs beside each article. Raw reports, individual samples, logs, artifact manifests and frozen harnesses live in the relevant directory’s .data/, ignored by Git and excluded from the site. CSV downloads contain derived statistics and source summaries, rather than claiming to contain raw samples.

See the runner manual for reproduction and retention rules.

Downloads and reproduction

Derived table CSV · Evidence source summary · Comparison method · Runner manual