How many tool tasks fit in two cores and 2 GiB?¶
Main conclusions¶
Under a shared two-core, 2 GiB, zero-swap budget, stage and Podman complete all five rounds at 32 concurrent Python/Git tasks. pVisor VM completes all five at 16, but only three at 32. Idle stage probes reach 128 and Podman reaches 64; idle occupancy does not establish active-task capacity.
| Need | Selection implication |
|---|---|
| Parallel edits and Git checks | Stage and Podman reach the same tested task capacity; choose by isolation and review needs |
| Many idle environments | Stage reaches higher tested concurrency, using a host-process boundary |
| A VM boundary per task | Reserve more resources; do not size active tasks from idle counts |
Motivation¶
When multiple Agents execute together, completed work matters more than created environments. Capacity planning must include runtimes, helpers, page cache and VM backing; neither single-process RSS nor configured RAM can establish capacity.
Experiment design¶
Compare native processes, pVisor stage, pVisor VM and Podman 5.8.7 on Linux/x86_64. Each batch has a fresh cgroup with a shared two-core quota, CPU 0/1 affinity, a 2 GiB memory limit and zero swap. The coordinator, payloads and helpers are inside the budget; Podman payload/conmon membership is checked. Each VM has 2 vCPU/256 MiB, still constrained by the common total budget.
Scan concurrency 1, 2, 4, 8, 16, 32, 64 and 128, with five fresh batches per cell and no warmups. Conditions are randomized within each round. All tasks wait at a stdin readiness barrier before release. Idle tasks use the same Python environment without private data or edits. Useful tasks touch and checksum all 32 MiB of private data, edit four files in a 64-file Git repository, run git status and verify all contents. Stage/VM preserve the original workspace and match results to independent Run records.
The primary memory metric is whole-cgroup physical accounting, including charged page cache, kernel memory and private backing. Prepared shared tool/image caches may be charged to a parent cgroup, so this does not rank net whole-machine memory. Failed, unknown and OOM outcomes remain counted; only entirely verified, no-OOM batches enter memory/timing summaries. All slow valid samples remain. Five rounds do not establish long-term reliability or tail latency. This workload excludes inference, large builds and real Agent CLIs.
Data and analysis¶
Measured on 2026-10-06 local time. There are 320 batches: 263 recorded as fully passed and 57 failed; all contribute to capacity statistics. Memory is readiness-barrier memory.current P50 in MiB, from complete no-OOM batches only. “—” denotes no such batch.
Python/Git tool tasks¶
| Mode | Concurrency | Fully valid batches | Verified tasks / attempted | Unknown results | OOM batches | Barrier memory P50, MiB |
|---|---|---|---|---|---|---|
| Native process | 16 | 5/5 | 80/80 | 0 | 0 | 640.4 |
| Native process | 32 | 5/5 | 160/160 | 0 | 0 | 1268.3 |
| Native process | 64 | 0/5 | 233/320 | 0 | 5 | — |
| Native process | 128 | 0/5 | 244/640 | 0 | 5 | — |
| pVisor stage | 16 | 5/5 | 80/80 | 0 | 0 | 723.0 |
| pVisor stage | 32 | 5/5 | 160/160 | 0 | 0 | 1434.0 |
| pVisor stage | 64 | 0/5 | 204/320 | 0 | 5 | — |
| pVisor stage | 128 | 0/5 | 183/640 | 0 | 5 | — |
| pVisor VM | 16 | 5/5 | 80/80 | 0 | 0 | 2047.9 |
| pVisor VM | 32 | 3/5 | 138/160 | 0 | 0 | 2047.8 |
| pVisor VM | 64 | 0/5 | 1/320 | 0 | 0 | — |
| pVisor VM | 128 | 0/5 | 0/640 | 0 | 5 | — |
| Podman | 16 | 5/5 | 80/80 | 0 | 0 | 969.3 |
| Podman | 32 | 5/5 | 160/160 | 0 | 0 | 1931.8 |
| Podman | 64 | 0/5 | 120/320 | 64 | 5 | — |
| Podman | 128 | 0/5 | 128/640 | 128 | 5 | — |
Stage and Podman each complete 160/160 tasks at concurrency 32. Both encounter OOM at 64/128; partial completions do not establish those capacities. VM completes 138/160 tasks at 32 with only 3/5 complete batches. Launch/completion failures at higher concurrency remain visible and cannot all be attributed to OOM. Unknown results lack retained verifiable completion evidence; they are neither successes nor zero-duration tasks.
Idle environments¶
| Mode | Concurrency | Fully valid batches | Verified tasks / attempted | Unknown results | OOM batches | Barrier memory P50, MiB |
|---|---|---|---|---|---|---|
| Native process | 32 | 5/5 | 160/160 | 0 | 0 | 242.0 |
| Native process | 64 | 5/5 | 320/320 | 0 | 0 | 471.9 |
| Native process | 128 | 5/5 | 640/640 | 0 | 0 | 931.9 |
| pVisor stage | 32 | 5/5 | 160/160 | 0 | 0 | 407.8 |
| pVisor stage | 64 | 5/5 | 320/320 | 0 | 0 | 804.5 |
| pVisor stage | 128 | 5/5 | 640/640 | 0 | 0 | 1600.4 |
| pVisor VM | 32 | 5/5 | 160/160 | 0 | 0 | 2047.9 |
| pVisor VM | 64 | 0/5 | 7/320 | 0 | 0 | — |
| pVisor VM | 128 | 0/5 | 0/640 | 0 | 5 | — |
| Podman | 32 | 5/5 | 160/160 | 0 | 0 | 902.8 |
| Podman | 64 | 5/5 | 320/320 | 0 | 0 | 1806.9 |
| Podman | 128 | 0/5 | 141/640 | 384 | 5 | — |
Stage passes five rounds at idle concurrency 128, but adding private data and Git work reduces its highest wholly successful tested level to 32. Plan the two workloads separately; discrete levels also do not establish an exact maximum.
Compressed parked snapshots versus container pause have not completed validation. Snapshot file size or single-VM reclamation cannot establish capacity. Firecracker/QEMU density under this budget is unmeasured; their single-task latency is in the end-to-end tasks. Matching macOS capacity is unmeasured.
Downloads and reproduction¶
Complete concurrency scan CSV · Artifacts, budget and evidence summary · Comparison method · Runner manual
Derived tables retain attempts, complete batches, failure reasons, unknowns, OOM, observed ranges and source digests for every condition. Raw reports, logs, input manifests and source/binaries stay in local .data/; failed batches remain in capacity denominators.