An execution layer for agentic RL rollouts and evaluation¶
Treat each rollout as an independent task: prepare a workspace, run the agent, collect its trajectory and file result, calculate reward, then retain or discard its changes. pVisor supplies execution and evidence; the trainer owns sampling, reward, and scheduling.
First validate result collection with an offline task:
pvisor run --safe --overlaynet-deny-all --stdio capture \
--stage ../rollout-001 -- /bin/sh -c 'printf "candidate\n" > answer.txt'
pvisor status --review --json ../rollout-001 > rollout-001.bundle.json
pvisor inspect ../rollout-001 -- cat answer.txt
Expect a completed task, exit code zero, and a staged answer.txt. Have the evaluator read the candidate in the Stage, then run pvisor drop ../rollout-001 after scoring. Training candidates do not need to be applied to the baseline.
How to assemble one rollout today¶
An evaluation sample today can be one command execution, its model traffic capture, and the native agent trajectory. Prepare a clean workspace for the sample, then run the command and keep the Job ID, Run Bundle, Event Journal, and native agent trajectory. Reward computation, the task queue, model training, and cross-node scheduling stay with your existing framework.
| Artifact | Purpose | Does not replace |
|---|---|---|
| Run Bundle | Outcome, controls, and file changes | Training-framework reward and dataset metadata |
| Gateway Journal | Model calls through the Gateway | Uncaptured traffic and native agent sessions |
| Native agent trajectory | Tool-prefix replay for the matching adapter | Process memory snapshots |
| Logical checkpoint | Forking staged file state | Full environment and external service state |
Pin the sample identity¶
Record the task ID, pVisor commit, agent/model versions, initial repository commit, image digest, sampling parameters, policy, and executor for every rollout. Track task failure, isolation refusal, recording failure, and continuation quality separately; never record an infrastructure failure as "the model was incapable".
Replay reruns tools in a new environment and repeats their side effects. Use test APIs/databases and a separate output directory; validate the format with --prepare-only, validate the tool prefix with --replay-only, and then start a real continuation. See replay for the pinned adapter versions and boundary semantics. Cluster throughput and hostile multi-tenancy are still unproven.
Score and retain a training sample¶
Use this minimal evaluator for the preceding offline task. It checks execution first, then candidate content. The example requires jq.
jq -e '.schema_version == 4 and .run.state == "completed" and .run.exit_code == 0' \
rollout-001.bundle.json
pvisor inspect ../rollout-001 -- cat answer.txt > rollout-001.answer.txt
reward=0
if [ "$(cat rollout-001.answer.txt)" = candidate ]; then
reward=1
fi
jq -n --arg task_id sample-001 --argjson reward "$reward" \
'{task_id: $task_id, reward: $reward, bundle: "rollout-001.bundle.json"}' \
> rollout-001.score.json
pvisor drop ../rollout-001
This reward checks content only; replace it with unit tests, environment scores, or human labels for real tasks. If execution checks or file reading fail, retain diagnostics as a failed sample and stop normal scoring. In CI, use set -e or explicitly check every return value.
Configure the agent's native trajectory output using Agent integrations; follow Model traffic capture for Gateway recording. Associate trajectories, Bundles, and scores with one task/attempt ID. For batches, reuse the parallel workspace workflow with a fresh Stage for every sample.