I would like to propose an optional, default-off offline integration that captures a minimal
UniRL rollout trace, converts it to an existing AISimulate / Dynamo Replay interface, and
compares the replay against the matched real execution.
The initial target is a fixed-policy window on one supported engine configuration. This is a
trace adapter and validation tool, not a new rollout engine, not a source of synthetic training
samples, and not a complete RL training simulator.
Motivation
A UniRL-specific adapter can preserve the relationship between backend requests, Sample
lineage, trajectories, GRPO groups and the actual rollout batch boundary. That gives a
reproducible workload for evaluating rollout timing models before attempting configuration
comparisons. Existing Replay functionality should be reused wherever it already supports
request dependencies and external waits; the contribution should be the UniRL mapping and
matched validation, not another scheduler or discrete-event simulator.
Initial scope
- Pin UniRL and a compatible simulator revision/API; verify a common backend/model/profile
configuration rather than assuming one.
- Export a versioned manifest and workload records, with measured outcomes stored separately.
- Start with observed-arrival, single-engine open-loop replay using actual token lengths and
an explicit cache state.
- Add CPU contract tests and a real upstream replay smoke test; validate on the target
hardware only when a matching profile and runtime configuration are available.
The simulator dependency would stay optional and isolated from normal training. A small
experimental module reusing existing capture hooks and artifact conventions is preferred.
Trace and result contract
The adapter preserves stable request identities, actual input/output lengths, fixed-policy
identity, engine/worker mapping, timing boundaries and configuration provenance. Measured
service times are never supplied to the predictor. Raw prompts, outputs and tool content are
excluded by default.
Request-level comparisons require request-level completion output. If the selected runner only
provides aggregate summaries, the initial result would be explicitly aggregate-only, and group
completion times would not be reconstructed from percentiles.
Agentic follow-up
For supported agentic traces, subsequent submission would depend on simulated predecessor
completion plus a separately measured external delay, never on observed next-turn arrivals.
Two boundaries need explicit validation:
- A trajectory slot may stay occupied while its thread waits on the environment; it is not an
in-flight model request slot.
- A GRPO group is not necessarily the collector barrier; group completion and rollout-batch
readiness must stay distinct.
The adapter would reuse upstream dependency/workload-driver facilities and reject unsupported
resource or lifecycle semantics. It would not insert fake model requests as barrier events or
add an artificial global barrier between independent engines.
Non-goals
No changes to sampling, training scheduling, placement, loss or weight synchronisation. No
backward/optimizer, offload, policy evolution, diffusion/video, full RL-step or online-control
model. Partial/async cross-policy resume is outside the initial scope.
Backend support would be the intersection of real UniRL execution and simulator/profile
coverage; this proposal does not include implementing another serving backend.
Validation
Tests cover ID mapping, multiple Sample rows/branches, timestamp units, measured-outcome
isolation, dependency ordering, zero versus missing waits, trajectory slot limits, batch
barriers and explicit unsupported cases.
Hardware evidence would separate calibration and holdout workloads, report capture overhead and
measurement variability, and compare only matched timing boundaries. Synthetic clock fixtures
would not be presented as calibrated predictions.
Status of the first slice
Prepared against the public interface only: the trace contract, the converters and their CPU
tests are written and pass (18/18 not-FAIL, 16 PASS / 2 NOT_RUN). No simulator package is
installed in the environment used, so the trace parser, the CLI, request-level result export,
trajectory-slot admission and the collector barrier are unvalidated and are marked as such.
Maintainer feedback requested
Would this optional tooling scope be useful, and is an overlapping adapter already in progress?
Which existing capture hook and directory should it use? Which backend/model/profile and pinned
Replay interface should be the first supported combination?
I would like to propose an optional, default-off offline integration that captures a minimal
UniRL rollout trace, converts it to an existing AISimulate / Dynamo Replay interface, and
compares the replay against the matched real execution.
The initial target is a fixed-policy window on one supported engine configuration. This is a
trace adapter and validation tool, not a new rollout engine, not a source of synthetic training
samples, and not a complete RL training simulator.
Motivation
A UniRL-specific adapter can preserve the relationship between backend requests,
Samplelineage, trajectories, GRPO groups and the actual rollout batch boundary. That gives a
reproducible workload for evaluating rollout timing models before attempting configuration
comparisons. Existing Replay functionality should be reused wherever it already supports
request dependencies and external waits; the contribution should be the UniRL mapping and
matched validation, not another scheduler or discrete-event simulator.
Initial scope
configuration rather than assuming one.
an explicit cache state.
hardware only when a matching profile and runtime configuration are available.
The simulator dependency would stay optional and isolated from normal training. A small
experimental module reusing existing capture hooks and artifact conventions is preferred.
Trace and result contract
The adapter preserves stable request identities, actual input/output lengths, fixed-policy
identity, engine/worker mapping, timing boundaries and configuration provenance. Measured
service times are never supplied to the predictor. Raw prompts, outputs and tool content are
excluded by default.
Request-level comparisons require request-level completion output. If the selected runner only
provides aggregate summaries, the initial result would be explicitly aggregate-only, and group
completion times would not be reconstructed from percentiles.
Agentic follow-up
For supported agentic traces, subsequent submission would depend on simulated predecessor
completion plus a separately measured external delay, never on observed next-turn arrivals.
Two boundaries need explicit validation:
in-flight model request slot.
readiness must stay distinct.
The adapter would reuse upstream dependency/workload-driver facilities and reject unsupported
resource or lifecycle semantics. It would not insert fake model requests as barrier events or
add an artificial global barrier between independent engines.
Non-goals
No changes to sampling, training scheduling, placement, loss or weight synchronisation. No
backward/optimizer, offload, policy evolution, diffusion/video, full RL-step or online-control
model. Partial/async cross-policy resume is outside the initial scope.
Backend support would be the intersection of real UniRL execution and simulator/profile
coverage; this proposal does not include implementing another serving backend.
Validation
Tests cover ID mapping, multiple
Samplerows/branches, timestamp units, measured-outcomeisolation, dependency ordering, zero versus missing waits, trajectory slot limits, batch
barriers and explicit unsupported cases.
Hardware evidence would separate calibration and holdout workloads, report capture overhead and
measurement variability, and compare only matched timing boundaries. Synthetic clock fixtures
would not be presented as calibrated predictions.
Status of the first slice
Prepared against the public interface only: the trace contract, the converters and their CPU
tests are written and pass (18/18 not-FAIL, 16 PASS / 2 NOT_RUN). No simulator package is
installed in the environment used, so the trace parser, the CLI, request-level result export,
trajectory-slot admission and the collector barrier are unvalidated and are marked as such.
Maintainer feedback requested
Would this optional tooling scope be useful, and is an overlapping adapter already in progress?
Which existing capture hook and directory should it use? Which backend/model/profile and pinned
Replay interface should be the first supported combination?