aorta

Turning on nightly perf gating for TokenSpeed serving

tokenspeed_serve reports TTFT, TPOT, ITL and throughput, and the serving metric names are already in the nightly’s gating allowlist. A serving recipe is in config/ci/nightly_eval_matrix.yaml, and this document is the sequence that took it from record-only to gated, plus the reasoning behind the numbers it picks.

[!IMPORTANT] Current state: two metrics are gated, nine are not. config/ci/regression_baselines.yaml carries median_tpot_ms and p99_itl_ms as max ceilings on both tokenspeed_serve_smoke cells (baseline and no-scratch-reclaim), derived from the ten-night window 2026-09-08..09-17. A run over either fails the nightly, and the breaching observation is still recorded and charted like any other night’s, so the trend reads straight through the failure.

Nine auto-gateable metrics stay record-only and still only charted: median_ttft_ms, p99_ttft_ms, p99_tpot_ms, median_e2el_ms, p99_e2el_ms, output_throughput, request_throughput, total_token_throughput and tokens_per_sec. So does step_time_ms.max, which was pruned deliberately rather than overlooked: it is what --perf-gate always writes, it is not a recorded metric at all — no per-night series accumulates for it, it is synthesised at bless time from one night’s mean — and the measured step-0 compile excursion (2825 ms) clears any ceiling a clean night gives it. median_itl_ms is absent for a different reason again: _NO_AUTO_GATE keeps the refresher from ever writing it, because it measures ~0 here.

Steps 1–6 below are done. Step 7 — promoting the nine on evidence — is not, and the nine are not waiting on the same thing. None of them is waiting on nights that have not happened yet — each group starts from evidence the completed window already holds, and a second window is a possible outcome for the three p99_* metrics only if analysing the first shows it is needed:

record-only metric what step 7 is waiting for
median_ttft_ms, output_throughput Nothing. The evidence is in hand. Their only blocker was the step-0 compile excursion, and the completed 2026-09-08..09-17 window clears it (step 4). They are deferred to a separate PR so the first armed gate stays attributable, not deferred for data.
p99_ttft_ms, p99_tpot_ms, p99_e2el_ms Analysis of a window, and possibly a second one. The blocker is that a p99 over 32 requests is the 32nd of 32 order statistics with no repeat measurement behind it. The completed window recorded them — the nightly harvests every allowlisted metric — but step 3 only asked for five metrics to be written down, so nobody has computed their spread. Read the existing ten nights first; another window is needed only if that spread is too wide to size a margin from.
median_e2el_ms, request_throughput, total_token_throughput, tokens_per_sec No window will unblock these. They are held back on redundancy, not on evidence: each is determined by a metric already gated, so arming them adds reason lines to one event rather than detection. Promoting them needs the argument in the per-metric table to change, not more nights.

Nothing else in the matrix changed: every other entry is unchanged and still correctness-only.

The reason it was a document rather than a commit is that there was no window to derive thresholds from. A threshold derived from a single observation encodes whichever night it was taken on, and the nightly then fails on the difference between two healthy runs. That failure is worse than no gate: it trains everyone to ignore the alert, and the first real regression arrives into a channel nobody reads. The window has since been taken — 2026-09-08..09-17, ten nights, all twenty cell-runs — and the two bounds it sized are in config/ci/regression_baselines.yaml. The argument is kept because it is the one step 7 still has to satisfy for the nine record-only metrics, and because it is the procedure for the next workload’s first bless.

Measuring the cell rather than reasoning about it has since made that case stronger and the first bless smaller. One cell-run in thirteen carries a five-second compile excursion in its first measured step, which three of the four originally-proposed gates cannot survive — see the smoke cell section.

See ci-nightly-eval.md for how the nightly works in general; this only covers what is specific to serving.

Record-only is the absence of a baseline, not a matrix field

Worth stating plainly, because the phrase suggests a setting somewhere and there isn’t one. nightly_eval.py looks each (entry, cell) up in config/ci/regression_baselines.yaml by the key entry::cell; a cell with no key gets compare_to_baseline(harvested, None), which returns record. So an entry added to the matrix is record-only by construction, and stops being record-only the moment a refresh-baselines PR merges a key for it.

Two consequences that decide the shape of this rollout:

What we actually know about the variance

Every number below is measured, on one gfx950 (MI355X) node against the image digest the recipes pin. There is no synthetic data here, and no more of it than this — that is the whole problem.

Two caveats that apply to every number in this section, both established later in the document rather than assumed here. They were taken at warmup_steps: 1 and the recipe now sets 2, so they are rationale and not a baseline; and they were taken on an MI355X, whereas the CI runner smci350-rck-g03-f16-12 is an MI350X, so their absolute level does not transfer to the runner even though their spread is the thing being argued about. See Verification status.

Bring-up is the noisy one, and it is very noisy. From tokenspeed.md:

Startup to /health is the dominant cost and it is noisy: 189, 276, 285, 291, 316 and 319 seconds across six runs of the same recipe on the same node — a 1.7× spread with nothing changed between them. So do not treat a slow bring-up as a signal without repeats, and do not tune timeout_per_trial close to an observed number; the recipe’s 1800 s leaves deliberate headroom.

The multi-model sweep adds a seventh observation for the same 0.6B model at 379 s, and tokenspeed-serving.md declines to treat the column as a measurement at all:

Read the startup column as a floor, not a measurement: it is dominated by weight loading and Triton compilation against whatever the node’s caches already hold, which is why the smallest model here posts the largest number. It is reported because ready_timeout_sec has to cover it, not because it scales with anything.

The four-model table is the evidence for that last clause: 379, 283, 289 and 328 seconds for Qwen3 at 0.6B, 1.7B, 4B and 8B. Bring-up does not even order by model size.

The compile cache is a second, larger excursion, and one mitigation stands between it and the metrics:

On Qwen3-0.6B the first bench invocation against a fresh server took 6.2s against 1.1s for every later one. Rolled into the metrics that one outlier dominates the mean step time and inflates TTFT tenfold (465ms vs 47ms), so a cell looks like a regression purely because it went first.

warmup_steps: 1 is supposed to discard that step. It does not always, and that is the single most important number in this document — see the next section, which measures it on the cell the staged entry actually runs. That measurement is why the recipe now sets warmup_steps: 2.

Steady-state serving metrics, by contrast, reproduce well. The docs do not say this in one sentence, but they contain two independent repeats:

Metric Qwen3-0.6B, two sweeps spread gpt-oss-20b, two sweeps spread
median_ttft_ms 46.3 / 45.9 0.87% 67.1 / 67.2 0.15%
median_tpot_ms 1.94 / 1.91 1.57% 7.61 / 7.63 0.26%
output_throughput 3538 / 3631 2.63% 994 / 991 0.30%
server_startup_sec 379 / — — 316 / 199 59%

The same node reports serving rates to within 3% across sweeps and bring-up time to within 59%. Those two facts are what the per-metric table below is derived from.

There is also a useful contrast from the kernel side of the integration, where the repo already distinguishes a tight metric from a loose one: the GEMM probe’s “spread across trials was under 1.5 µs in every cell”. Not everything TokenSpeed measures is noisy; bring-up is.

The smoke cell is not unconditionally clean (measured, 13 cell-runs)

Everything above is a claim quoted from another document. This section is a measurement of the exact cell the staged nightly entry runs, taken from the step_times_ms and metrics_summary recorded in every matrix.json on disk for TOKENSPEED-SERVE-SMOKE: six sweeps, 13 cell-runs with step times, 39 steps.

These numbers were all taken at warmup_steps: 1, which the recipe no longer uses. The recipe now sets warmup_steps: 2, precisely because of what this section measures. So read everything below as the rationale for that change, not as the baseline to bless against: the table is the evidence that a step-0 excursion reaches the metrics at warmup_steps: 1, and it is retained for that purpose. It was not a prediction of what the record-only window would show, because the window was taken at a different setting — one whose whole purpose is to remove the excursion these thirteen cell-runs contain. Nothing here is deleted or restated at the new setting. The measurement this was waiting for is the ten-night window at warmup_steps: 2, 2026-09-08..09-17, and it lives in step 4 of the rollout sequence rather than here.

Twelve of the thirteen are very clean. Tighter than the cross-sweep numbers above, because these are the same recipe on the same node:

Metric Clean range over 12 cell-runs Spread
median_ttft_ms 43.65 – 46.87 7.37%
median_tpot_ms 1.89 – 1.95 3.08%
output_throughput 3502.53 – 3646.20 4.10%
p99_itl_ms 34.00 – 35.84 5.43%
step time 1108 – 1178 ms 6.3%

The thirteenth is the compile-cache excursion, and warmup_steps: 1 did not catch it. serve-smoke3::no-scratch-reclaim recorded its three measured steps, in order, as:

6193.4 ms,  1140.2 ms,  1142.6 ms

The excursion is the first measured step — after the warmup step was already discarded. Its cell then reported median_ttft_ms 465.30 against a clean 43.65– 46.87, which is 10.26× the clean mean and the same 465 vs 47 the serving doc quotes. So the 10× TTFT excursion is not a hypothetical a recipe edit could introduce; it is present, once, in the data we already have, at 1 cell-run in 13 (7.7%).

What it does to the four gates the table below proposes, anchoring each on the worst of the twelve clean runs and applying the plan’s own margins:

Gate Anchor Threshold Excursion Verdict
step_time_ms.max 1169.50 1461.88 2825.40 fail
median_ttft_ms 46.87 58.59 465.30 fail
output_throughput 3502.53 2977.15 2612.83 fail
median_tpot_ms 1.95 2.44 1.93 pass
p99_itl_ms 35.84 44.80 34.96 pass

Three of the four proposed gates fire on a run that is not a regression. Those verdicts are compare_to_baseline’s, not arithmetic done here — tests/ci/test_eval_lib.py runs the measured numbers through the real comparator.

Why the latency-per-token metrics survive it, and why that is the useful signal. The excursion is a fixed ~5 s of compilation added to one step. It therefore lands on anything derived from step duration — the step-time mean, and throughput, which is tokens over that duration — and on TTFT, because the first request of that step waits for the compile. It does not touch median_tpot_ms (1.01× clean) or p99_itl_ms (1.00× clean), because those are measured between tokens, after the compile has happened. A metric whose definition excludes the excursion is immune to it by construction, not by luck.

This is a different failure from the concurrency-64 one, and the difference decides the fix. The matrix work found ~8% of bench steps at concurrency 64 stalling by a fixed ~0.92 s with no positional pattern — the serve-load::conc-64 cell on disk reads 1621, 2286, 2595 ms, bimodal in a way no warmup setting can remove. The smoke cell’s excursion is at position 0 every time it appears, which makes it deterministically avoidable: it is cache warming, and one more discarded step removes it. The two look alike in a summary statistic (both make a three-step mean untrustworthy) and are opposite in what to do about them.

Finally, the mitigation A/B is consistent with all of the above. On gpt-oss the hsa_no_scratch_reclaim cell landed at TTFT 66.7 vs 67.1 baseline, TPOT 7.49 vs 7.61, throughput 1008 vs 994 — the doc calls it “within noise of baseline, as it was at every smaller size”, and the size of that “noise” is the 0.3–2.6% band above, not the 59% one.

Per-metric: gate, record-only, or never

Direction is fixed by _METRIC_POLICIES in scripts/ci/eval_lib.py (max for latencies, min for throughputs). What this table decides is which of them get a bound on the first bless.

Only that. The verdicts below are about the nightly’s blessed baselines, which are derived by applying a margin to an observed value, and they say nothing about the workload’s own gates: block (_GATE_SPECS in src/aorta/workloads/tokenspeed_serve.py), which enforces absolute numbers a recipe writes out per trial. The two differ exactly where the margin does the damage: max_p99_itl_ms: 50 is a stated ceiling on tail stalls and is a perfectly good gate, whereas a baseline ITL ceiling is observed × 1.25, which for an observation near zero is near zero. So “Never” in this table means “never armed automatically from a baseline”, not “not gateable” — and a metric listed here as record-only can still carry a hand-written per-trial bound today.

Revised by measurement. The four-gate set below was derived before the step-0 excursion was measured on the smoke cell. Three of those four fire on it. The First bless column now reflects that; the Why column keeps the original reasoning, because it was not wrong about the noise — it was wrong about warmup_steps: 1 making the excursion unreachable.

Metric Policy First bless Why
median_tpot_ms max Gate Reproduced to 1.57% and 0.26% across sweeps, 3.08% over 12 same-node cell-runs, and the docs already name it “the better-behaved per-token metric”. It is steady-state decode cost with no queueing term, which is why it is the tightest number we have — and it is measured between tokens, so the step-0 compile excursion does not enter it (1.01× on the excursion run). The one gate the measurement leaves standing.
p99_itl_ms max Gate Promoted from record-only for the same reason: 1.00× on the excursion run, 5.43% clean spread. It is the useful half of the ITL pair, it catches tail stalls that a median cannot, and it is the only other metric whose definition excludes the excursion. Gating it and median_tpot_ms together covers per-token latency at both the centre and the tail without touching anything duration-derived.
mean_step_time_ms (step_time_ms.max) max Record-only (was: gate) The bench step duration, and therefore the metric the excursion hits hardest: 2825 ms against a 1462 ms ceiling, 2.4× over. It is still the bound --perf-gate always writes, so arming it is the default — which is exactly why the rollout sequence below prunes it by hand. The window has since shown warmup_steps: 2 covers the excursion, and it stays pruned anyway: it is not a recorded metric, so no per-night series of it exists for a window to measure (step 4).
output_throughput min Record-only (was: gate) Reproduced to 2.63% and 0.30%, and it is the headline number — but it is tokens over the step duration, so the excursion drags it to 0.73× clean and through a 0.85 floor. Nothing is wrong with the metric; it simply cannot be gated while a five-second compile can land inside the window it divides by.
median_ttft_ms max Record-only (was: gate) The original entry said “if one gate flaps, expect it to be this one”, and that was right for the wrong reason — not a flap but a 10.26× excursion, 465.30 against a 58.59 ceiling. It carries queueing delay the others do not, and its first request is the one that waits for the compile.
p99_ttft_ms max Record-only A p99 over 32 requests is the 32nd of 32 order statistics — effectively the maximum, and we have no repeat measurement of it. The load sweep shows it moving 194 → 2139 ms across shapes and 315 → 426 ms for a 2× concurrency change, so it is responsive to things a gate should not fire on. Promote on evidence from the record-only window.
p99_tpot_ms, p99_e2el_ms max Record-only Same order-statistic argument, same absence of repeat data. (p99_itl_ms was in this row and has been promoted to the gate set above — it now has 13 same-node cell-runs behind it, not zero repeats, and it is one of the two metrics the step-0 excursion leaves alone.)
median_e2el_ms max Record-only End-to-end latency is TTFT plus OSL × TPOT, so gating it adds a third alarm for an event two gates already catch, and its reason line is the least specific of the three.
request_throughput min Record-only Under ignore_eos at fixed OSL it is output_throughput / output_len — fully determined by a metric already gated.
total_token_throughput, tokens_per_sec min Record-only Restatements of output_throughput at fixed ISL/OSL (tokens_per_sec is documented as an alias of it). Gating them costs nothing but produces three reason lines for one event, which makes triage slower rather than safer.
median_itl_ms max Never Measured at ~0, because the gateway delivers several tokens per SSE chunk. A margin is multiplicative, so an observation of 0.0 blesses a ceiling of 0.0 and every later run with any inter-token gap at all fails. eval_lib already warned about this in a comment; it is now enforced by _NO_AUTO_GATE.
server_startup_sec — Never 189–379 s in the docs; 180–415 s measured over the 13 smoke cell-runs, so the spread is wider than the quoted one, not narrower. It does not order by model size. Not in the allowlist and must stay out: the allowlist is what --perf-gate arms from, so adding it is gating it.
container_elapsed_sec — Never Dominated by bring-up; same argument.
duration, total_input_tokens, total_output_tokens — Never Work-done counters. The token totals are pinned by the recipe, so a bound on them restates the configuration rather than measuring the stack.
completed_total, failed_total — Never Already enforced, harder, elsewhere: the bench script and the workload independently require failed == 0 and completed == num_prompts, so a shortfall fails the cell and reddens the nightly with no baseline involved. A metric bound here would be a third and weaker copy of a check that already fails closed.
max_output_tokens_per_s, max_concurrent_requests — Never Single-sample maxima — the noisiest available summary of a distribution.
mean_*_ms, std_*_ms, p50_*, p90_* — Never Deliberately absent from the allowlist already (“means are the noisiest summary of a latency distribution and the least useful thing to gate on”). Keep them absent.

Two gates, then — median_tpot_ms and p99_itl_ms, the per-token pair — and everything else charted, including three metrics that would have been gated before the excursion was measured. That is a deliberately smaller first bless than the plan originally proposed: the two that remain are the two whose definitions exclude the failure mode we can actually demonstrate. Of the dropped ones, median_ttft_ms and output_throughput can be promoted from the record-only window once it shows the excursion is gone, which the 2026-09-08..09-17 window has (step 4). step_time_ms.max does not follow them: it is not a recorded metric, so no window measures it. The gap between “gated” and “invisible” is covered by the dashboard’s What changed view, which reports any metric that moved more than 10% between the two most recent runs without failing the job. A 10% throughput drift is real, is not a gate breach under these margins, and is exactly what that view exists to surface.

How many record-only runs, and what the threshold comes from

Take ten nightlies before blessing. Derive each bound from the window’s extremum — the maximum for a max metric, the minimum for a min metric — not from its mean, and not from one night’s observation.

The statistic matters more than the count, so take that first. Deriving from the mean is the intuitive choice and it is wrong here, because the variation we have measured is not jitter around a centre. The startup series is 189 against a cluster of 276–379: a single environmental excursion, not a spread. A mean is precisely the statistic that hides one, and a bound built on it is breached by the next occurrence. The extremum is the statistic that asks the question we actually care about — how bad has a healthy night ever been?

That also fixes the current tooling’s real defect, which is not the margin size but the anchor. refresh_baselines.py --perf-gate derives value × 1.25 from whatever single run was under way, so the bound depends on which night the operator pressed the button. Feeding the seven known startup observations through compare_to_baseline shows what that costs:

Blessed on Ceiling (×1.25) Later runs breaching
189 s 236.2 6 of 6
276 s 345.0 1 of 6
285 s 356.2 1 of 6
291 s 363.8 1 of 6
316 s / 319 s / 379 s 395.0 / 398.8 / 473.8 0 of 6

Four of the seven possible blessing nights produce a gate that breaches, and one of them reddens every subsequent run. The window maximum (379 → 473.8) breaches none — but only because the window contained the excursion, which is the argument for the count.

Ten is chosen from that, not from convention. Scoring every n-run subset of the series against the runs it did not contain gives the residual breach rate on an unseen night:

Window size Mean breach rate on unseen runs Worst window
1 21.4% 6 breaches
2 5.7% 1 breach
3 2.9% 1 breach
4 1.0% 1 breach
5 and up 0% none

The nightly runs roughly 250 times a year, so one false alarm per quarter needs a per-run breach rate under about 1.1% — which this series reaches at n=4, on the noisiest metric the workload produces. The metrics we are actually gating are 10–16× tighter than that one. Zero at n=5 is an artefact of a seven-point sample rather than a real floor, so the honest reading is that the risk is small by 5 and the remaining reason to go further is calendar coverage: ten nightlies is two full weeks, long enough to contain a runner reimage, a Docker Hub re-pull or a cold HF cache — the once-a-week-ish events that produce excursions in the first place. Five is the floor; below it the breach rate on known-noisy data is 2.9% per run, about nine false alarms a quarter.

A stack bump does not get absorbed by the window — it invalidates it

An earlier version of the paragraph above listed “a Dependabot ROCm bump” beside the runner reimage and the cold cache, as one more thing ten nights is long enough to contain. That was wrong, and it is worth correcting rather than quietly dropping, because the two categories look alike and are opposites.

A reimage, a re-pull and a cold cache are environmental events: they perturb a run and the stack underneath is the same before and after, so a window that contains one has measured a genuine bad night and the extremum is doing exactly its job. A digest bump is not that. It replaces the thing being measured half-way through measuring it, so the window no longer describes one population. You can sample a stack bump or derive a stable ceiling across the window, but not both — the extremum then answers “how bad has a healthy night been on either of two stacks”, which is a number about neither.

The counter-example is not hypothetical, and it happened while this branch was open. sanitizers-nightly.yml went red on 2026-09-02 and stayed red, because #411 moved Dockerfile.ci-gpu from ROCm 7.2.4 to ROCm 10 and the f32 GEMM code objects the gate scans are extracted from the image’s own Tensile bundle — so the committed expectation was describing objects the image no longer ships (#453). That is a correctness baseline invalidated by a mid-flight digest change. A perf baseline is not better protected; it is worse, because a shifted number still looks like a number and nothing about it announces that the stack moved. A mid-window bump would do to the perf baseline precisely what that one did to the sanitizer baseline.

Operational condition, for the duration of the window: hold any PR that changes either pinned digest. Concretely, the FROM … @sha256 in docker/Dockerfile.ci-gpu (currently rocm/pytorch:rocm10.0_ubuntu26.04_py3.14_pytorch_release_2.13.0@sha256:3174cb70…) and the engine digest in recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml (lightseekorg/tokenspeed-amd@sha256:60c12e37…). If one has to land, restart the count; do not average across it.

This is preventable by policy rather than by luck, which is the useful part. .github/dependabot.yml runs the docker ecosystem against /docker on a weekly schedule and its own comment says the job “bumps the FROM ... @sha256 digests”, so an in-window bump is not merely possible but scheduled. There is no auto-merge — the same file says so explicitly, and a human therefore has to approve and merge one for it to reach the window. Ten nights is long enough for roughly two of these to be opened, so the hold is a real decision someone will be asked to make twice, not a theoretical one.

Checked at the time of writing: the only open PR in the docker ecosystem is #309 (bump ubuntu from 22.04 to 26.04 in /docker), and it is clear — it touches Dockerfile.rocm-ubuntu-ebpf, Dockerfile.rocm70_2-ubuntu-nan and Dockerfile.rocm70_2-ubuntu-pytorch, none of which is Dockerfile.ci-gpu, and it does not move the engine digest. The other open dependabot/* PRs (#310–#314) are github-actions bumps. So nothing currently open needs holding; re-check before the window starts, since the ecosystem is on a weekly timer.

Everything else in the ten-night justification stands. The residual-breach-rate analysis is unaffected — it is a statement about sampling a single population, which is what this condition exists to preserve — and so is the calendar-coverage argument for the three environmental events it still lists.

The extremum anchor has no good answer on a bimodal cell

That whole derivation assumes the window’s extremum is a healthy worst case. The step-0 excursion breaks the assumption, and it is worth being explicit about why, because it is the reason the gate set shrank rather than the window growing.

With a bimodal cell the ten-night window either contains an excursion or it does not, and both outcomes are bad:

Window median_ttft_ms anchor Ceiling Consequence
No excursion (12 of 13 nights) 46.87 58.59 The excursion, when it lands, reads as a 10× regression. False alarm.
Contains one (1 in 13) 465.30 581.63 A genuine 2× TTFT regression — 47 → 94 ms — passes comfortably. No detection power at all.

Enlarging the window does not resolve this; it only makes the second row more likely. The extremum is the right statistic for a unimodal metric with occasional environmental excursions, which is what the startup series is. It is the wrong statistic for a metric with two modes, because “the worst a healthy night has been” is not a single number any more.

So the correct response to a bimodal cell is not a cleverer threshold. It is either to gate a metric the second mode does not reach — which is what median_tpot_ms and p99_itl_ms are — or to remove the second mode. For the step-0 excursion the second option is genuinely available, because the excursion is positional: raising warmup_steps from 1 to 2 discards it by construction.

That has now been done — the recipe sets warmup_steps: 2 — which is what lets median_ttft_ms and output_throughput be promoted later. (step_time_ms.max, the third metric the excursion blocked, stays pruned for a reason warmup_steps does not touch; see step 4.) The cost was that it changed the measurement: every number taken before the change was at warmup_steps: 1, so none of them could be a baseline, and the ten-night record-only window had to be taken afresh at the new setting before anything was blessed. It was: the 2026-09-08..09-17 window is at warmup_steps: 2, and the two gates are blessed from it (step 4). That was a deliberate trade. Carrying the excursion into the window instead would have meant either blessing a bimodal cell or spending ten nights establishing a distribution we already intended to change.

For the concurrency-64 stall no such fix exists — it has no position to discard — which is why that cell stays out of the nightly entirely rather than being gated on a narrower metric.

The margins themselves need no change. --step-time-margin 0.25 and --throughput-margin 0.15 are already right for these metrics, and the simulation says so: against the measured cross-sweep spread, every one of the twelve bless-one-run-check-the-other combinations passes. Against the window extremum the separation is wide — worst observed TPOT noise 1.57% versus a 25% detection threshold, a factor of 16 — while a regression the size of a genuine stack change is caught:

Scenario median_tpot_ms vs 2.425 output_throughput vs 3007.3
the other measured night 1.910 → pass 3631 → pass
+10% / −10% 2.134 → pass 3184 → pass
+30% / −25% 2.522 → fail 2654 → fail
+94% / −46% (the 0.6B → 4B step) 3.770 → fail 1905 → fail

Those rows are in tests/ci/test_eval_lib.py, run against the real comparator, so the claim is checked rather than asserted.

What blocked this, and what remains

Nothing blocks the first bless any more: the matrix entry is live, the socket is signed off for the nightly lane, nightly-eval.yml sets docker_socket: true, and the ten-night measurement is taken (2026-09-08..09-17, all ten scheduled workflow_run events, all ten green). The two ceilings that window sized are in config/ci/regression_baselines.yaml.

What remains is step 7, and not one of the nine record-only metrics is blocked on a measurement this window failed to take. Two of them (median_ttft_ms, output_throughput) have their evidence already and are deferred only so the first armed gate stays attributable; three are waiting on the analysis of nights already recorded; four are held back on a redundancy argument no number will settle. The table in the current-state box at the top says which is which — read it before opening a step-7 PR, because “wait for another window” is the first answer for none of the nine.

This section is kept because the shape of the plumbing is what the sign-off was given against, and because the argument is worth being able to re-read.

nightly_eval.py runs inside the aorta-ci-gpu container (eval-reusable.yml → docker_cmd.sh exec), and tokenspeed_serve runs the TokenSpeed engine in a container of its own. So the entry needs a Docker client and a route to a daemon in there. It had neither, and the workload’s setup() raised 'docker' not on PATH before anything else happened. Record-only does not help — it defers perf bounds, not failures.

Both pieces now exist, the image has been built, and a cell has been run to completion from inside the container (see Verification). The client half is therefore done and proven. The daemon half is a per-lane opt-in, granted to the nightly lane and to nothing else — see Security.

The entry lives in entries with needs_docker_daemon: true, which is what keeps that a configuration choice rather than a load-bearing one: on any lane without the socket the entry skips with the reason recorded, rather than failing. That still matters after the sign-off, because it is what keeps the socket scoped to one lane — refresh-baselines.yml, bump-validate.yml and every other consumer of the matrix can run without it. tests/ci/test_nightly_eval.py covers both halves on the nightly path, and refresh_baselines.build_baselines honours the same flag, with both halves covered too.

The enabling change

1. A Docker client in the CI image. docker/install_docker_cli.py, invoked from Dockerfile.ci-gpu, fetches the pinned static tarball, checks its sha256 before extracting, and extracts exactly one member — docker/docker. The same tarball ships dockerd, containerd and runc; naming the member is what keeps a daemon out of the image, and a build-time check fails the build if any of them reaches PATH. The engine container is a sibling, started by the host daemon next to aorta-ci-gpu, not nested inside it.

2. A route to the daemon, opt-in per lane. docker-compose.docker-socket.yaml is an override, which is the mechanism the base compose already documents for optional mounts (“Do not add a volume here”), so the default container still has no route to the daemon. rocm-ci-setup takes a docker-socket input, default false, and adds the override’s -f only when it is true.

That input is reachable from a caller workflow through a matching docker_socket workflow_call input on eval-reusable.yml, also defaulting to false and forwarded to the setup step. The default is what matters here: eval-reusable.yml is shared by the nightly and by bump-validate.yml, so the choice has to be the caller’s rather than the reusable workflow’s — setting it centrally would grant the socket to a PR-triggered lane as a side effect of enabling the nightly. bump-validate.yml therefore pins it false explicitly, and sanitizers-nightly.yml never sees it, using rocm-ci-setup directly with no docker-socket argument. nightly-eval.yml is the only lane that sets it true — see the security note below for the sign-off and its basis.

3. work_dir must be an explicitly configured shared path. This is the detail most likely to be missed, because nothing reports it as a path problem.

With the socket mounted, the -v sources in the docker run the workload builds are strings the host daemon resolves against the host filesystem. The workload builds them from work_dir — <work_dir>/u<uid>/{scripts,out,hf} — using paths as seen from inside aorta-ci-gpu. Those are different namespaces. A missing bind source is not an error to the daemon: it creates the directory on the host and mounts that. So the engine container comes up with an empty /ts-scripts and dies on a missing script, or — worse — with an empty writable /ts-out, and the harvest reports no results for a run that really happened.

The default work_dir does not fix this by being /tmp/ts-work-serve on both sides. /tmp inside the CI container is the container’s own /tmp, not the host’s, so the two are different directories that share a name — which is the failure above with the confusing property that the path looks right in every log.

What the nightly needs, concretely:

Measured, and more specific than the above was. Demonstrating this on a gfx950 node produced two failures worth writing down, because neither is what the guidance as written would have led you to expect.

Every bind source must be resolvable by the daemon, not only work_dir. The first attempt put AORTA_WORKSPACE on an autofs-mounted NFS home and the container never started:

Error response from daemon: error while creating mount source path
'/home/.../wt-gating': mkdir /home/...: permission denied

The daemon resolves all -v sources in the host namespace, so a checkout on an automounted or root-squashed filesystem fails the same way a work_dir there would — and it fails at container create, before any of the workload’s own checks can produce a better message. On a GitHub runner the workspace is local disk and this does not arise; on any node where /home is networked, the repo has to be staged locally first. Worth knowing before debugging it as a socket problem, which is what it looks like.

The work root must be owned by the uid the container runs as, and the obvious way to create it gets that wrong. aorta-ci-gpu runs as root, so the workload checks u0 and requires the root itself to be root-owned. A work_dir created by an ordinary host-side run is owned by that user, and the containerised run then refuses it — the node used here already had a /tmp/ts-work-serve at uid=100550 mode=1777 left by earlier host-side sweeps, which uid 0 must reject. The demonstration used a fresh path created through the daemon so it landed root-owned, which is exactly what the workload’s own error message advises an administrator to do.

Creating it is not the same as repairing it, and the runner needs the second one. mkdir -p is a no-op on a directory that already exists: run as root over a leftover user-owned root it returns 0 and changes neither owner nor mode, so the next run fails the same ownership check with the remedy apparently already applied. Say what the end state is instead — the sequence below is idempotent whether or not the directory is there, and is what to run on the runner before the nightly first executes the entry:

docker run --rm -v /tmp:/mnt busybox:1.37 sh -c \
  "mkdir -p /mnt/ts-work-serve && chown 0:0 /mnt/ts-work-serve && chmod 1777 /mnt/ts-work-serve"
docker run --rm -v /tmp:/mnt busybox:1.37 stat -c "%n %u:%g %a" /mnt/ts-work-serve

chown the root and nothing under it: the per-uid scratch directories beneath it belong to whichever uid created them, and -R would take those from a co-tenant. The second line is worth running — the check the workload makes is on the root’s owner, so 0:0 1777 is the whole acceptance criterion.

The operational trap in that is worth stating on its own: the ten record-only runs cannot share a work_dir with the containerised nightly if they are taken by hand as an ordinary user. Either take them as root, or give the two a separate work_dir and accept the cold HF cache on the first nightly.

The uid has to agree too, and this is a second way the same boundary bites. The workload uses a per-uid scratch root and then verifies it: <work_dir>/u<uid> must be a real directory (not a symlink), owned by the running uid, and not group- or world-writable. Inside aorta-ci-gpu the process is root (user: root in the base compose), so it writes and checks u0. That works — a root-owned directory created through the shared mount is root-owned on the host too, so the host daemon and the checking process agree. But it holds only because both sides are uid 0. Change user: in the base compose, run the harness as a non-root uid, or turn on userns-remap on the daemon, and <work_dir>/u<uid> is created by one uid and inspected by another: the check fails with an ownership error that does not mention containers, uids-across-a-boundary, or the mount. Anyone changing either should expect that error to be the symptom.

One consequence worth stating: with run_as_current_user defaulting true, the engine container also runs as uid 0, so exports land on the host owned by root, outside the workspace that eval-reusable.yml’s “Reclaim results ownership” step chowns. They are cleaned per trial unless keep_work_dir is set, but a runner that fills /tmp with root-owned scratch is a plausible future complaint.

Security: the marginal risk, and what actually bounds it

Signed off for the nightly lane on 2026-09-03, and nightly-eval.yml sets docker_socket: true accordingly. The reasoning is recorded here rather than only in a review thread, because it is the durable part: a future reader deciding whether a second lane may have the socket needs the argument, not the verdict.

Stated plainly first, because it is the part most easily lost in a diff: a process that can talk to the Docker daemon can ask it for a privileged container that bind-mounts /. That is root on the host, with no exploit involved. So the mount is genuinely root-equivalent, and nothing below disputes that.

What it is not is a new capability on this runner. Before any container exists, eval-reusable.yml’s first step runs

docker run --rm -v "$ws":/ws "$busybox" chown -R "$uidgid" /ws

as a plain host step, falling back to sudo -n docker info and then to sudo chown -R. refresh-baselines.yml opens the same way. So the runner account already has direct daemon access and passwordless sudo, and any job step that can execute on this runner is already host-root capable — with or without the socket. The socket makes that reachable from inside aorta-ci-gpu as well as from the job shell, which is a convenience, not a new privilege. Reasoning about it as though the container were the boundary overstates what the container was ever doing.

What bounds the risk is write access to the repository, and that is checkable. Three things, each of which can be re-verified by grep rather than taken on trust:

So the population that can reach the daemon through this mount is exactly the population that can push a branch to this repository — which is the population that can already run sudo on the runner through any workflow step. That is the whole of the marginal risk, and it is close to zero.

The “single-tenant runner” framing is withdrawn, because it was doing work the argument does not need and cannot currently support. ROCm/aorta and ROCm/aorta-internal hold separate repo-level runner registrations (smci350-rck-g03-f16-12 and smci350-internal), and the internal registration’s July jobs reported machine sharkmi300x-3 — an MI300X box — despite carrying an mi350x label. Whether the two registrations share a physical host is not confirmed as of today. It does not need to be: the argument above rests on what the runner account can already do and on who can reach it, neither of which changes with tenancy. Co-tenancy would matter for measurement (a neighbour’s load moving the numbers), which is a separate question handled by the ten-night window and the two-cell control, not by this grant.

The alternatives, kept for the case where a future lane wants the socket and this argument does not carry over. Run the serving workload outside the CI container entirely — as a step on the runner host, which already has a client and the socket, publishing its results JSON into the harness’s results directory. A socket proxy allowlisting just the container create/start/logs/remove calls is a middle option; rootless Docker is a third, and it is the only one that is a real reduction rather than a re-arrangement, precisely because of the sudo point above. Nested docker-in-docker is not on the list: it needs privileged too, so it trades no privilege away, and it gives the engine a different filesystem view, which breaks exactly the bind mounts described above.

What is still missing

Nothing in the plumbing. Two operational preconditions before the entry first runs, both on the runner rather than in this repository:

  1. /tmp/ts-work-serve must exist root-owned and 1777 — the sequence is under work_dir above, and note that mkdir -p alone does not repair a pre-existing user-owned root.
  2. Both pinned digests must hold for the whole window — the base image in docker/Dockerfile.ci-gpu and the engine image in the recipe. See the stack-bump section.

The entry is in entries rather than pending_entries, carrying needs_docker_daemon: true, so on any lane without the socket — the baseline refresher, bump-validate — it skips with the reason in the results instead of failing on Cannot connect to the Docker daemon. Promoting it without that flag would have been the mistake the matrix file’s own header warns about: an entry may only be added once it can actually pass on the runner.

Verification status

Built and run on a gfx950 (MI355X) dev node on 2026-09-02 — not on the CI runner, which is an MI350X. Read the hardware note further down this section before treating any absolute number here as something the nightly should reproduce.

The image builds and is client-only. docker compose --env-file .env.ci -f docker-compose.build.yaml build succeeds, producing a 51.9 GB aorta:ci-gpu. The build’s own RUN checks pass — docker client check: OK -- /usr/local/bin/docker, no daemon on PATH — and the assertion holds when re-checked against the built image rather than during it: dockerd, containerd, runc, containerd-shim-runc-v2, docker-proxy and docker-init are all absent from PATH, and docker --version reports 29.7.2. That is what makes “client only” a property of the artifact and not of the Dockerfile’s intent.

The client reaches the host daemon from inside the container. With the override mounted, docker exec aorta-ci-gpu docker ps returns the running container list, exit 0.

A cell completes. Two sweeps of tokenspeed-serve-bench-smoke.yaml were driven from inside aorta-ci-gpu, each starting the TokenSpeed engine as a sibling container: four cell-runs, all passed, failed_total 0 and completed_total 96 throughout, engine containers cleaned up afterwards. The twelve measured steps were 1108–1149 ms with no step-0 excursion, and every metric landed within 5% of the host-side clean envelope — so the container boundary does not move the measurement, which is what lets a nightly baseline be compared against the host-side runs the variance analysis is built from.

[!IMPORTANT] The 1108–1149 ms envelope is not a night-one acceptance check. It was measured on an MI355X dev node. The CI runner smci350-rck-g03-f16-12 is an MI350X — same gfx950/CDNA4 ISA, but a lower power budget and air rather than liquid cooling. Correctness is unaffected by that difference: the ISA is identical, so the kernels, the compiler output and every verdict the harness produces are the same. Absolute step times are a different matter and will plausibly differ — a priori, at least; one MI350X cell has since been measured and did not show an offset, which narrows the expectation without licensing the envelope as a check. See the subsection below.

So do not use this envelope to accept or reject night one, in either direction. A first nightly landing outside 1108–1149 ms is not evidence of a problem, and one landing inside it is not evidence that the lane is healthy.

What has to reproduce across the window is the spread, not the level. The ceilings are derived from the ten CI nights themselves — max × 1.25 and min × 0.85 over that window, on that runner — so a systematic offset between the dev node and the runner is absorbed by construction and is expected, harmless and not worth investigating. What would matter is the shape: a window spread materially wider than the measured few percent, or a step-0 excursion, and step 4 already checks both. Judge night one against the two checks in step 4 and against the previous nights on the same runner, never against a number taken on other hardware.

One MI350X cell at warmup_steps: 2 (2026-09-03) — a data point, not a spread

Both caveats above — MI355X hardware, and warmup_steps: 1 — are now partly addressed by one measurement, and it is recorded here with its limits stated first because it is easy to over-read.

This is a single cell-run. It is not a variance estimate and must not be used as one, and it does not replace anything above. The 13-cell-run table stays exactly as it is — it remains the rationale for warmup_steps: 2, and one cell cannot restate it. The ten-night window is still the only thing that produces a spread at the new setting.

Run on cv350-rck-g03-c16-18 (Slurm partition meta64), which is an MI350X: 1000 W package power cap, VBIOS 113-M350-01-1K5-000C, gfx950, 288 GiB HBM. The 1000 W cap and the air-cooled VBIOS are what identify it as an MI350X rather than an MI355X, and it is the same hardware class as the CI runner — which is the point. Deliberately not on smci350-rck-g03-f16-12, which is reserved for the baseline window. One baseline cell of tokenspeed-serve-bench-smoke.yaml with every measurement-relevant setting as committed (image digest, Qwen3-0.6B, ISL 512 / OSL 128, concurrency 8, 32 prompts, 3 measured steps, num_warmups: 1, warmup_steps: 2, ignore_eos, seed 0), on a private work_dir so it cannot touch the CI lane’s. Passed, failed_total 0, completed_total 96.

Metric MI350X, 1 cell, warmup_steps: 2 MI355X clean range, warmup_steps: 1
step times, in order 1135.03 / 1136.39 / 1136.51 ms 1108 – 1178 ms
median_ttft_ms 45.00 43.65 – 46.87
median_tpot_ms 1.910 1.89 – 1.95
p99_itl_ms 34.64 34.00 – 35.84
output_throughput 3605.7 3502.53 – 3646.20
server_startup_sec 282 180 – 415

Two things worth taking from it, and one worth not taking.

The MI350X level is not detectably different at this workload size. Every metric lands inside the MI355X clean range, step time included. That is a plausible result rather than a surprising one: Qwen3-0.6B at concurrency 8 is a small, latency-bound load that does not approach either card’s power ceiling, so the MI355X’s higher budget and liquid cooling have nothing to bite on. It does not license using the MI355X envelope as an acceptance check — one cell cannot establish agreement, and the rule above stands unchanged — but it does mean a large night-one offset would itself be worth a look, rather than being shrugged off as expected hardware difference.

No step-0 excursion, with the excursion’s own signature absent. The first measured step is the fastest of the three, so there is no positional outlier at all. This is the first observation at warmup_steps: 2 and n=1, so it is consistent with the change working and nowhere near sufficient to conclude it; step 4’s bimodality check over twenty cell-runs is what settles that.

What not to take from it: the three step times span 0.13%, and that number is not comparable to the 6.3% step-time or 3.08%/5.43% metric spreads in the table above. Those are across twelve cell-runs — separate server bring-ups on separate invocations, which is where the variance lives. 0.13% is three consecutive steps against one already-warm server, and it is the within-cell figure that the earlier table never reported. Reading it as a 48× variance improvement would be a category error. Re-taking the spread needs the window.

Two configuration requirements were discovered by doing this rather than by reading, and are written up under work_dir above: all bind sources must be resolvable by the daemon (an autofs NFS checkout is not), and the work root must be root-owned because the container runs as uid 0.

Also verified earlier, and still true: the installer rejects a wrong sha256 without leaving a binary behind; the client negotiates against an older (29.1.3) daemon; docker compose config resolves all three mounts with the scratch mount identical on both sides.

Still not verified, and it is now the first thing the nightly will establish: the entry has never run under nightly_eval.py on the CI runner, only under aorta sweep run in the container by hand. Night one is therefore a bring-up observation as much as a measurement — treat a failure there as plumbing until shown otherwise, and see the two operational preconditions under What is still missing before counting it. The skip path is covered by tests rather than by a runner.

To reproduce:

cd docker
bash ../scripts/ci/docker_compose.sh --env-file .env.ci -f docker-compose.build.yaml build
docker run --rm aorta:ci-gpu docker --version

# The work root must be root-owned, and `mkdir -p` on its own does not make it
# so: on a node where an earlier host-side sweep left a /tmp/ts-work-serve owned
# by that user, `mkdir -p` as root sees an existing directory and returns 0
# having changed nothing -- so the run still fails the ownership check, with the
# remedy apparently already applied. State the owner and the mode explicitly;
# the sequence is idempotent whether or not the directory is there.
#
# `chown` on the root only, never `-R`: the per-uid scratch beneath it belongs
# to whoever created it, and the check the workload makes is on the root's owner.
docker run --rm -v /tmp:/mnt busybox:1.37 sh -c \
  "mkdir -p /mnt/ts-work-serve && chown 0:0 /mnt/ts-work-serve && chmod 1777 /mnt/ts-work-serve"

export TS_SERVE_WORK_DIR=/tmp/ts-work-serve
bash ../scripts/ci/docker_compose.sh --env-file .env.ci \
  -f docker-compose.build.yaml -f docker-compose.docker-socket.yaml up -d
docker exec aorta-ci-gpu docker ps
docker exec aorta-ci-gpu aorta sweep run \
  --recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml \
  --output-dir "$TS_SERVE_WORK_DIR/out" --strict

The rollout sequence

Steps 1–6 are done: the entry is live, the nightly lane has the socket, the ten-night window was taken 2026-09-08..09-17, and the two bounds it sized are in config/ci/regression_baselines.yaml. Step 7 is the only one outstanding.

The steps are kept written as instructions rather than rewritten as history: they are the procedure for the next bless (step 7 here, and any other workload’s first bless), and what each one was checked against is the part worth being able to re-read.

1. Promote the entry. Done. It is in entries with min_gpus: 1, timeout_sec: 3600 and needs_docker_daemon: true.

timeout_sec: 3600 rather than the 1800 default: two bring-ups at the observed spread plus the teardown VRAM drain is around 15 minutes before the image pull and any cold-cache weight download, and a timeout is an unconditional entry failure rather than a slow record. Check the job’s own timeout-minutes: 150 in eval-reusable.yml still has room. Measured end to end from inside the container, a full two-cell sweep took 12 minutes on a warm cache, so the budget is right but most of it is bring-up.

2. Verify it before merging. Done, and rerun these after any change:

python -m pytest tests/ci -q --timeout=180
python -m pytest tests/workloads/test_tokenspeed_serve.py -q
aorta sweep run --recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml --dry-run

2a. Enable the socket on the nightly lane. Done, signed off 2026-09-03 — see Security for the basis. It is one line in .github/workflows/nightly-eval.yml:

    uses: ./.github/workflows/eval-reusable.yml
    with:
      docker_socket: true

The docker_socket input on eval-reusable.yml defaults to false and is set per caller, so this arms the nightly lane and nothing else. bump-validate.yml sets it false explicitly: that lane runs PR head code on the same self-hosted runner and must keep no route to the host daemon. sanitizers-nightly.yml does not go through eval-reusable.yml at all — it uses the rocm-ci-setup action directly and passes no docker-socket, so it inherits the action default of false. Setting the input inside eval-reusable.yml instead of per caller would hand the socket to every lane at once and is the mistake this shape exists to prevent.

The grant is per lane, and it should stay that way. Any other lane wanting the socket is a fresh decision, and the argument recorded under Security does not transfer to a lane whose trigger is reachable from a fork or from an unreviewed branch — that argument is what bounds this one. The dashboard also makes the absence visible rather than silent: on a lane without the socket the entry reports skip with needs a docker daemon: ... in its reasons.

3. Let it record for ten nightlies, at warmup_steps: 2. Done — 2026-09-08 to 09-17, ten consecutive scheduled workflow_run events, every one green, no workflow_dispatch inside the range. While it ran it reported recording on the dashboard.

The window is only valid at the setting the gate will run at, and that setting is now warmup_steps: 2 in recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml. Confirm that before counting nights, and restart the count if it changes underneath the window. None of the numbers earlier in this document can substitute for any of these ten: they were measured at warmup_steps: 1, and the change was made specifically to alter the behaviour they describe. Expect the window to be cleaner than the 13-cell-run table — that is the change working — but do not assume it; the point of the window is to measure it rather than predict it.

And do not accept or reject night one against any number in this document. All of them were taken on an MI355X dev node; the runner is an MI350X. A systematic offset in the absolute level is expected and harmless, because the ceilings come from these ten nights on this runner. It is the spread that has to reproduce — see the hardware note under Verification status, and the single MI350X data point recorded there.

Also confirm before counting that neither pinned digest is about to move: an in-window bump invalidates the window rather than being absorbed by it, which is its own section and the reason for the digest hold.

Watch the Workloads view and write the numbers down; the ten values of median_tpot_ms and p99_itl_ms per cell are the input to step 4, and the ten of median_ttft_ms and output_throughput are what decides whether those two record-only metrics can be promoted later. The ten of mean_step_time_ms are what step 4 reads to tell whether the excursion is gone. Two cells, so twenty observations of each. A fail during this window is a real failure — the entry is unbaselined but the harness is fail-closed — and must be fixed rather than waited out.

4. Check the window before blessing, and check it for bimodality first. Done, and both checks passed — with one number that looks alarming and is not, recorded below the checks. Compute, per cell and per metric, the extremum and the ratio of extremum to median. Two separate checks:

What the 2026-09-08..09-17 window actually showed, per cell and per gated metric:

cell metric min median max max/median full range
baseline median_tpot_ms 1.7373 1.8906 1.9138 1.0122 9.33%
baseline p99_itl_ms 34.1259 34.7533 35.4296 1.0195 3.75%
no-scratch-reclaim median_tpot_ms 1.8604 1.8945 1.9196 1.0132 3.13%
no-scratch-reclaim p99_itl_ms 33.9241 34.9653 35.8309 1.0248 5.45%

The 9.33% is not the spread-check failing, and the reason generalises. Both metrics are max policy, so the number that sizes a ceiling and the number that can breach it is the upside deviation, not the full range. That 9.33% is a single low outlier on 09-13 — the cell ran faster than usual — against a ceiling it therefore cannot approach. Upside deviation never exceeds 2.48% anywhere in the window, which is what the 10% check is about. Read the max/median column, not full range, when applying this check to a max metric; the reverse for a min one.

Bimodality: clean. mean_step_time_ms stayed within 1113.98–1148.66 ms across all twenty cell-runs and the step-0 compile excursion did not recur once, so warmup_steps: 2 did what it was changed to do.

If the window shows no excursion in twenty cell-runs, that is reasonable evidence it has stopped happening, and median_ttft_ms and output_throughput can be promoted with the same margins.

It did, and they were still not promoted in the first bless — deliberately. The window clears the excursion blocker on median_ttft_ms and output_throughput, so the evidence for promoting them is now in hand and promoting them is an ordinary step-7 PR. step_time_ms.max is not in that list, and a clean window does not put it there: it is not a recorded metric, no per-night series exists for it, and a bound synthesised from one night’s mean at bless time is not something this window measured. Arming the first gate on the two metrics whose definitions exclude the failure mode we can demonstrate, and promoting median_ttft_ms and output_throughput on their own evidence in their own diff, keeps a gate that fires attributable to the change that armed it.

This window is the first measurement at warmup_steps: 2, so it is also the test of whether that change did what it was supposed to. Two outcomes worth telling apart: no excursion in twenty cell-runs is the expected result and clears median_ttft_ms and output_throughput for promotion in step 7; an excursion that still appears at position 0 means two discarded steps are not enough to cover the compile, which is new information and should be investigated rather than absorbed by raising warmup_steps again.

5. Bless, scoped to this entry only. Done, by hand — see the note immediately below for why a refresher run was never an option for this entry.

[!IMPORTANT] For tokenspeed_serve_smoke, this is a hand-written baseline, not a refresher run. refresh-baselines.yml does not mount the docker socket — the sign-off is scoped to the nightly lane — so on that lane the entry declares a capability the lane does not have and skips. Dispatching a refresh with perf_gate_entry: tokenspeed_serve_smoke is therefore refused outright, with a message saying so, rather than producing a correctness-only file that reads in the PR diff exactly like a successful bless.

This costs less than it sounds like, because step 6 below replaces every number the refresher derives with one computed from the ten-night window, and deletes most of the keys it writes. What the refresher would contribute is a skeleton. So write the two keys into config/ci/regression_baselines.yaml directly, in a PR, using the values from step 6 — same review, same file, fewer moving parts, and no GPU runner time.

That is safe now in a way it was not before: a later refresh of any other entry carries these bounds over untouched, in both the scoped and the default correctness-only modes. Two options if the refresher’s convenience is wanted for this entry later: extend the sign-off to the refresh lane (a dispatch-only lane, so the fork-guard part of the argument is trivially satisfied — but it is a separate decision), or run the sweep on the runner host and hand-transcribe. Neither is on the critical path.

The mechanism itself, for the entries that can use it:

Actions -> Refresh baselines -> Run workflow

with the dispatch form filled in as

Input Value
perf_gate true
perf_gate_entry the entry under test, e.g. inference_offline
step_time_margin 0.25 (default)
throughput_margin 0.15 (default)

which the workflow turns into

python scripts/ci/refresh_baselines.py --perf-gate \
  --perf-gate-entry inference_offline

The scope is not optional. Without it the same PR arms step-time ceilings for every other entry in the matrix from that one run, so leaving perf_gate_entry empty with perf_gate: true logs a warning on the job. It is a warning rather than a hard failure because an unscoped refresh is the legitimate end state once every workload has variance data behind it — but during rollout it is not what you want. Three ways of getting the scope wrong all fail rather than degrade: a misspelled entry name is rejected rather than silently scoping to nothing; a scope without perf_gate fails immediately instead of after the wheel install; and a value that is non-empty but is only delimiters (,, a stray tab) is rejected too, because it would otherwise parse to no names and leave a bare --perf-gate — the unscoped global refresh, reached by an operator who thought they had scoped it, and without even the unscoped warning.

What the scope does not do any more is disarm anything. Out-of-scope entries keep their existing step_time_ms and metric bounds verbatim while their correctness data is refreshed, which is what makes per-workload rollout additive rather than a game of whack-a-mole. The same holds in the default, correctness-only mode: omitting --perf-gate means “do not arm new gates”, not “disarm the ones already blessed”.

perf_gate_entry accepts several names, comma- or space-separated, mapping to one --perf-gate-entry each. For this rollout it should be exactly one.

6. Write the two bounds, then merge. Done — the four keys now in config/ci/regression_baselines.yaml are max × 1.25 on the window maximum of each cell’s median_tpot_ms and p99_itl_ms, and nothing else was written. Whether the numbers come from a refresher diff or are written by hand (as they are for this entry — see the note in step 5), the same two edits apply, because the refresher derives bounds from the single run it just did and from every auto-gateable metric it observed:

step_time_ms.max is the one that needs attention, because it is the bound --perf-gate always writes, the only one on this list that is not a metric key, and — when hand-writing — the one most easily added out of a sense of completeness. Including it arms the gate the measured excursion breaches hardest: 2825 ms against a 1462 ms ceiling. It is the most likely mistake in this whole sequence. The end state is exactly two metric keys per cell, median_tpot_ms and p99_itl_ms, alongside passed: true.

This hand-editing is the honest cost of the current tooling, and it is bounded: ten keys across two cells, on a PR a human reviews anyway, and it is meant to shrink as the record-only metrics are promoted on evidence. It got larger rather than smaller when the excursion was measured, which strengthens the case below for a per-metric scope.

7. Promote the record-only metrics on evidence, not on schedule. The nine are blocked on three different things, and none of the groups is waiting on nights that have not happened yet: each starts from evidence the completed window already holds, and a second window is an outcome the middle group’s analysis might reach, not a precondition for doing it. The groups are the ones in the current-state box at the top of this file; this is what to do about each.

Do the middle group rather than skipping it, and do not let the easy group crowd it out: a gated p99_ttft_ms is worth more than a gated median_ttft_ms, because tail latency is what a serving regression damages first. The first group is the one with no work left in it, not the one with the most value in it.

step_time_ms.max is not one of the nine and does not belong to any of these groups, and a clean window does not clear it: it is not a recorded metric, no per-night series accumulates for it, and a bound synthesised from one night’s mean at bless time is not something a window measured. Promoting it needs that to change, not more observations.

If the hand-editing recurs every refresh rather than converging, the fix is a per-metric scope alongside --perf-gate-entry. That is the point to build it — not before, when the right metric set is still a guess.

When a gate fires

The alert names the cell, the metric, the observed value and the bound. Three things it can be, and they are distinguishable without a GPU:

A stack change. Check the dashboard’s What changed line first: it reports whether PyTorch, ROCm or HIP moved that night. If one did, and the direction of the metric move is plausible for it, this is a stack effect and the answer is either an upstream bug report or a re-bless. Do not re-bless silently — a bump that costs 20% of decode throughput is exactly the result the nightly exists to produce, and docs/ci-nightly-eval.md is explicit that the baseline diff is the bump’s impact.

A noise event. The two cells of this recipe differ only by mitigation and run minutes apart on the same node, which is the reason this recipe was chosen. Use it: if baseline and no-scratch-reclaim both moved by a similar amount, the node or the stack changed. If exactly one moved, and the mitigation has never shown an effect before at any model size, that is a single-cell excursion — a co-tenant, a thermal event, a cold cache. Confirm by looking at the previous nights on the dashboard’s run history; a noise event is one red cell in a column of green, a regression is a column that turns red and stays red. Also check server_startup_sec for that cell: it is not gated, but a bring-up at the top of its range is a good indicator that the node was busy.

A real regression. Everything else — both cells moved, no toolchain change, and the next night reproduces it. Reproduce locally with the recipe as committed and bisect against the pinned image digest. Note that the image is pinned by digest precisely so this case cannot be caused by TokenSpeed changing underneath the baseline.

If you cannot tell which of the three within one working day, revert to record-only and keep collecting. An unexplained gate is worth less than the recording state it came from.

Reverting to record-only

Delete the metrics and step_time_ms keys for the affected cells from config/ci/regression_baselines.yaml, leaving passed: true. That returns those cells to record-only immediately — the comparator treats an absent bound as nothing to check — while keeping correctness gating and the fail-closed behaviour intact. It is a small, reviewable, obviously-correct diff, which is what you want at the point where a gate is misbehaving.

Do not revert by deleting the matrix entry. That stops the measurement as well as the gate, and the record-only data is the thing needed to size a better bound.

Do not revert by widening the margin as a first move. A margin wide enough to absorb an unexplained excursion is usually too wide to catch a regression, and the widening tends to be permanent. Go back to record-only, find out what moved, then re-bless from a longer window.

One whole-file caveat, now much narrower than it was: any later refresh-baselines run rewrites regression_baselines.yaml completely, but a run that is not deriving bounds for a cell carries that cell’s existing bounds over untouched — so a hand-reverted cell stays reverted through a scoped refresh of another entry and through a default correctness-only refresh. The one thing that re-arms it is an unscoped --perf-gate refresh, which re-derives every cell it can. Scope those runs with --perf-gate-entry, or check the diff for cells you had deliberately demoted.

“Every cell it can” is the precise claim, and the exception matters for this rollout specifically: a bound --perf-gate cannot produce is never deleted by it either, because there would be nothing to put back. That covers a hand-written median_itl_ms ceiling — the metric is allowlisted but _NO_AUTO_GATE, so the refresher skips it — and any spec naming a metric off the allowlist entirely. Both survive a re-bless of their own cell. So if the window later justifies gating median_itl_ms by hand, a subsequent refresh of this same entry will not quietly undo it.

Assumptions

The user referenced Aorta planning & Updates.docx, which is not reachable from this machine. Everything above was derived from the repository, and where that was not enough a decision was made and is recorded here.