tokenspeed_serve reports TTFT, TPOT, ITL and
throughput, and the serving metric names are already in the nightly’s gating
allowlist. A serving recipe is in config/ci/nightly_eval_matrix.yaml, and this
document is the sequence that took it from record-only to gated, plus the
reasoning behind the numbers it picks.
[!IMPORTANT] Current state: two metrics are gated, nine are not.
config/ci/regression_baselines.yamlcarriesmedian_tpot_msandp99_itl_msasmaxceilings on bothtokenspeed_serve_smokecells (baselineandno-scratch-reclaim), derived from the ten-night window 2026-09-08..09-17. A run over either fails the nightly, and the breaching observation is still recorded and charted like any other night’s, so the trend reads straight through the failure.Nine auto-gateable metrics stay record-only and still only charted:
median_ttft_ms,p99_ttft_ms,p99_tpot_ms,median_e2el_ms,p99_e2el_ms,output_throughput,request_throughput,total_token_throughputandtokens_per_sec. So doesstep_time_ms.max, which was pruned deliberately rather than overlooked: it is what--perf-gatealways writes, it is not a recorded metric at all — no per-night series accumulates for it, it is synthesised at bless time from one night’s mean — and the measured step-0 compile excursion (2825 ms) clears any ceiling a clean night gives it.median_itl_msis absent for a different reason again:_NO_AUTO_GATEkeeps the refresher from ever writing it, because it measures ~0 here.Steps 1–6 below are done. Step 7 — promoting the nine on evidence — is not, and the nine are not waiting on the same thing. None of them is waiting on nights that have not happened yet — each group starts from evidence the completed window already holds, and a second window is a possible outcome for the three
p99_*metrics only if analysing the first shows it is needed:
record-only metric what step 7 is waiting for median_ttft_ms,output_throughputNothing. The evidence is in hand. Their only blocker was the step-0 compile excursion, and the completed 2026-09-08..09-17 window clears it (step 4). They are deferred to a separate PR so the first armed gate stays attributable, not deferred for data. p99_ttft_ms,p99_tpot_ms,p99_e2el_msAnalysis of a window, and possibly a second one. The blocker is that a p99 over 32 requests is the 32nd of 32 order statistics with no repeat measurement behind it. The completed window recorded them — the nightly harvests every allowlisted metric — but step 3 only asked for five metrics to be written down, so nobody has computed their spread. Read the existing ten nights first; another window is needed only if that spread is too wide to size a margin from. median_e2el_ms,request_throughput,total_token_throughput,tokens_per_secNo window will unblock these. They are held back on redundancy, not on evidence: each is determined by a metric already gated, so arming them adds reason lines to one event rather than detection. Promoting them needs the argument in the per-metric table to change, not more nights. Nothing else in the matrix changed: every other entry is unchanged and still correctness-only.
The reason it was a document rather than a commit is that there was no window to
derive thresholds from. A threshold derived from a single observation encodes
whichever night it was taken on, and the nightly then fails on the difference
between two healthy runs. That failure is worse than no gate: it trains everyone
to ignore the alert, and the first real regression arrives into a channel nobody
reads. The window has since been taken — 2026-09-08..09-17, ten nights, all
twenty cell-runs — and the two bounds it sized are in
config/ci/regression_baselines.yaml. The argument is kept because it is the
one step 7 still has to satisfy for the nine record-only metrics, and because it
is the procedure for the next workload’s first bless.
Measuring the cell rather than reasoning about it has since made that case stronger and the first bless smaller. One cell-run in thirteen carries a five-second compile excursion in its first measured step, which three of the four originally-proposed gates cannot survive — see the smoke cell section.
See ci-nightly-eval.md for how the nightly works in general; this only covers what is specific to serving.
Worth stating plainly, because the phrase suggests a setting somewhere and there
isn’t one. nightly_eval.py looks each (entry, cell) up in
config/ci/regression_baselines.yaml by the key entry::cell; a cell with no
key gets compare_to_baseline(harvested, None), which returns record. So an
entry added to the matrix is record-only by construction, and stops being
record-only the moment a refresh-baselines PR merges a key for it.
Two consequences that decide the shape of this rollout:
fail with or without a baseline — the harness is fail-closed by design, and
docs/ci-nightly-eval.md says so. An entry may therefore only be added to the
matrix once it can actually pass on the runner. Adding one that cannot is not
a soft landing; it is a nightly that is red every night. That constraint is
what shaped this entry: it declares needs_docker_daemon: true so that a
runner without the socket skips it rather than failing — see What blocked
this, and what remains.refresh_baselines.py rebuilds the whole baseline file and --perf-gate armed
every entry in it. Running it to bless serving would, in the same PR, derive
step-time ceilings for gpu_smoke, inference_offline, training_ddp,
training_fsdp, race and llm_determinism — sixteen cells, each from a
single observation, none of which anyone has variance data for. That is the
failure this document is about, inflicted on six unrelated workloads as a side
effect. --perf-gate-entry (added with this document) scopes it.Every number below is measured, on one gfx950 (MI355X) node against the image digest the recipes pin. There is no synthetic data here, and no more of it than this — that is the whole problem.
Two caveats that apply to every number in this section, both established later
in the document rather than assumed here. They were taken at warmup_steps: 1
and the recipe now sets 2, so they are rationale and not a baseline; and they
were taken on an MI355X, whereas the CI runner smci350-rck-g03-f16-12 is an
MI350X, so their absolute level does not transfer to the runner even though
their spread is the thing being argued about. See
Verification status.
Bring-up is the noisy one, and it is very noisy. From tokenspeed.md:
Startup to
/healthis the dominant cost and it is noisy: 189, 276, 285, 291, 316 and 319 seconds across six runs of the same recipe on the same node — a 1.7× spread with nothing changed between them. So do not treat a slow bring-up as a signal without repeats, and do not tunetimeout_per_trialclose to an observed number; the recipe’s 1800 s leaves deliberate headroom.
The multi-model sweep adds a seventh observation for the same 0.6B model at 379 s, and tokenspeed-serving.md declines to treat the column as a measurement at all:
Read the startup column as a floor, not a measurement: it is dominated by weight loading and Triton compilation against whatever the node’s caches already hold, which is why the smallest model here posts the largest number. It is reported because
ready_timeout_sechas to cover it, not because it scales with anything.
The four-model table is the evidence for that last clause: 379, 283, 289 and 328 seconds for Qwen3 at 0.6B, 1.7B, 4B and 8B. Bring-up does not even order by model size.
The compile cache is a second, larger excursion, and one mitigation stands between it and the metrics:
On Qwen3-0.6B the first bench invocation against a fresh server took 6.2s against 1.1s for every later one. Rolled into the metrics that one outlier dominates the mean step time and inflates TTFT tenfold (465ms vs 47ms), so a cell looks like a regression purely because it went first.
warmup_steps: 1 is supposed to discard that step. It does not always, and
that is the single most important number in this document — see the next
section, which measures it on the cell the staged entry actually runs. That
measurement is why the recipe now sets warmup_steps: 2.
Steady-state serving metrics, by contrast, reproduce well. The docs do not say this in one sentence, but they contain two independent repeats:
tokenspeed-serve-models.yaml::qwen3-0.6b and
tokenspeed-serve-load.yaml::conc-8 run a byte-identical measurement
configuration — Qwen3-0.6B, ISL 512 / OSL 128, concurrency 8, 32 prompts, 3
measured steps, 1 warmup step, ignore_eos, seed 0 — in two separate sweeps.tokenspeed-serve-gptoss.yaml::baseline and
tokenspeed-serve-gptoss-tp.yaml::tp1, of which the doc says “TP=1 reproduces
the single-GPU numbers above to within a percent, which is the control this
axis needs”.| Metric | Qwen3-0.6B, two sweeps | spread | gpt-oss-20b, two sweeps | spread |
|---|---|---|---|---|
median_ttft_ms |
46.3 / 45.9 | 0.87% | 67.1 / 67.2 | 0.15% |
median_tpot_ms |
1.94 / 1.91 | 1.57% | 7.61 / 7.63 | 0.26% |
output_throughput |
3538 / 3631 | 2.63% | 994 / 991 | 0.30% |
server_startup_sec |
379 / — | — | 316 / 199 | 59% |
The same node reports serving rates to within 3% across sweeps and bring-up time to within 59%. Those two facts are what the per-metric table below is derived from.
There is also a useful contrast from the kernel side of the integration, where the repo already distinguishes a tight metric from a loose one: the GEMM probe’s “spread across trials was under 1.5 µs in every cell”. Not everything TokenSpeed measures is noisy; bring-up is.
Everything above is a claim quoted from another document. This section is a
measurement of the exact cell the staged nightly entry runs, taken from the
step_times_ms and metrics_summary recorded in every matrix.json on disk for
TOKENSPEED-SERVE-SMOKE: six sweeps, 13 cell-runs with step times, 39 steps.
These numbers were all taken at
warmup_steps: 1, which the recipe no longer uses. The recipe now setswarmup_steps: 2, precisely because of what this section measures. So read everything below as the rationale for that change, not as the baseline to bless against: the table is the evidence that a step-0 excursion reaches the metrics atwarmup_steps: 1, and it is retained for that purpose. It was not a prediction of what the record-only window would show, because the window was taken at a different setting — one whose whole purpose is to remove the excursion these thirteen cell-runs contain. Nothing here is deleted or restated at the new setting. The measurement this was waiting for is the ten-night window atwarmup_steps: 2, 2026-09-08..09-17, and it lives in step 4 of the rollout sequence rather than here.
Twelve of the thirteen are very clean. Tighter than the cross-sweep numbers above, because these are the same recipe on the same node:
| Metric | Clean range over 12 cell-runs | Spread |
|---|---|---|
median_ttft_ms |
43.65 – 46.87 | 7.37% |
median_tpot_ms |
1.89 – 1.95 | 3.08% |
output_throughput |
3502.53 – 3646.20 | 4.10% |
p99_itl_ms |
34.00 – 35.84 | 5.43% |
| step time | 1108 – 1178 ms | 6.3% |
The thirteenth is the compile-cache excursion, and warmup_steps: 1 did not
catch it. serve-smoke3::no-scratch-reclaim recorded its three measured steps,
in order, as:
6193.4 ms, 1140.2 ms, 1142.6 ms
The excursion is the first measured step — after the warmup step was already
discarded. Its cell then reported median_ttft_ms 465.30 against a clean 43.65–
46.87, which is 10.26× the clean mean and the same 465 vs 47 the serving doc
quotes. So the 10× TTFT excursion is not a hypothetical a recipe edit could
introduce; it is present, once, in the data we already have, at 1 cell-run in
13 (7.7%).
What it does to the four gates the table below proposes, anchoring each on the worst of the twelve clean runs and applying the plan’s own margins:
| Gate | Anchor | Threshold | Excursion | Verdict |
|---|---|---|---|---|
step_time_ms.max |
1169.50 | 1461.88 | 2825.40 | fail |
median_ttft_ms |
46.87 | 58.59 | 465.30 | fail |
output_throughput |
3502.53 | 2977.15 | 2612.83 | fail |
median_tpot_ms |
1.95 | 2.44 | 1.93 | pass |
p99_itl_ms |
35.84 | 44.80 | 34.96 | pass |
Three of the four proposed gates fire on a run that is not a regression. Those
verdicts are compare_to_baseline’s, not arithmetic done here —
tests/ci/test_eval_lib.py runs the measured numbers through the real comparator.
Why the latency-per-token metrics survive it, and why that is the useful
signal. The excursion is a fixed ~5 s of compilation added to one step. It
therefore lands on anything derived from step duration — the step-time mean,
and throughput, which is tokens over that duration — and on TTFT, because the
first request of that step waits for the compile. It does not touch
median_tpot_ms (1.01× clean) or p99_itl_ms (1.00× clean), because those are
measured between tokens, after the compile has happened. A metric whose
definition excludes the excursion is immune to it by construction, not by luck.
This is a different failure from the concurrency-64 one, and the difference
decides the fix. The matrix work found ~8% of bench steps at concurrency 64
stalling by a fixed ~0.92 s with no positional pattern — the serve-load::conc-64
cell on disk reads 1621, 2286, 2595 ms, bimodal in a way no warmup setting can
remove. The smoke cell’s excursion is at position 0 every time it appears, which
makes it deterministically avoidable: it is cache warming, and one more
discarded step removes it. The two look alike in a summary statistic (both make a
three-step mean untrustworthy) and are opposite in what to do about them.
Finally, the mitigation A/B is consistent with all of the above. On gpt-oss the
hsa_no_scratch_reclaim cell landed at TTFT 66.7 vs 67.1 baseline, TPOT 7.49 vs
7.61, throughput 1008 vs 994 — the doc calls it “within noise of baseline, as it
was at every smaller size”, and the size of that “noise” is the 0.3–2.6% band
above, not the 59% one.
Direction is fixed by _METRIC_POLICIES in scripts/ci/eval_lib.py (max for
latencies, min for throughputs). What this table decides is which of them get a
bound on the first bless.
Only that. The verdicts below are about the nightly’s blessed baselines, which
are derived by applying a margin to an observed value, and they say nothing
about the workload’s own gates: block (_GATE_SPECS in
src/aorta/workloads/tokenspeed_serve.py), which enforces absolute numbers a
recipe writes out per trial. The two differ exactly where the margin does the
damage: max_p99_itl_ms: 50 is a stated ceiling on tail stalls and is a
perfectly good gate, whereas a baseline ITL ceiling is observed × 1.25,
which for an observation near zero is near zero. So “Never” in this table means
“never armed automatically from a baseline”, not “not gateable” — and a metric
listed here as record-only can still carry a hand-written per-trial bound today.
Revised by measurement. The four-gate set below was derived before the step-0 excursion was measured on the smoke cell. Three of those four fire on it. The
First blesscolumn now reflects that; theWhycolumn keeps the original reasoning, because it was not wrong about the noise — it was wrong aboutwarmup_steps: 1making the excursion unreachable.
| Metric | Policy | First bless | Why |
|---|---|---|---|
median_tpot_ms |
max | Gate | Reproduced to 1.57% and 0.26% across sweeps, 3.08% over 12 same-node cell-runs, and the docs already name it “the better-behaved per-token metric”. It is steady-state decode cost with no queueing term, which is why it is the tightest number we have — and it is measured between tokens, so the step-0 compile excursion does not enter it (1.01× on the excursion run). The one gate the measurement leaves standing. |
p99_itl_ms |
max | Gate | Promoted from record-only for the same reason: 1.00× on the excursion run, 5.43% clean spread. It is the useful half of the ITL pair, it catches tail stalls that a median cannot, and it is the only other metric whose definition excludes the excursion. Gating it and median_tpot_ms together covers per-token latency at both the centre and the tail without touching anything duration-derived. |
mean_step_time_ms (step_time_ms.max) |
max | Record-only (was: gate) | The bench step duration, and therefore the metric the excursion hits hardest: 2825 ms against a 1462 ms ceiling, 2.4× over. It is still the bound --perf-gate always writes, so arming it is the default — which is exactly why the rollout sequence below prunes it by hand. The window has since shown warmup_steps: 2 covers the excursion, and it stays pruned anyway: it is not a recorded metric, so no per-night series of it exists for a window to measure (step 4). |
output_throughput |
min | Record-only (was: gate) | Reproduced to 2.63% and 0.30%, and it is the headline number — but it is tokens over the step duration, so the excursion drags it to 0.73× clean and through a 0.85 floor. Nothing is wrong with the metric; it simply cannot be gated while a five-second compile can land inside the window it divides by. |
median_ttft_ms |
max | Record-only (was: gate) | The original entry said “if one gate flaps, expect it to be this one”, and that was right for the wrong reason — not a flap but a 10.26× excursion, 465.30 against a 58.59 ceiling. It carries queueing delay the others do not, and its first request is the one that waits for the compile. |
p99_ttft_ms |
max | Record-only | A p99 over 32 requests is the 32nd of 32 order statistics — effectively the maximum, and we have no repeat measurement of it. The load sweep shows it moving 194 → 2139 ms across shapes and 315 → 426 ms for a 2× concurrency change, so it is responsive to things a gate should not fire on. Promote on evidence from the record-only window. |
p99_tpot_ms, p99_e2el_ms |
max | Record-only | Same order-statistic argument, same absence of repeat data. (p99_itl_ms was in this row and has been promoted to the gate set above — it now has 13 same-node cell-runs behind it, not zero repeats, and it is one of the two metrics the step-0 excursion leaves alone.) |
median_e2el_ms |
max | Record-only | End-to-end latency is TTFT plus OSL × TPOT, so gating it adds a third alarm for an event two gates already catch, and its reason line is the least specific of the three. |
request_throughput |
min | Record-only | Under ignore_eos at fixed OSL it is output_throughput / output_len — fully determined by a metric already gated. |
total_token_throughput, tokens_per_sec |
min | Record-only | Restatements of output_throughput at fixed ISL/OSL (tokens_per_sec is documented as an alias of it). Gating them costs nothing but produces three reason lines for one event, which makes triage slower rather than safer. |
median_itl_ms |
max | Never | Measured at ~0, because the gateway delivers several tokens per SSE chunk. A margin is multiplicative, so an observation of 0.0 blesses a ceiling of 0.0 and every later run with any inter-token gap at all fails. eval_lib already warned about this in a comment; it is now enforced by _NO_AUTO_GATE. |
server_startup_sec |
— | Never | 189–379 s in the docs; 180–415 s measured over the 13 smoke cell-runs, so the spread is wider than the quoted one, not narrower. It does not order by model size. Not in the allowlist and must stay out: the allowlist is what --perf-gate arms from, so adding it is gating it. |
container_elapsed_sec |
— | Never | Dominated by bring-up; same argument. |
duration, total_input_tokens, total_output_tokens |
— | Never | Work-done counters. The token totals are pinned by the recipe, so a bound on them restates the configuration rather than measuring the stack. |
completed_total, failed_total |
— | Never | Already enforced, harder, elsewhere: the bench script and the workload independently require failed == 0 and completed == num_prompts, so a shortfall fails the cell and reddens the nightly with no baseline involved. A metric bound here would be a third and weaker copy of a check that already fails closed. |
max_output_tokens_per_s, max_concurrent_requests |
— | Never | Single-sample maxima — the noisiest available summary of a distribution. |
mean_*_ms, std_*_ms, p50_*, p90_* |
— | Never | Deliberately absent from the allowlist already (“means are the noisiest summary of a latency distribution and the least useful thing to gate on”). Keep them absent. |
Two gates, then — median_tpot_ms and p99_itl_ms, the per-token pair — and
everything else charted, including three metrics that would have been gated
before the excursion was measured. That is a deliberately smaller first bless
than the plan originally proposed: the two that remain are the two whose
definitions exclude the failure mode we can actually demonstrate. Of the
dropped ones, median_ttft_ms and output_throughput can be promoted from the
record-only window once it shows the excursion is gone, which the
2026-09-08..09-17 window has (step 4).
step_time_ms.max does not follow them: it is not a recorded metric, so no
window measures it. The gap between “gated” and
“invisible” is covered by the dashboard’s What changed view, which reports any
metric that moved more than 10% between the two most recent runs without failing
the job. A 10% throughput drift is real, is not a gate breach under these
margins, and is exactly what that view exists to surface.
Take ten nightlies before blessing. Derive each bound from the window’s
extremum — the maximum for a max metric, the minimum for a min metric — not
from its mean, and not from one night’s observation.
The statistic matters more than the count, so take that first. Deriving from the mean is the intuitive choice and it is wrong here, because the variation we have measured is not jitter around a centre. The startup series is 189 against a cluster of 276–379: a single environmental excursion, not a spread. A mean is precisely the statistic that hides one, and a bound built on it is breached by the next occurrence. The extremum is the statistic that asks the question we actually care about — how bad has a healthy night ever been?
That also fixes the current tooling’s real defect, which is not the margin size
but the anchor. refresh_baselines.py --perf-gate derives value × 1.25 from
whatever single run was under way, so the bound depends on which night the
operator pressed the button. Feeding the seven known startup observations through
compare_to_baseline shows what that costs:
| Blessed on | Ceiling (×1.25) | Later runs breaching |
|---|---|---|
| 189 s | 236.2 | 6 of 6 |
| 276 s | 345.0 | 1 of 6 |
| 285 s | 356.2 | 1 of 6 |
| 291 s | 363.8 | 1 of 6 |
| 316 s / 319 s / 379 s | 395.0 / 398.8 / 473.8 | 0 of 6 |
Four of the seven possible blessing nights produce a gate that breaches, and one of them reddens every subsequent run. The window maximum (379 → 473.8) breaches none — but only because the window contained the excursion, which is the argument for the count.
Ten is chosen from that, not from convention. Scoring every n-run subset of the series against the runs it did not contain gives the residual breach rate on an unseen night:
| Window size | Mean breach rate on unseen runs | Worst window |
|---|---|---|
| 1 | 21.4% | 6 breaches |
| 2 | 5.7% | 1 breach |
| 3 | 2.9% | 1 breach |
| 4 | 1.0% | 1 breach |
| 5 and up | 0% | none |
The nightly runs roughly 250 times a year, so one false alarm per quarter needs a per-run breach rate under about 1.1% — which this series reaches at n=4, on the noisiest metric the workload produces. The metrics we are actually gating are 10–16× tighter than that one. Zero at n=5 is an artefact of a seven-point sample rather than a real floor, so the honest reading is that the risk is small by 5 and the remaining reason to go further is calendar coverage: ten nightlies is two full weeks, long enough to contain a runner reimage, a Docker Hub re-pull or a cold HF cache — the once-a-week-ish events that produce excursions in the first place. Five is the floor; below it the breach rate on known-noisy data is 2.9% per run, about nine false alarms a quarter.
An earlier version of the paragraph above listed “a Dependabot ROCm bump” beside the runner reimage and the cold cache, as one more thing ten nights is long enough to contain. That was wrong, and it is worth correcting rather than quietly dropping, because the two categories look alike and are opposites.
A reimage, a re-pull and a cold cache are environmental events: they perturb a run and the stack underneath is the same before and after, so a window that contains one has measured a genuine bad night and the extremum is doing exactly its job. A digest bump is not that. It replaces the thing being measured half-way through measuring it, so the window no longer describes one population. You can sample a stack bump or derive a stable ceiling across the window, but not both — the extremum then answers “how bad has a healthy night been on either of two stacks”, which is a number about neither.
The counter-example is not hypothetical, and it happened while this branch was
open. sanitizers-nightly.yml went red on 2026-09-02 and stayed red, because
#411 moved Dockerfile.ci-gpu from
ROCm 7.2.4 to ROCm 10 and the f32 GEMM code objects the gate scans are extracted
from the image’s own Tensile bundle — so the committed expectation was describing
objects the image no longer ships
(#453). That is a correctness
baseline invalidated by a mid-flight digest change. A perf baseline is not
better protected; it is worse, because a shifted number still looks like a
number and nothing about it announces that the stack moved. A mid-window bump
would do to the perf baseline precisely what that one did to the sanitizer
baseline.
Operational condition, for the duration of the window: hold any PR that
changes either pinned digest. Concretely, the FROM … @sha256 in
docker/Dockerfile.ci-gpu (currently
rocm/pytorch:rocm10.0_ubuntu26.04_py3.14_pytorch_release_2.13.0@sha256:3174cb70…)
and the engine digest in
recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml
(lightseekorg/tokenspeed-amd@sha256:60c12e37…). If one has to land, restart
the count; do not average across it.
This is preventable by policy rather than by luck, which is the useful part.
.github/dependabot.yml runs the docker ecosystem against /docker on a
weekly schedule and its own comment says the job “bumps the FROM ... @sha256
digests”, so an in-window bump is not merely possible but scheduled. There is
no auto-merge — the same file says so explicitly, and a human therefore has
to approve and merge one for it to reach the window. Ten nights is long enough
for roughly two of these to be opened, so the hold is a real decision someone
will be asked to make twice, not a theoretical one.
Checked at the time of writing: the only open PR in the docker ecosystem is
#309 (bump ubuntu from 22.04 to
26.04 in /docker), and it is clear — it touches
Dockerfile.rocm-ubuntu-ebpf, Dockerfile.rocm70_2-ubuntu-nan and
Dockerfile.rocm70_2-ubuntu-pytorch, none of which is Dockerfile.ci-gpu, and
it does not move the engine digest. The other open dependabot/* PRs (#310–#314)
are github-actions bumps. So nothing currently open needs holding; re-check
before the window starts, since the ecosystem is on a weekly timer.
Everything else in the ten-night justification stands. The residual-breach-rate analysis is unaffected — it is a statement about sampling a single population, which is what this condition exists to preserve — and so is the calendar-coverage argument for the three environmental events it still lists.
That whole derivation assumes the window’s extremum is a healthy worst case. The step-0 excursion breaks the assumption, and it is worth being explicit about why, because it is the reason the gate set shrank rather than the window growing.
With a bimodal cell the ten-night window either contains an excursion or it does not, and both outcomes are bad:
| Window | median_ttft_ms anchor |
Ceiling | Consequence |
|---|---|---|---|
| No excursion (12 of 13 nights) | 46.87 | 58.59 | The excursion, when it lands, reads as a 10× regression. False alarm. |
| Contains one (1 in 13) | 465.30 | 581.63 | A genuine 2× TTFT regression — 47 → 94 ms — passes comfortably. No detection power at all. |
Enlarging the window does not resolve this; it only makes the second row more likely. The extremum is the right statistic for a unimodal metric with occasional environmental excursions, which is what the startup series is. It is the wrong statistic for a metric with two modes, because “the worst a healthy night has been” is not a single number any more.
So the correct response to a bimodal cell is not a cleverer threshold. It is
either to gate a metric the second mode does not reach — which is what
median_tpot_ms and p99_itl_ms are — or to remove the second mode. For the
step-0 excursion the second option is genuinely available, because the excursion
is positional: raising warmup_steps from 1 to 2 discards it by construction.
That has now been done — the recipe sets warmup_steps: 2 — which is what
lets median_ttft_ms and output_throughput be promoted later.
(step_time_ms.max, the third metric the excursion blocked, stays pruned for a
reason warmup_steps does not touch; see step 4.) The cost was that
it changed the measurement: every number taken before the change was at
warmup_steps: 1, so none of them could be a baseline, and the ten-night
record-only window had to be taken afresh at the new setting before anything was
blessed. It was: the 2026-09-08..09-17 window is at warmup_steps: 2, and the
two gates are blessed from it (step 4). That was a
deliberate trade. Carrying the excursion into the window instead would have
meant either blessing a bimodal cell or spending ten nights establishing a
distribution we already intended to change.
For the concurrency-64 stall no such fix exists — it has no position to discard — which is why that cell stays out of the nightly entirely rather than being gated on a narrower metric.
The margins themselves need no change. --step-time-margin 0.25 and
--throughput-margin 0.15 are already right for these metrics, and the
simulation says so: against the measured cross-sweep spread, every one of the
twelve bless-one-run-check-the-other combinations passes. Against the window
extremum the separation is wide — worst observed TPOT noise 1.57% versus a 25%
detection threshold, a factor of 16 — while a regression the size of a genuine
stack change is caught:
| Scenario | median_tpot_ms vs 2.425 |
output_throughput vs 3007.3 |
|---|---|---|
| the other measured night | 1.910 → pass | 3631 → pass |
| +10% / −10% | 2.134 → pass | 3184 → pass |
| +30% / −25% | 2.522 → fail | 2654 → fail |
| +94% / −46% (the 0.6B → 4B step) | 3.770 → fail | 1905 → fail |
Those rows are in tests/ci/test_eval_lib.py, run against the real comparator,
so the claim is checked rather than asserted.
Nothing blocks the first bless any more: the matrix entry is live, the socket is
signed off for the nightly lane, nightly-eval.yml sets docker_socket: true,
and the ten-night measurement is taken (2026-09-08..09-17, all ten scheduled
workflow_run events, all ten green). The two ceilings that window sized are in
config/ci/regression_baselines.yaml.
What remains is step 7, and not one of the nine
record-only metrics is blocked on a measurement this window failed to take.
Two of them (median_ttft_ms, output_throughput) have their evidence already
and are deferred only so the first armed gate stays attributable; three are
waiting on the analysis of nights already recorded; four are held back on a
redundancy argument no number will settle. The table in the
current-state box at
the top says which is which — read it before opening a step-7 PR, because
“wait for another window” is the first answer for none of the nine.
This section is kept because the shape of the plumbing is what the sign-off was given against, and because the argument is worth being able to re-read.
nightly_eval.py runs inside the aorta-ci-gpu container
(eval-reusable.yml → docker_cmd.sh exec), and tokenspeed_serve runs the
TokenSpeed engine in a container of its own. So the entry needs a Docker client
and a route to a daemon in there. It had neither, and the workload’s setup()
raised 'docker' not on PATH before anything else happened. Record-only does not
help — it defers perf bounds, not failures.
Both pieces now exist, the image has been built, and a cell has been run to completion from inside the container (see Verification). The client half is therefore done and proven. The daemon half is a per-lane opt-in, granted to the nightly lane and to nothing else — see Security.
The entry lives in entries with needs_docker_daemon: true, which is what
keeps that a configuration choice rather than a load-bearing one: on any lane
without the socket the entry skips with the reason recorded, rather than failing.
That still matters after the sign-off, because it is what keeps the socket
scoped to one lane — refresh-baselines.yml, bump-validate.yml and every
other consumer of the matrix can run without it.
tests/ci/test_nightly_eval.py covers both halves on the nightly path, and
refresh_baselines.build_baselines honours the same flag, with both halves
covered too.
1. A Docker client in the CI image. docker/install_docker_cli.py, invoked
from Dockerfile.ci-gpu, fetches the pinned static tarball, checks its sha256
before extracting, and extracts exactly one member — docker/docker. The same
tarball ships dockerd, containerd and runc; naming the member is what keeps
a daemon out of the image, and a build-time check fails the build if any of them
reaches PATH. The engine container is a sibling, started by the host daemon
next to aorta-ci-gpu, not nested inside it.
2. A route to the daemon, opt-in per lane. docker-compose.docker-socket.yaml
is an override, which is the mechanism the base compose already documents for
optional mounts (“Do not add a volume here”), so the default container still has
no route to the daemon. rocm-ci-setup takes a docker-socket input, default
false, and adds the override’s -f only when it is true.
That input is reachable from a caller workflow through a matching docker_socket
workflow_call input on eval-reusable.yml, also defaulting to false and
forwarded to the setup step. The default is what matters here: eval-reusable.yml
is shared by the nightly and by bump-validate.yml, so the choice has to be the
caller’s rather than the reusable workflow’s — setting it centrally would grant
the socket to a PR-triggered lane as a side effect of enabling the nightly.
bump-validate.yml therefore pins it false explicitly, and
sanitizers-nightly.yml never sees it, using rocm-ci-setup directly with no
docker-socket argument. nightly-eval.yml is the only lane that sets it
true — see the security note below for the sign-off and its basis.
3. work_dir must be an explicitly configured shared path. This is the detail
most likely to be missed, because nothing reports it as a path problem.
With the socket mounted, the -v sources in the docker run the workload builds
are strings the host daemon resolves against the host filesystem. The
workload builds them from work_dir — <work_dir>/u<uid>/{scripts,out,hf} — using
paths as seen from inside aorta-ci-gpu. Those are different namespaces. A
missing bind source is not an error to the daemon: it creates the directory on
the host and mounts that. So the engine container comes up with an empty
/ts-scripts and dies on a missing script, or — worse — with an empty writable
/ts-out, and the harvest reports no results for a run that really happened.
The default work_dir does not fix this by being /tmp/ts-work-serve on both
sides. /tmp inside the CI container is the container’s own /tmp, not the
host’s, so the two are different directories that share a name — which is the
failure above with the confusing property that the path looks right in every log.
What the nightly needs, concretely:
${TS_SERVE_WORK_DIR:-/tmp/ts-work-serve} at
the same path inside the container. Source and target are identical on
purpose, and a test asserts they stay identical; that is what makes one string
name one directory on both sides of the boundary.work_dir to that same path explicitly. It happens to equal
the workload default, but a default that must agree with a mount is a
coincidence, not a configuration: change either one alone and the run breaks
quietly.Measured, and more specific than the above was. Demonstrating this on a gfx950 node produced two failures worth writing down, because neither is what the guidance as written would have led you to expect.
Every bind source must be resolvable by the daemon, not only work_dir. The
first attempt put AORTA_WORKSPACE on an autofs-mounted NFS home and the
container never started:
Error response from daemon: error while creating mount source path
'/home/.../wt-gating': mkdir /home/...: permission denied
The daemon resolves all -v sources in the host namespace, so a checkout on an
automounted or root-squashed filesystem fails the same way a work_dir there
would — and it fails at container create, before any of the workload’s own
checks can produce a better message. On a GitHub runner the workspace is local
disk and this does not arise; on any node where /home is networked, the repo
has to be staged locally first. Worth knowing before debugging it as a socket
problem, which is what it looks like.
The work root must be owned by the uid the container runs as, and the obvious
way to create it gets that wrong. aorta-ci-gpu runs as root, so the workload
checks u0 and requires the root itself to be root-owned. A work_dir created
by an ordinary host-side run is owned by that user, and the containerised run
then refuses it — the node used here already had a /tmp/ts-work-serve at
uid=100550 mode=1777 left by earlier host-side sweeps, which uid 0 must reject.
The demonstration used a fresh path created through the daemon so it landed
root-owned, which is exactly what the workload’s own error message advises an
administrator to do.
Creating it is not the same as repairing it, and the runner needs the second
one. mkdir -p is a no-op on a directory that already exists: run as root over
a leftover user-owned root it returns 0 and changes neither owner nor mode, so
the next run fails the same ownership check with the remedy apparently already
applied. Say what the end state is instead — the sequence below is idempotent
whether or not the directory is there, and is what to run on the runner before
the nightly first executes the entry:
docker run --rm -v /tmp:/mnt busybox:1.37 sh -c \
"mkdir -p /mnt/ts-work-serve && chown 0:0 /mnt/ts-work-serve && chmod 1777 /mnt/ts-work-serve"
docker run --rm -v /tmp:/mnt busybox:1.37 stat -c "%n %u:%g %a" /mnt/ts-work-serve
chown the root and nothing under it: the per-uid scratch directories beneath
it belong to whichever uid created them, and -R would take those from a
co-tenant. The second line is worth running — the check the workload makes is on
the root’s owner, so 0:0 1777 is the whole acceptance criterion.
The operational trap in that is worth stating on its own: the ten record-only
runs cannot share a work_dir with the containerised nightly if they are taken
by hand as an ordinary user. Either take them as root, or give the two a separate
work_dir and accept the cold HF cache on the first nightly.
The uid has to agree too, and this is a second way the same boundary bites.
The workload uses a per-uid scratch root and then verifies it: <work_dir>/u<uid>
must be a real directory (not a symlink), owned by the running uid, and not group-
or world-writable. Inside aorta-ci-gpu the process is root (user: root in
the base compose), so it writes and checks u0. That works — a root-owned
directory created through the shared mount is root-owned on the host too, so the
host daemon and the checking process agree. But it holds only because both sides
are uid 0. Change user: in the base compose, run the harness as a non-root uid,
or turn on userns-remap on the daemon, and <work_dir>/u<uid> is created by one
uid and inspected by another: the check fails with an ownership error that does
not mention containers, uids-across-a-boundary, or the mount. Anyone changing
either should expect that error to be the symptom.
One consequence worth stating: with run_as_current_user defaulting true, the
engine container also runs as uid 0, so exports land on the host owned by root,
outside the workspace that eval-reusable.yml’s “Reclaim results ownership” step
chowns. They are cleaned per trial unless keep_work_dir is set, but a runner
that fills /tmp with root-owned scratch is a plausible future complaint.
Signed off for the nightly lane on 2026-09-03, and nightly-eval.yml sets
docker_socket: true accordingly. The reasoning is recorded here rather than
only in a review thread, because it is the durable part: a future reader deciding
whether a second lane may have the socket needs the argument, not the verdict.
Stated plainly first, because it is the part most easily lost in a diff: a
process that can talk to the Docker daemon can ask it for a privileged container
that bind-mounts /. That is root on the host, with no exploit involved. So the
mount is genuinely root-equivalent, and nothing below disputes that.
What it is not is a new capability on this runner. Before any container
exists, eval-reusable.yml’s first step runs
docker run --rm -v "$ws":/ws "$busybox" chown -R "$uidgid" /ws
as a plain host step, falling back to sudo -n docker info and then to sudo
chown -R. refresh-baselines.yml opens the same way. So the runner account
already has direct daemon access and passwordless sudo, and any job step that
can execute on this runner is already host-root capable — with or without the
socket. The socket makes that reachable from inside aorta-ci-gpu as well as
from the job shell, which is a convenience, not a new privilege. Reasoning about
it as though the container were the boundary overstates what the container was
ever doing.
What bounds the risk is write access to the repository, and that is checkable. Three things, each of which can be re-verified by grep rather than taken on trust:
pull_request_target anywhere in .github/ — so no workflow
runs fork code with repository secrets or on a privileged trigger.gpu-tests.yml (twice) and bump-validate.yml gate on
github.event.pull_request.head.repo.full_name == github.repository. A fork PR
therefore never reaches this runner at all.gpu-tests.yml does not
go through eval-reusable.yml and passes no socket argument; bump-validate.yml
pins docker_socket: false explicitly rather than inheriting a default that a
later edit could flip. sanitizers-nightly.yml calls rocm-ci-setup directly
and passes nothing.So the population that can reach the daemon through this mount is exactly the
population that can push a branch to this repository — which is the population
that can already run sudo on the runner through any workflow step. That is the
whole of the marginal risk, and it is close to zero.
The “single-tenant runner” framing is withdrawn, because it was doing work
the argument does not need and cannot currently support. ROCm/aorta and
ROCm/aorta-internal hold separate repo-level runner registrations
(smci350-rck-g03-f16-12 and smci350-internal), and the internal
registration’s July jobs reported machine sharkmi300x-3 — an MI300X box —
despite carrying an mi350x label. Whether the two registrations share a
physical host is not confirmed as of today. It does not need to be: the
argument above rests on what the runner account can already do and on who can
reach it, neither of which changes with tenancy. Co-tenancy would matter for
measurement (a neighbour’s load moving the numbers), which is a separate
question handled by the ten-night window and the two-cell control, not by this
grant.
The alternatives, kept for the case where a future lane wants the socket and
this argument does not carry over. Run the serving workload outside the CI
container entirely — as a step on the runner host, which already has a client and
the socket, publishing its results JSON into the harness’s results directory. A
socket proxy allowlisting just the container create/start/logs/remove calls is a
middle option; rootless Docker is a third, and it is the only one that is a real
reduction rather than a re-arrangement, precisely because of the sudo point
above. Nested docker-in-docker is not on the list: it needs privileged too, so
it trades no privilege away, and it gives the engine a different filesystem view,
which breaks exactly the bind mounts described above.
Nothing in the plumbing. Two operational preconditions before the entry first runs, both on the runner rather than in this repository:
/tmp/ts-work-serve must exist root-owned and 1777 — the sequence is
under work_dir above, and note that mkdir -p alone
does not repair a pre-existing user-owned root.docker/Dockerfile.ci-gpu and the engine image in the recipe. See
the stack-bump section.The entry is in entries rather than pending_entries, carrying
needs_docker_daemon: true, so on any lane without the socket — the baseline
refresher, bump-validate — it skips with the reason in the results instead
of failing on Cannot connect to the Docker daemon. Promoting it without that
flag would have been the mistake the matrix file’s own header warns about: an
entry may only be added once it can actually pass on the runner.
Built and run on a gfx950 (MI355X) dev node on 2026-09-02 — not on the CI runner, which is an MI350X. Read the hardware note further down this section before treating any absolute number here as something the nightly should reproduce.
The image builds and is client-only. docker compose --env-file .env.ci -f
docker-compose.build.yaml build succeeds, producing a 51.9 GB aorta:ci-gpu.
The build’s own RUN checks pass — docker client check: OK --
/usr/local/bin/docker, no daemon on PATH — and the assertion holds when
re-checked against the built image rather than during it: dockerd,
containerd, runc, containerd-shim-runc-v2, docker-proxy and
docker-init are all absent from PATH, and docker --version reports
29.7.2. That is what makes “client only” a property of the artifact and not
of the Dockerfile’s intent.
The client reaches the host daemon from inside the container. With the
override mounted, docker exec aorta-ci-gpu docker ps returns the running
container list, exit 0.
A cell completes. Two sweeps of tokenspeed-serve-bench-smoke.yaml were
driven from inside aorta-ci-gpu, each starting the TokenSpeed engine as a
sibling container: four cell-runs, all passed, failed_total 0 and
completed_total 96 throughout, engine containers cleaned up afterwards. The
twelve measured steps were 1108–1149 ms with no step-0 excursion, and every
metric landed within 5% of the host-side clean envelope — so the container
boundary does not move the measurement, which is what lets a nightly baseline be
compared against the host-side runs the variance analysis is built from.
[!IMPORTANT] The 1108–1149 ms envelope is not a night-one acceptance check. It was measured on an MI355X dev node. The CI runner
smci350-rck-g03-f16-12is an MI350X — same gfx950/CDNA4 ISA, but a lower power budget and air rather than liquid cooling. Correctness is unaffected by that difference: the ISA is identical, so the kernels, the compiler output and every verdict the harness produces are the same. Absolute step times are a different matter and will plausibly differ — a priori, at least; one MI350X cell has since been measured and did not show an offset, which narrows the expectation without licensing the envelope as a check. See the subsection below.So do not use this envelope to accept or reject night one, in either direction. A first nightly landing outside 1108–1149 ms is not evidence of a problem, and one landing inside it is not evidence that the lane is healthy.
What has to reproduce across the window is the spread, not the level. The ceilings are derived from the ten CI nights themselves —
max × 1.25andmin × 0.85over that window, on that runner — so a systematic offset between the dev node and the runner is absorbed by construction and is expected, harmless and not worth investigating. What would matter is the shape: a window spread materially wider than the measured few percent, or a step-0 excursion, and step 4 already checks both. Judge night one against the two checks in step 4 and against the previous nights on the same runner, never against a number taken on other hardware.
warmup_steps: 2 (2026-09-03) — a data point, not a spreadBoth caveats above — MI355X hardware, and warmup_steps: 1 — are now partly
addressed by one measurement, and it is recorded here with its limits stated
first because it is easy to over-read.
This is a single cell-run. It is not a variance estimate and must not be used
as one, and it does not replace anything above. The
13-cell-run table
stays exactly as it is — it remains the rationale for warmup_steps: 2, and one
cell cannot restate it. The ten-night window is still the only thing that
produces a spread at the new setting.
Run on cv350-rck-g03-c16-18 (Slurm partition meta64), which is an
MI350X: 1000 W package power cap, VBIOS 113-M350-01-1K5-000C, gfx950,
288 GiB HBM. The 1000 W cap and the air-cooled VBIOS are what identify it as an
MI350X rather than an MI355X, and it is the same hardware class as the CI
runner — which is the point. Deliberately not on
smci350-rck-g03-f16-12, which is reserved for the baseline window. One
baseline cell of tokenspeed-serve-bench-smoke.yaml with every
measurement-relevant setting as committed (image digest, Qwen3-0.6B, ISL 512 /
OSL 128, concurrency 8, 32 prompts, 3 measured steps, num_warmups: 1,
warmup_steps: 2, ignore_eos, seed 0), on a private work_dir so it cannot
touch the CI lane’s. Passed, failed_total 0, completed_total 96.
| Metric | MI350X, 1 cell, warmup_steps: 2 |
MI355X clean range, warmup_steps: 1 |
|---|---|---|
| step times, in order | 1135.03 / 1136.39 / 1136.51 ms | 1108 – 1178 ms |
median_ttft_ms |
45.00 | 43.65 – 46.87 |
median_tpot_ms |
1.910 | 1.89 – 1.95 |
p99_itl_ms |
34.64 | 34.00 – 35.84 |
output_throughput |
3605.7 | 3502.53 – 3646.20 |
server_startup_sec |
282 | 180 – 415 |
Two things worth taking from it, and one worth not taking.
The MI350X level is not detectably different at this workload size. Every metric lands inside the MI355X clean range, step time included. That is a plausible result rather than a surprising one: Qwen3-0.6B at concurrency 8 is a small, latency-bound load that does not approach either card’s power ceiling, so the MI355X’s higher budget and liquid cooling have nothing to bite on. It does not license using the MI355X envelope as an acceptance check — one cell cannot establish agreement, and the rule above stands unchanged — but it does mean a large night-one offset would itself be worth a look, rather than being shrugged off as expected hardware difference.
No step-0 excursion, with the excursion’s own signature absent. The first
measured step is the fastest of the three, so there is no positional outlier at
all. This is the first observation at warmup_steps: 2 and n=1, so it is
consistent with the change working and nowhere near sufficient to conclude it;
step 4’s bimodality check over twenty cell-runs is what settles that.
What not to take from it: the three step times span 0.13%, and that number is not comparable to the 6.3% step-time or 3.08%/5.43% metric spreads in the table above. Those are across twelve cell-runs — separate server bring-ups on separate invocations, which is where the variance lives. 0.13% is three consecutive steps against one already-warm server, and it is the within-cell figure that the earlier table never reported. Reading it as a 48× variance improvement would be a category error. Re-taking the spread needs the window.
Two configuration requirements were discovered by doing this rather than by reading, and are written up under work_dir above: all bind sources must be resolvable by the daemon (an autofs NFS checkout is not), and the work root must be root-owned because the container runs as uid 0.
Also verified earlier, and still true: the installer rejects a wrong sha256
without leaving a binary behind; the client negotiates against an older (29.1.3)
daemon; docker compose config resolves all three mounts with the scratch mount
identical on both sides.
Still not verified, and it is now the first thing the nightly will establish:
the entry has never run under nightly_eval.py on the CI runner, only under
aorta sweep run in the container by hand. Night one is therefore a bring-up
observation as much as a measurement — treat a failure there as plumbing until
shown otherwise, and see the two operational preconditions under
What is still missing before counting it. The skip
path is covered by tests rather than by a runner.
To reproduce:
cd docker
bash ../scripts/ci/docker_compose.sh --env-file .env.ci -f docker-compose.build.yaml build
docker run --rm aorta:ci-gpu docker --version
# The work root must be root-owned, and `mkdir -p` on its own does not make it
# so: on a node where an earlier host-side sweep left a /tmp/ts-work-serve owned
# by that user, `mkdir -p` as root sees an existing directory and returns 0
# having changed nothing -- so the run still fails the ownership check, with the
# remedy apparently already applied. State the owner and the mode explicitly;
# the sequence is idempotent whether or not the directory is there.
#
# `chown` on the root only, never `-R`: the per-uid scratch beneath it belongs
# to whoever created it, and the check the workload makes is on the root's owner.
docker run --rm -v /tmp:/mnt busybox:1.37 sh -c \
"mkdir -p /mnt/ts-work-serve && chown 0:0 /mnt/ts-work-serve && chmod 1777 /mnt/ts-work-serve"
export TS_SERVE_WORK_DIR=/tmp/ts-work-serve
bash ../scripts/ci/docker_compose.sh --env-file .env.ci \
-f docker-compose.build.yaml -f docker-compose.docker-socket.yaml up -d
docker exec aorta-ci-gpu docker ps
docker exec aorta-ci-gpu aorta sweep run \
--recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml \
--output-dir "$TS_SERVE_WORK_DIR/out" --strict
Steps 1–6 are done: the entry is live, the nightly lane has the socket, the
ten-night window was taken 2026-09-08..09-17, and the two bounds it sized are in
config/ci/regression_baselines.yaml. Step 7 is the only one outstanding.
The steps are kept written as instructions rather than rewritten as history: they are the procedure for the next bless (step 7 here, and any other workload’s first bless), and what each one was checked against is the part worth being able to re-read.
1. Promote the entry. Done. It is in entries with min_gpus: 1,
timeout_sec: 3600 and needs_docker_daemon: true.
timeout_sec: 3600 rather than the 1800 default: two bring-ups at the observed
spread plus the teardown VRAM drain is around 15 minutes before the image pull
and any cold-cache weight download, and a timeout is an unconditional entry
failure rather than a slow record. Check the job’s own timeout-minutes: 150 in
eval-reusable.yml still has room. Measured end to end from inside the
container, a full two-cell sweep took 12 minutes on a warm cache, so the
budget is right but most of it is bring-up.
2. Verify it before merging. Done, and rerun these after any change:
python -m pytest tests/ci -q --timeout=180
python -m pytest tests/workloads/test_tokenspeed_serve.py -q
aorta sweep run --recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml --dry-run
2a. Enable the socket on the nightly lane. Done, signed off 2026-09-03
— see Security for
the basis. It is one line in .github/workflows/nightly-eval.yml:
uses: ./.github/workflows/eval-reusable.yml
with:
docker_socket: true
The docker_socket input on eval-reusable.yml defaults to false and is set
per caller, so this arms the nightly lane and nothing else. bump-validate.yml
sets it false explicitly: that lane runs PR head code on the same self-hosted
runner and must keep no route to the host daemon. sanitizers-nightly.yml does
not go through eval-reusable.yml at all — it uses the rocm-ci-setup action
directly and passes no docker-socket, so it inherits the action default of
false. Setting the input inside eval-reusable.yml instead of per caller would
hand the socket to every lane at once and is the mistake this shape exists to
prevent.
The grant is per lane, and it should stay that way. Any other lane wanting the
socket is a fresh decision, and the argument recorded under Security does not
transfer to a lane whose trigger is reachable from a fork or from an unreviewed
branch — that argument is what bounds this one. The dashboard also makes the
absence visible rather than silent: on a lane without the socket the entry
reports skip with needs a docker daemon: ... in its reasons.
3. Let it record for ten nightlies, at Done —
2026-09-08 to 09-17, ten consecutive scheduled warmup_steps: 2.workflow_run events, every one
green, no workflow_dispatch inside the range. While it ran it reported
recording on the dashboard.
The window is only valid at the setting the gate will run at, and that setting is now
warmup_steps: 2inrecipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml. Confirm that before counting nights, and restart the count if it changes underneath the window. None of the numbers earlier in this document can substitute for any of these ten: they were measured atwarmup_steps: 1, and the change was made specifically to alter the behaviour they describe. Expect the window to be cleaner than the 13-cell-run table — that is the change working — but do not assume it; the point of the window is to measure it rather than predict it.
And do not accept or reject night one against any number in this document. All of them were taken on an MI355X dev node; the runner is an MI350X. A systematic offset in the absolute level is expected and harmless, because the ceilings come from these ten nights on this runner. It is the spread that has to reproduce — see the hardware note under Verification status, and the single MI350X data point recorded there.
Also confirm before counting that neither pinned digest is about to move: an in-window bump invalidates the window rather than being absorbed by it, which is its own section and the reason for the digest hold.
Watch the Workloads view and write the numbers down; the ten values
of median_tpot_ms and p99_itl_ms per cell are the input to step 4, and the
ten of median_ttft_ms and output_throughput are what decides whether those
two record-only metrics can be promoted later. The ten of mean_step_time_ms
are what step 4 reads to tell whether the excursion is gone. Two cells, so
twenty observations of each. A fail during this window is a real failure —
the entry is unbaselined but the harness is fail-closed — and must be fixed
rather than waited out.
4. Check the window before blessing, and check it for bimodality first.
Done, and both checks passed — with one number that looks alarming and is not,
recorded below the checks. Compute, per cell and per metric, the extremum and
the ratio of extremum to median. Two separate checks:
median_ttft_ms and output_throughput
stay record-only regardless of how good their spread looks — a clean ten-night
window over a bimodal cell is the exact situation where the extremum anchor
gives a threshold with no detection power. See
the bimodal-cell section.What the 2026-09-08..09-17 window actually showed, per cell and per gated metric:
| cell | metric | min | median | max | max/median | full range |
|---|---|---|---|---|---|---|
baseline |
median_tpot_ms |
1.7373 | 1.8906 | 1.9138 | 1.0122 | 9.33% |
baseline |
p99_itl_ms |
34.1259 | 34.7533 | 35.4296 | 1.0195 | 3.75% |
no-scratch-reclaim |
median_tpot_ms |
1.8604 | 1.8945 | 1.9196 | 1.0132 | 3.13% |
no-scratch-reclaim |
p99_itl_ms |
33.9241 | 34.9653 | 35.8309 | 1.0248 | 5.45% |
The 9.33% is not the spread-check failing, and the reason generalises. Both metrics are
maxpolicy, so the number that sizes a ceiling and the number that can breach it is the upside deviation, not the full range. That 9.33% is a single low outlier on 09-13 — the cell ran faster than usual — against a ceiling it therefore cannot approach. Upside deviation never exceeds 2.48% anywhere in the window, which is what the 10% check is about. Read themax/mediancolumn, notfull range, when applying this check to amaxmetric; the reverse for aminone.
Bimodality: clean. mean_step_time_ms stayed within 1113.98–1148.66 ms across
all twenty cell-runs and the step-0 compile excursion did not recur once, so
warmup_steps: 2 did what it was changed to do.
If the window shows no excursion in twenty cell-runs, that is reasonable evidence
it has stopped happening, and median_ttft_ms and output_throughput can be
promoted with the same margins.
It did, and they were still not promoted in the first bless — deliberately. The window clears the excursion blocker on
median_ttft_msandoutput_throughput, so the evidence for promoting them is now in hand and promoting them is an ordinary step-7 PR.step_time_ms.maxis not in that list, and a clean window does not put it there: it is not a recorded metric, no per-night series exists for it, and a bound synthesised from one night’s mean at bless time is not something this window measured. Arming the first gate on the two metrics whose definitions exclude the failure mode we can demonstrate, and promotingmedian_ttft_msandoutput_throughputon their own evidence in their own diff, keeps a gate that fires attributable to the change that armed it.
This window is the first measurement at warmup_steps: 2, so it is also the
test of whether that change did what it was supposed to. Two outcomes worth
telling apart: no excursion in twenty cell-runs is the expected result and clears
median_ttft_ms and output_throughput for promotion in step 7; an excursion that
still appears at position 0 means two discarded steps are not enough to cover
the compile, which is new information and should be investigated rather than
absorbed by raising warmup_steps again.
5. Bless, scoped to this entry only. Done, by hand — see the note
immediately below for why a refresher run was never an option for this entry.
[!IMPORTANT] For
tokenspeed_serve_smoke, this is a hand-written baseline, not a refresher run.refresh-baselines.ymldoes not mount the docker socket — the sign-off is scoped to the nightly lane — so on that lane the entry declares a capability the lane does not have and skips. Dispatching a refresh withperf_gate_entry: tokenspeed_serve_smokeis therefore refused outright, with a message saying so, rather than producing a correctness-only file that reads in the PR diff exactly like a successful bless.This costs less than it sounds like, because step 6 below replaces every number the refresher derives with one computed from the ten-night window, and deletes most of the keys it writes. What the refresher would contribute is a skeleton. So write the two keys into
config/ci/regression_baselines.yamldirectly, in a PR, using the values from step 6 — same review, same file, fewer moving parts, and no GPU runner time.That is safe now in a way it was not before: a later refresh of any other entry carries these bounds over untouched, in both the scoped and the default correctness-only modes. Two options if the refresher’s convenience is wanted for this entry later: extend the sign-off to the refresh lane (a dispatch-only lane, so the fork-guard part of the argument is trivially satisfied — but it is a separate decision), or run the sweep on the runner host and hand-transcribe. Neither is on the critical path.
The mechanism itself, for the entries that can use it:
Actions -> Refresh baselines -> Run workflow
with the dispatch form filled in as
| Input | Value |
|---|---|
perf_gate |
true |
perf_gate_entry |
the entry under test, e.g. inference_offline |
step_time_margin |
0.25 (default) |
throughput_margin |
0.15 (default) |
which the workflow turns into
python scripts/ci/refresh_baselines.py --perf-gate \
--perf-gate-entry inference_offline
The scope is not optional. Without it the same PR arms step-time ceilings for
every other entry in the matrix from that one run, so leaving perf_gate_entry
empty with perf_gate: true logs a warning on the job. It is a warning rather
than a hard failure because an unscoped refresh is the legitimate end state once
every workload has variance data behind it — but during rollout it is not what
you want. Three ways of getting the scope wrong all fail rather than degrade: a
misspelled entry name is rejected rather than silently scoping to nothing; a
scope without perf_gate fails immediately instead of after the wheel install;
and a value that is non-empty but is only delimiters (,, a stray tab) is
rejected too, because it would otherwise parse to no names and leave a bare
--perf-gate — the unscoped global refresh, reached by an operator who thought
they had scoped it, and without even the unscoped warning.
What the scope does not do any more is disarm anything. Out-of-scope entries
keep their existing step_time_ms and metric bounds verbatim while their
correctness data is refreshed, which is what makes per-workload rollout additive
rather than a game of whack-a-mole. The same holds in the default,
correctness-only mode: omitting --perf-gate means “do not arm new gates”, not
“disarm the ones already blessed”.
perf_gate_entry accepts several names, comma- or space-separated, mapping to
one --perf-gate-entry each. For this rollout it should be exactly one.
6. Write the two bounds, then merge. Done — the four keys now in
config/ci/regression_baselines.yaml are max × 1.25 on the window maximum of
each cell’s median_tpot_ms and p99_itl_ms, and nothing else was written.
Whether the numbers come from a refresher diff or are written by hand (as they
are for this entry — see the note in step 5), the same two edits apply, because
the refresher derives bounds from the single run it just did and from every
auto-gateable metric it observed:
max × 1.25 for a
max metric, min × 0.85 for a min metric.p99_ttft_ms,
p99_tpot_ms, median_e2el_ms, p99_e2el_ms, request_throughput,
total_token_throughput, tokens_per_sec, and now also median_ttft_ms,
output_throughput and the step_time_ms.max entry — leaving
median_tpot_ms and p99_itl_ms. median_itl_ms will not be there;
_NO_AUTO_GATE keeps the refresher from writing it.step_time_ms.max is the one that needs attention, because it is the bound
--perf-gate always writes, the only one on this list that is not a metric key,
and — when hand-writing — the one most easily added out of a sense of
completeness. Including it arms the gate the measured excursion breaches
hardest: 2825 ms against a 1462 ms ceiling. It is the most likely mistake in
this whole sequence. The end state is exactly two metric keys per cell,
median_tpot_ms and p99_itl_ms, alongside passed: true.
This hand-editing is the honest cost of the current tooling, and it is bounded: ten keys across two cells, on a PR a human reviews anyway, and it is meant to shrink as the record-only metrics are promoted on evidence. It got larger rather than smaller when the excursion was measured, which strengthens the case below for a per-metric scope.
7. Promote the record-only metrics on evidence, not on schedule. The nine are blocked on three different things, and none of the groups is waiting on nights that have not happened yet: each starts from evidence the completed window already holds, and a second window is an outcome the middle group’s analysis might reach, not a precondition for doing it. The groups are the ones in the current-state box at the top of this file; this is what to do about each.
Evidence in hand — median_ttft_ms, output_throughput. Nothing to
wait for. Their only blocker was the step-0 excursion and the completed
2026-09-08..09-17 window cleared it in all twenty cell-runs
(step 4). They are unarmed because the first bless
was kept attributable, not because the measurement is missing. Promoting
them is an ordinary PR against the window already recorded here: step 6’s
derivations, on the table in step 4, with no new nights.
Mind the policy. median_ttft_ms is max and takes max × 1.25 on the
window maximum; output_throughput is min and takes min × 0.85 on the
window minimum. Both derivations are checked against this document by
test_each_blessed_bound_is_its_windows_extremum_times_the_policys_margin.
Analysis first, and possibly a second window — p99_ttft_ms,
p99_tpot_ms, p99_e2el_ms. Not a shortage of nights: the nightly
harvests every allowlisted metric, so the completed window recorded these too.
Step 3 only asked for five metrics to be written
down, so nobody has computed their spread — that is the gap, and it is an
afternoon with the dashboard rather than ten more nights. Read the ten
nights that exist, then decide: promote any whose spread is comparable to
its median’s, and take a second window only if a p99 over 32 requests turns
out too wide to size a margin from. That is the real risk here — a p99 over
32 requests is the 32nd of 32 order statistics — but it is a finding to
make, not an assumption to act on.
No window will unblock these — median_e2el_ms, request_throughput,
total_token_throughput, tokens_per_sec. Held back on redundancy, not
on evidence: each is determined by a metric already gated, so arming them
adds reason lines to one event rather than detection. What has to change is
the argument in the per-metric table,
not the number of nights.
Do the middle group rather than skipping it, and do not let the easy group
crowd it out: a gated p99_ttft_ms is worth more than a gated median_ttft_ms,
because tail latency is what a serving regression damages first. The first
group is the one with no work left in it, not the one with the most value in
it.
step_time_ms.max is not one of the nine and does not belong to any of these
groups, and a clean window does not clear it: it is not a recorded metric, no
per-night series accumulates for it, and a bound synthesised from one night’s
mean at bless time is not something a window measured. Promoting it needs that
to change, not more observations.
If the hand-editing recurs every refresh rather than converging, the fix is a
per-metric scope alongside --perf-gate-entry. That is the point to build it —
not before, when the right metric set is still a guess.
The alert names the cell, the metric, the observed value and the bound. Three things it can be, and they are distinguishable without a GPU:
A stack change. Check the dashboard’s What changed line first: it reports
whether PyTorch, ROCm or HIP moved that night. If one did, and the direction of
the metric move is plausible for it, this is a stack effect and the answer is
either an upstream bug report or a re-bless. Do not re-bless silently — a bump
that costs 20% of decode throughput is exactly the result the nightly exists to
produce, and docs/ci-nightly-eval.md is explicit that the baseline diff is
the bump’s impact.
A noise event. The two cells of this recipe differ only by mitigation and run
minutes apart on the same node, which is the reason this recipe was chosen. Use
it: if baseline and no-scratch-reclaim both moved by a similar amount, the
node or the stack changed. If exactly one moved, and the mitigation has never
shown an effect before at any model size, that is a single-cell excursion — a
co-tenant, a thermal event, a cold cache. Confirm by looking at the previous
nights on the dashboard’s run history; a noise event is one red cell in a column
of green, a regression is a column that turns red and stays red. Also check
server_startup_sec for that cell: it is not gated, but a bring-up at the top of
its range is a good indicator that the node was busy.
A real regression. Everything else — both cells moved, no toolchain change, and the next night reproduces it. Reproduce locally with the recipe as committed and bisect against the pinned image digest. Note that the image is pinned by digest precisely so this case cannot be caused by TokenSpeed changing underneath the baseline.
If you cannot tell which of the three within one working day, revert to
record-only and keep collecting. An unexplained gate is worth less than the
recording state it came from.
Delete the metrics and step_time_ms keys for the affected cells from
config/ci/regression_baselines.yaml, leaving passed: true. That returns those
cells to record-only immediately — the comparator treats an absent bound as
nothing to check — while keeping correctness gating and the fail-closed behaviour
intact. It is a small, reviewable, obviously-correct diff, which is what you want
at the point where a gate is misbehaving.
Do not revert by deleting the matrix entry. That stops the measurement as well as the gate, and the record-only data is the thing needed to size a better bound.
Do not revert by widening the margin as a first move. A margin wide enough to absorb an unexplained excursion is usually too wide to catch a regression, and the widening tends to be permanent. Go back to record-only, find out what moved, then re-bless from a longer window.
One whole-file caveat, now much narrower than it was: any later
refresh-baselines run rewrites regression_baselines.yaml completely, but a
run that is not deriving bounds for a cell carries that cell’s existing bounds
over untouched — so a hand-reverted cell stays reverted through a scoped refresh
of another entry and through a default correctness-only refresh. The one thing
that re-arms it is an unscoped --perf-gate refresh, which re-derives every
cell it can. Scope those runs with --perf-gate-entry, or check the diff for
cells you had deliberately demoted.
“Every cell it can” is the precise claim, and the exception matters for this
rollout specifically: a bound --perf-gate cannot produce is never deleted by
it either, because there would be nothing to put back. That covers a hand-written
median_itl_ms ceiling — the metric is allowlisted but _NO_AUTO_GATE, so the
refresher skips it — and any spec naming a metric off the allowlist entirely.
Both survive a re-bless of their own cell. So if the window later justifies
gating median_itl_ms by hand, a subsequent refresh of this same entry will not
quietly undo it.
The user referenced Aorta planning & Updates.docx, which is not reachable from
this machine. Everything above was derived from the repository, and where that
was not enough a decision was made and is recorded here.
tokenspeed-serve-bench-smoke.yaml was chosen over the other
four serving recipes as the cheapest that still answers something — Qwen3-0.6B
is ~1.2 GB of weights against ~40 GB for the gpt-oss pair, and two cells rather
than four or six. The tie-breaker was triage rather than cost: it is the only
serving recipe whose cells differ only by mitigation, so the pair is a
same-night control. tokenspeed-serve-gptoss.yaml is the recipe whose numbers
matter most (it is TokenSpeed’s canonical AMD benchmark) and is the right
second entry once this one has been gated for a month.--perf-gate-entry over hand-pruning. Scoping was added as code because
the alternative recurs on every refresh, for sixteen cells belonging to other
people’s workloads, in a diff where the omission looks identical to the
inclusion. The per-metric pruning in step 6 was left manual for the mirror-image
reason: it is ten keys in one entry, on a PR a human reviews anyway, and
building a flag for it now would fix a metric set we are explicitly planning to
change.needs_docker_daemon over leaving the entry staged. A launch was
demonstrated, so the entry had earned promotion; but the socket was still off
at the time, and an entry in entries that cannot reach a daemon fails every
night. Rather than choose between a stale pending_entries row and a red
nightly, the capability was made declarable, exactly as min_gpus already is
for GPU count. The socket has since been signed off for the nightly lane,
which does not make the field redundant — it is what keeps the grant scoped
to that lane, since the baseline refresher and bump-validate run the same
matrix without it and now skip the entry instead of failing on it. Both
consumers honour the flag; only the nightly did at first, which broke every
baseline refresh and was caught in review.warmup_steps raised to 2. Reversed from “left at 1” earlier in this
branch. The step-0 excursion is positional, so one more discarded step removes
it by construction, and it is measured at 1 cell-run in 13 — frequent enough
that a ten-night window taken at 1 would probably contain one and produce
either a false alarm or a threshold with no detection power. The objection to
bundling it was that it changes the measurement and invalidates every number
here; that is true and is now the accepted cost, because those numbers were
never going to be the bless baseline — the window is, and the window has not
been taken yet. Changing the setting before the window costs nothing;
changing it after would have cost ten nights. The variance table is kept as
the rationale.median_itl_ms. Kept in the allowlist and excluded from auto-blessing,
rather than removed. Removing it would make a legitimate hand-written bound
impossible; the problem is only ever with a bound derived from a margin.