Before you start
- One gfx950 (MI355X) or gfx1250 GPU — the TokenSpeed image targets these
- A working docker client and daemon: the engine runs in its own container
- A node-local work_dir — an NFS home under root-squash cannot be bind-mounted
- Egress to the Hugging Face Hub, or a pre-populated cache plus hf_offline
1. Check out the dashboard commit and install the matching AORTA wheel
git clone https://github.com/ROCm/aorta.git
cd aorta
git checkout e45a9ae778dfbb0796da23ad516befc1939f7266
python3 -m pip install --upgrade pip
python3 -m pip install --upgrade --pre 'amd-aorta[hw-queue]==0.3.1rc20260929' -f https://github.com/ROCm/aorta/releases/expanded_assets/dev-wheels
# If that exact wheel is unavailable, pick the closest dev-wheel build
# and compare `aorta --version` with the dashboard header before reproducing.
aorta --help
2. Confirm ROCm + PyTorch see enough GPUs
python3 -c "import torch; n=torch.cuda.device_count(); assert torch.cuda.is_available() and n>0, 'no CUDA/HIP devices'; assert n>=1, f'need 1 GPU(s), have {n}'; print(f'{n} GPU(s), HIP', torch.version.hip)"3. Serving-specific: pull the pinned engine image
docker pull lightseekorg/tokenspeed-amd@sha256:60c12e37c01496891053b9c30c4204e5d1cf9b4b641859d3aadcbd95bccc7c78
4. Serving-specific: pre-warm the model cache as the running uid
# The path is load-bearing. The workload derives its cache from
# workload_config.hf_home, falling back to <work_dir>/u<uid>/hf --
# it does not read an inherited HF_HOME, so a download into any
# other directory populates a cache the run never mounts and the
# first cell downloads the weights again regardless.
# run_as_current_user defaults to true, so download as the uid
# that will run the sweep: a cache populated by a root container
# leaves the trial failing with PermissionError.
mkdir -p /tmp/ts-work-serve/u$(id -u)/hf
HF_HOME=/tmp/ts-work-serve/u$(id -u)/hf hf download Qwen/Qwen3-0.6B
5. Validate the recipe YAML (no GPU execution)
aorta sweep run --recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml --dry-run
6. Run the recipe with --strict (same flag as nightly CI)
Run from the repository root after setup completes. A standalone --strict sweep only fails cells that error or never run. Dashboard pass/record/fail comes separately from nightly_eval.py comparing matrix.json against config/ci/regression_baselines.yaml; run that harness to apply the dashboard gates.
aorta sweep run --recipe recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml --output-dir triage_results/repro/tokenspeed_serve_smoke --strict
7. Inspect serving latency and throughput artifacts
RUN_DIR="$(find triage_results/repro/tokenspeed_serve_smoke -mindepth 3 -maxdepth 3 -type d -printf '%T@ %p\n' 2>/dev/null | sort -rn | head -1 | cut -d' ' -f2-)"
if [ -z "$RUN_DIR" ] || [ ! -d "$RUN_DIR" ]; then echo "No run directory found"; exit 1; fi
echo "Using run directory: $RUN_DIR"
cat "$RUN_DIR/matrix.md"
python3 -m json.tool "$RUN_DIR/matrix.json" | head -80
if [ -f "$RUN_DIR/perf.md" ]; then head -60 "$RUN_DIR/perf.md"; else echo "(no perf.md)"; fi
# Standalone --strict only catches errored or not-run cells; dashboard pass/record/fail comes from nightly_eval.py separately comparing matrix.json to config/ci/regression_baselines.yaml
grep -m 10 -E "ttft|tpot|output_throughput|completed_total" "$RUN_DIR/perf.md" 2>/dev/null || python3 -c "import json,sys; d=json.load(open(sys.argv[1])); print(json.dumps(d.get(\"cells\",[]), indent=2)[:2000])" "$RUN_DIR/matrix.json"
Success criteria
Both cells pass with completed_total == num_prompts * steps and failed_total == 0. median_ttft_ms / median_tpot_ms / output_throughput appear in perf.md. Two metrics are gated on both cells — median_tpot_ms and p99_itl_ms, ceilings from the 2026-09-08..09-17 window — so a run over either fails the nightly, and the breaching observation is still recorded and charted in this cell's history for diagnosis. Every other serving metric is still record-only, including step_time_ms.max; see docs/tokenspeed-gating-rollout.md for which metrics get bounds and why bring-up time never does (189-379s on one node with nothing changed).
Compare with nightly numbers
Match the aorta, PyTorch, ROCm, and HIP versions shown in the dashboard header when comparing numbers. Small differences on different hardware are expected; regressions vs the blessed baseline are what nightly CI flags.
Recipe YAML: recipes/tokenspeed/tokenspeed-serve-bench-smoke.yaml · Multi-node launcher guide