aorta nightly CI
recording
Every night AORTA installs the freshly built wheel on an
MI350 runner and replays its workload sweep, comparing each result
against a blessed baseline. This page is that run.
AORTA documentation ·
sanitizer nightly ·
repository
PyTorch
2.9.1+rocm7.2.0.git7e1940d4
Run of 2026-08-10 11:35 UTC · commit b79ffc3 · workflow run · amd_aorta-0.2.2rc20260810-py3-none-any.whl
No baselines are blessed yet, so nothing was graded pass or fail — this run recorded metrics for all 16 results to become the reference. Every result reports: no baseline (record-only).
What changed since 2026-08-09 training_ddp::baseline-local step_time_p50 ▲ +20.9% 1.632 → 1.973 ms training_ddp::baseline-local step_time_p99 ▲ +16.1% 2.923 → 3.392 ms training_fsdp::baseline-local step_time_p99 ▼ -14.4% 56.58 → 48.42 ms training_fsdp_8gpu::baseline-local step_time_p99 ▼ -10.1% 52.17 → 46.93 ms
Run history Each row is one nightly run, newest first; each column is a workload. A cell shows the worst verdict among that workload's results (✓ pass · ✗ fail · ◆ record · ○ skip · ? unrecognised) and how many there were — ✗ 1/4 means one of four failed. Hover for the full breakdown.
run status gpu_smoke inference_offline training_ddp training_fsdp race llm_determinism training_ddp_8gpu training_fsdp_8gpu race_8gpu llm_determinism_8gpu 2026-08-10 0.2.2rc20260810 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-09 0.2.2rc20260809 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-08 0.2.2rc20260808 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-07 0.2.2rc20260807 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-06 0.2.2rc20260806 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-05 0.2.2rc20260805 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-04 0.2.2rc20260804 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4 2026-08-03 0.2.2rc20260803 recording ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 1 ◆ 4 ◆ 1 ◆ 1 ◆ 1 ◆ 4
Workloads
Expand all
Collapse all
cell status step time trend (step ms)
gpu_smoke 1 result · 1 record · workload run 16s baseline-local record 263.2 ms 3 metrics recipe recipes/ci/gpu-smoke.yaml · 1 trial
metric policy latest trend expected trend only 8 n trend only 8 sum trend only 8
inference_offline 1 result · 1 record · workload run 24s baseline-local record 37.7 ms 11 metrics recipe recipes/inference/example-inference-smoke.yaml · 1 trial
metric policy latest trend batch_size trend only 4 decode_latency_ms max 0.9615 ms decoded_tokens trend only 496 generate_tokens trend only 32 logits_checksum equal 33,904,201 parameter_count trend only 33,170,432 prefill_latency_ms max 1.079 ms prompt_len trend only 128 step_time_p50 trend only 35.4 ms step_time_p99 trend only 45.08 ms tokens_per_sec min 4160 tok/s
llm_determinism 4 results · 4 record · workload run 111s baseline-bf16-24L record 12,491.3 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 24 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 2
bf16-12L record 12,451.2 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 12 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 2
moe4-bf16-12L record 13,583.4 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 4 num_layers trend only 12 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 2
tf32_off-bf16-24L record 12,602.4 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 24 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 2
llm_determinism_8gpu 4 results · 4 record · workload run 169s baseline-bf16-24L record 20,921.2 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 24 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 8
bf16-12L record 19,818.1 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 12 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 8
moe4-bf16-12L record 21,241.0 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 4 num_layers trend only 12 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 8
tf32_off-bf16-24L record 20,149.5 ms 6 metrics recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial
metric policy latest trend corruption_details_omitted trend only 0 num_experts trend only 1 num_layers trend only 24 rank trend only 0 ranks_with_divergence trend only 0 world_size trend only 8
race 1 result · 1 record · workload run 32s smoke record 2,410.6 ms 14 metrics recipe recipes/race/race_smoke.yaml · 1 trial
metric policy latest trend avg_step_time_ms trend only 2411 corruption_details_omitted trend only 0 declared_h2d_tensor_size trend only 1,000,000 eff_batch_size trend only 1 eff_ffn_size trend only 2,048 eff_num_heads trend only 4 eff_seq_len trend only 512 effective_h2d_tensor_size trend only 1,000,000 layer_checksum_mismatches trend only 0 layers_verified trend only 15 local_world_size trend only 2 node_count trend only 1 rank trend only 0 world_size trend only 2
race_8gpu 1 result · 1 record · workload run 47s smoke record 4,426.0 ms 14 metrics recipe recipes/race/race_smoke.yaml · 1 trial
metric policy latest trend avg_step_time_ms trend only 4426 corruption_details_omitted trend only 0 declared_h2d_tensor_size trend only 1,000,000 eff_batch_size trend only 1 eff_ffn_size trend only 2,048 eff_num_heads trend only 4 eff_seq_len trend only 512 effective_h2d_tensor_size trend only 1,000,000 layer_checksum_mismatches trend only 0 layers_verified trend only 15 local_world_size trend only 8 node_count trend only 1 rank trend only 0 world_size trend only 8
training_ddp 1 result · 1 record · workload run 37s baseline-local record 1,662.0 ms 6 metrics recipe recipes/training/example-training-ddp-smoke.yaml · 1 trial
metric policy latest trend final_loss trend only 6.937 parameter_count trend only 656,896 rank trend only 0 step_time_p50 trend only 1.973 ms step_time_p99 trend only 3.392 ms world_size trend only 2
training_ddp_8gpu 1 result · 1 record · workload run 47s baseline-local record 1,711.1 ms 6 metrics recipe recipes/training/example-training-ddp-smoke.yaml · 1 trial
metric policy latest trend final_loss trend only 6.937 parameter_count trend only 656,896 rank trend only 0 step_time_p50 trend only 1.426 ms step_time_p99 trend only 43.12 ms world_size trend only 8
training_fsdp 1 result · 1 record · workload run 32s baseline-local record 1,569.5 ms 6 metrics recipe recipes/training/example-training-fsdp-smoke.yaml · 1 trial
metric policy latest trend final_loss trend only 254.4 parameter_count trend only 2,361,856 rank trend only 0 step_time_p50 trend only 7.759 ms step_time_p99 trend only 48.42 ms world_size trend only 2
training_fsdp_8gpu 1 result · 1 record · workload run 47s baseline-local record 2,171.4 ms 6 metrics recipe recipes/training/example-training-fsdp-smoke.yaml · 1 trial
metric policy latest trend final_loss trend only 254.4 parameter_count trend only 2,361,856 rank trend only 0 step_time_p50 trend only 6.887 ms step_time_p99 trend only 46.93 ms world_size trend only 8
Verdicts: pass /fail = compared
against a blessed baseline · record = no baseline yet, metrics
captured as the future reference · skip = not enough GPUs on the
runner. One row is one recorded result: a workload skipped before it reached
its cells yields a single row for the whole workload, so a count of rows is
not a count of configured cells. Expand a row for its metrics and recipe.