aorta nightly CI

recording

Every night AORTA installs the freshly built wheel on an MI350 runner and replays its workload sweep, comparing each result against a blessed baseline. This page is that run.

aorta
0.2.2rc20260810
PyTorch
2.9.1+rocm7.2.0.git7e1940d4
ROCm
7.2.0
HIP
7.2.26015-fc0010cf6a

Run of 2026-08-10 11:35 UTC · commit b79ffc3 · workflow run · amd_aorta-0.2.2rc20260810-py3-none-any.whl

No baselines are blessed yet, so nothing was graded pass or fail — this run recorded metrics for all 16 results to become the reference. Every result reports: no baseline (record-only).
results
16
pass
0
fail
0
record
16
skip
0
pass-rate trend
n/a

What changed since 2026-08-09

Run history

Each row is one nightly run, newest first; each column is a workload. A cell shows the worst verdict among that workload's results (✓ pass · ✗ fail · ◆ record · ○ skip · ? unrecognised) and how many there were — ✗ 1/4 means one of four failed. Hover for the full breakdown.

runstatusgpu_smokeinference_offlinetraining_ddptraining_fsdpracellm_determinismtraining_ddp_8gputraining_fsdp_8gpurace_8gpullm_determinism_8gpu
2026-08-100.2.2rc20260810recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-090.2.2rc20260809recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-080.2.2rc20260808recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-070.2.2rc20260807recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-060.2.2rc20260806recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-050.2.2rc20260805recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-040.2.2rc20260804recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4
2026-08-030.2.2rc20260803recording◆ 1◆ 1◆ 1◆ 1◆ 1◆ 4◆ 1◆ 1◆ 1◆ 4

Workloads

cellstatusstep timetrend (step ms)
gpu_smoke 1 result · 1 record · workload run 16s
baseline-localrecord263.2 ms
3 metrics

recipe recipes/ci/gpu-smoke.yaml · 1 trial

metricpolicylatesttrend
expectedtrend only8
ntrend only8
sumtrend only8
inference_offline 1 result · 1 record · workload run 24s
baseline-localrecord37.7 ms
11 metrics

recipe recipes/inference/example-inference-smoke.yaml · 1 trial

metricpolicylatesttrend
batch_sizetrend only4
decode_latency_msmax0.9615 ms
decoded_tokenstrend only496
generate_tokenstrend only32
logits_checksumequal33,904,201
parameter_counttrend only33,170,432
prefill_latency_msmax1.079 ms
prompt_lentrend only128
step_time_p50trend only35.4 ms
step_time_p99trend only45.08 ms
tokens_per_secmin4160 tok/s
llm_determinism 4 results · 4 record · workload run 111s
baseline-bf16-24Lrecord12,491.3 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only24
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only2
bf16-12Lrecord12,451.2 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only12
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only2
moe4-bf16-12Lrecord13,583.4 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only4
num_layerstrend only12
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only2
tf32_off-bf16-24Lrecord12,602.4 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only24
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only2
llm_determinism_8gpu 4 results · 4 record · workload run 169s
baseline-bf16-24Lrecord20,921.2 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only24
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only8
bf16-12Lrecord19,818.1 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only12
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only8
moe4-bf16-12Lrecord21,241.0 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only4
num_layerstrend only12
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only8
tf32_off-bf16-24Lrecord20,149.5 ms
6 metrics

recipe recipes/llm-determinism/example-llm-determinism.yaml · 1 trial

metricpolicylatesttrend
corruption_details_omittedtrend only0
num_expertstrend only1
num_layerstrend only24
ranktrend only0
ranks_with_divergencetrend only0
world_sizetrend only8
race 1 result · 1 record · workload run 32s
smokerecord2,410.6 ms
14 metrics

recipe recipes/race/race_smoke.yaml · 1 trial

metricpolicylatesttrend
avg_step_time_mstrend only2411
corruption_details_omittedtrend only0
declared_h2d_tensor_sizetrend only1,000,000
eff_batch_sizetrend only1
eff_ffn_sizetrend only2,048
eff_num_headstrend only4
eff_seq_lentrend only512
effective_h2d_tensor_sizetrend only1,000,000
layer_checksum_mismatchestrend only0
layers_verifiedtrend only15
local_world_sizetrend only2
node_counttrend only1
ranktrend only0
world_sizetrend only2
race_8gpu 1 result · 1 record · workload run 47s
smokerecord4,426.0 ms
14 metrics

recipe recipes/race/race_smoke.yaml · 1 trial

metricpolicylatesttrend
avg_step_time_mstrend only4426
corruption_details_omittedtrend only0
declared_h2d_tensor_sizetrend only1,000,000
eff_batch_sizetrend only1
eff_ffn_sizetrend only2,048
eff_num_headstrend only4
eff_seq_lentrend only512
effective_h2d_tensor_sizetrend only1,000,000
layer_checksum_mismatchestrend only0
layers_verifiedtrend only15
local_world_sizetrend only8
node_counttrend only1
ranktrend only0
world_sizetrend only8
training_ddp 1 result · 1 record · workload run 37s
baseline-localrecord1,662.0 ms
6 metrics

recipe recipes/training/example-training-ddp-smoke.yaml · 1 trial

metricpolicylatesttrend
final_losstrend only6.937
parameter_counttrend only656,896
ranktrend only0
step_time_p50trend only1.973 ms
step_time_p99trend only3.392 ms
world_sizetrend only2
training_ddp_8gpu 1 result · 1 record · workload run 47s
baseline-localrecord1,711.1 ms
6 metrics

recipe recipes/training/example-training-ddp-smoke.yaml · 1 trial

metricpolicylatesttrend
final_losstrend only6.937
parameter_counttrend only656,896
ranktrend only0
step_time_p50trend only1.426 ms
step_time_p99trend only43.12 ms
world_sizetrend only8
training_fsdp 1 result · 1 record · workload run 32s
baseline-localrecord1,569.5 ms
6 metrics

recipe recipes/training/example-training-fsdp-smoke.yaml · 1 trial

metricpolicylatesttrend
final_losstrend only254.4
parameter_counttrend only2,361,856
ranktrend only0
step_time_p50trend only7.759 ms
step_time_p99trend only48.42 ms
world_sizetrend only2
training_fsdp_8gpu 1 result · 1 record · workload run 47s
baseline-localrecord2,171.4 ms
6 metrics

recipe recipes/training/example-training-fsdp-smoke.yaml · 1 trial

metricpolicylatesttrend
final_losstrend only254.4
parameter_counttrend only2,361,856
ranktrend only0
step_time_p50trend only6.887 ms
step_time_p99trend only46.93 ms
world_sizetrend only8

Verdicts: pass/fail = compared against a blessed baseline · record = no baseline yet, metrics captured as the future reference · skip = not enough GPUs on the runner. One row is one recorded result: a workload skipped before it reached its cells yields a single row for the whole workload, so a count of rows is not a count of configured cells. Expand a row for its metrics and recipe.