A ROCm / PyTorch debugging, reproducibility, and workload-triage toolkit for AMD GPUs.
AORTA wraps opaque launch commands, runs recipe-driven mitigation sweeps, captures
versioned environment snapshots, evaluates GPU hardware-queue scheduling, and
reproduces workload-specific issues (numerics, races, nondeterminism) that
micro-benchmarks miss. It ships a single aorta CLI plus a plugin system so
downstream packages can register their own workloads.
Every night AORTA’s own workload sweep runs on an MI350 runner; the results are published to the nightly CI dashboard.
mitigation × {environment | diagnostic} × trial matrix from
one command (aorta sweep run). Recipe-driven, with a confound detector for
speed regressions and a five-tier failure classifier for subprocess runs.jq diff
instead of a multi-day investigation (aorta env probe). Embedded automatically
into every sweep result.Install AORTA in the Python environment from which you will invoke aorta.
The core install provides the CLI, recipe support, and environment probe. It
does not install or require PyTorch. AORTA requires Python 3.10 or newer.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install amd-aorta
aorta --help
The distribution name is amd-aorta; the import package and command are
aorta.
Use uv for a fast editable setup:
git clone https://github.com/ROCm/aorta.git
cd aorta
uv venv
source .venv/bin/activate
uv pip install -e .
aorta --help
Plain pip also works: python -m pip install -e .. Contributors who need
the test and lint tools can use uv pip install -e ".[dev]".
An extra adds dependencies for a specific feature; it is not part of the minimal install. Install only the extras you need. For example:
# Published package:
python -m pip install "amd-aorta[hw-queue]"
# Editable source checkout:
uv pip install -e ".[hw-queue]"
Hardware-queue evaluation and most GPU training or inference workloads also require PyTorch. AORTA does not bundle it. When the selected workload requires PyTorch, install a build matching that environment’s ROCm version:
# Set this to the index URL for your ROCm release from
# https://pytorch.org/get-started/locally/
PYTORCH_ROCM_INDEX=https://download.pytorch.org/whl/nightly/rocmX.Y/
python -m pip install --pre torch \
--index-url "$PYTORCH_ROCM_INDEX"
Other optional extras include analysis, report, hw-queue-profiling,
agent, and ebpf. The ebpf extra has no additional Python packages; its
runtime needs bpftrace and the required permissions.
aorta runs in the active Python environment, and relative recipe, command,
and output paths are resolved from the current directory. The examples below
use paths from the AORTA repository root.
There is no global rule that AORTA must run on the host or inside a container:
aorta env probe in the environment you want to inspect, including
inside a container when appropriate. To inspect an image without installing
AORTA in it, use the package-mount workflow.docker run launch. For those workloads, follow
the plugin’s instructions; the CLI normally runs on a Docker-capable host,
while the workload image supplies its runtime dependencies.The core dispatcher does not execute docker run; Docker-aware workload
plugins may do so and own the launch.
# --- Environment snapshot ---
aorta env probe -o env.json # full snapshot to disk
aorta env probe --summary # one-screen brief, no file write
aorta env probe --field pytorch_build.git_commit # one field, JSON-typed
diff <(jq -S . env_a.json) <(jq -S . env_b.json) # diff two snapshots
# --- Unified sweep (recipe-driven) ---
aorta sweep run --recipe recipes/llm-determinism/example-llm-determinism.yaml --dry-run # validate only
aorta sweep run --recipe recipes/llm-determinism/example-llm-determinism.yaml # run the matrix
# Distributed workloads launch under torchrun (llm_determinism, race, fsdp):
torchrun --standalone --nproc_per_node=2 $(which aorta) \
sweep run --recipe recipes/training/example-fsdp-smoke.yaml
torchrun --standalone --nproc_per_node=1 $(which aorta) \
run --workload llm_determinism --trials 1 --steps 50
# --- Unified sweep (wrap an opaque launch command) ---
aorta sweep run --recipe recipes/probe/probe-template-bash.yaml \
--ticket ROCM-1234 -- bash launch.sh
# --- Inspect the registries / pattern catalogue ---
aorta sweep list-mitigations
aorta sweep list-environments
aorta sweep list-patterns
aorta mitigations list # standalone registry view
aorta environments list
# --- Single workload trial (no matrix) — training/inference only ---
aorta run --workload training --trials 1 --steps 50
aorta run --workload inference --trials 1 --steps 50
# --- Hardware queue evaluation (requires amd-aorta[hw-queue] + torch) ---
aorta bench hw_queue_eval list
aorta bench hw_queue_eval run hetero_kernels --streams 8
aorta bench hw_queue_eval sweep hetero_kernels --streams 2,4,8,16
aorta sweep run is the single front door for matrix runs. It auto-selects a flow:
mode: triage (or --workload flag mode) runs
a registered in-process workload across mitigations × environments × trials.mode: probe, or any trailing
-- <command>, wraps an opaque launch command and classifies each trial with the
five-tier failure detector.Both flows write matrix.md + matrix.json, embed a per-environment env.json
snapshot, and emit a replayable recipe.resolved.yaml. See
docs/probe/usage.md.
At the end of every run aorta sweep run prints a concise summary to stdout —
which cells failed vs. errored, the workload’s own failure hint, and the path to
each failing cell’s artifact directory (logs + per-trial JSON) — so you don’t
have to open matrix.md to find what broke. Pass -v (-vv) to also stream
live per-cell progress to stderr while a long matrix runs.
aorta env probe captures a versioned, schema-stable env.json. Diff two
snapshots with jq to localize cross-environment regressions. See
docs/env-probe.md.
aorta bench hw_queue_eval stress-tests GPU queue scheduling across many
concurrent streams. See docs/hw-queue-eval.md.
aorta run --workload <name> runs one workload directly (trials/steps,
environment overlay, mitigations) without building a matrix — handy for iterating
on a single reproducer. Note: llm_determinism and race require a distributed
environment and must be launched under torchrun; training and inference
self-bootstrap a single process.
In-tree workloads, registered via the aorta.workloads entry-point group:
| Workload | Description |
|---|---|
training |
Real DDP / FSDP training loop. |
inference |
Offline / continuous-serving inference loop. |
llm_determinism |
Bit-exact double-run check of a transformer step (FSDP2-aware, RCCL-safe, optional MoE). See docs/llm-determinism.md. |
race |
RCCL race / SDC reproducer (mode: default \| ddp \| fsdp). Distributed — launch under torchrun. |
_subprocess is a platform-internal workload that backs the subprocess flow; it is
not meant for direct aorta run use.
Downstream / private workloads register through the same aorta.workloads
entry-point group from their own pyproject.toml, so they appear in aorta sweep
without modifying this repo.
A recipe is the authoritative description of a sweep: which cells to run, per-cell trial/step counts, the ticket, and confound-detection config. Recipes are the primary interface; flag mode is an escape hatch.
recipes/README.md — schema reference and field semantics.recipes/README-running-recipes.md — how to
run them (including distributed launches under torchrun).recipes/ (e.g.
example-llm-determinism.yaml, example-fsdp-smoke.yaml,
probe-template-bash.yaml).Minimal workload recipe:
schema_version: 1
ticket: EXAMPLE-001
workload: training
trials: 2
steps: 100
cells:
- name: baseline-local
mitigations: [none]
environment: local
- name: tf32_off-local
mitigations: [tf32_off]
environment: local
Run it:
aorta sweep run --recipe my-recipe.yaml
aorta probe and aorta triage have merged into the unified aorta sweep
front door. The old commands still work as deprecated aliases — they delegate
to the same execution engine and print a one-line stderr notice — but new usage
should target aorta sweep.
| Deprecated command | Use instead |
|---|---|
aorta probe ... -- cmd |
aorta sweep run ... -- cmd |
aorta triage run ... |
aorta sweep run ... |
aorta triage list-mitigations |
aorta sweep list-mitigations |
aorta triage list-environments |
aorta sweep list-environments |
aorta probe --list-patterns |
aorta sweep list-patterns |
The standalone aorta mitigations list and aorta environments list groups are
not deprecated and remain available.
Note: probe-only runtime knobs (
--stop-after-events,--max-trials,--disable-detector) are not yet exposed onaorta sweep run. Until they land, keep usingaorta probefor those specific flags.
| Guide | Description |
|---|---|
| Getting Started | Installation, command location, and workload-specific prerequisites |
| Hardware Queue Eval | Workloads, CLI usage, metrics |
| Environment Probe | Capture / diff / query a versioned environment snapshot; jq cookbook |
aorta sweep |
Unified matrix runner — built-in workloads or opaque launch commands |
| LLM Determinism | Bit-exact double-run nondeterminism probe |
| Layer Numerics | Per-layer / per-stage NaN, magnitude, and out-of-range logger |
aorta agent |
Closed-loop mitigation search (optional LLM proposer) |
aorta bundle |
Package sweep artifacts with recipe-driven redaction |
| Recipes | Recipe schema and running recipes |
| Buck2 Build Reference | Build / run the AORTA CLI via Buck2 |
src/aorta/
├── cli/ # `aorta` CLI command groups (sweep, run, env, bundle, agent, ...)
├── workloads/ # In-tree workloads (training, inference, llm_determinism, race)
├── instrumentation/ # Environment probe (env.json) + layer_numerics NaN/OOB logger
├── registry/ # Mitigations + environments registry (extension points)
├── hw_queue_eval/ # Hardware queue evaluation framework
├── training/ # FSDP2 trainer with multi-stream overlap instrumentation
├── models/ # Synthetic ranking transformer
├── profiling/ # Stream profiler for overlap measurement
└── utils/ # Config loading, timing, device detection
recipes/ # Sweep recipes (examples + customer handout templates)
docs/ # Guides and reference
scripts/ # Launch, profiling, analysis tooling
uv pip install -e ".[dev]"
pre-commit install
pytest tests/
The FSDP2 overlap and hardware-queue workloads also run on NVIDIA CUDA for side-by-side comparison with ROCm.