aorta

Profiling Guide

This guide covers how to capture, analyze, and interpret profiling data from AORTA benchmark runs.

Profiling Outputs

Each rank writes artifacts/rank_<rank>_metrics.jsonl containing iteration-level telemetry:

Events are captured using torch.cuda.Event(enable_timing=True) for microsecond fidelity. Distributed collectives are monkey-patched at runtime so that all-reduce and reduce-scatter operations execute on dedicated streams and contribute to overlap calculation.

Torch Profiler Traces

Enable PyTorch’s profiler by toggling the profiling block in your config or via CLI override:

torchrun --nproc_per_node 4 train.py \
  --config config/default.yaml \
  --override profiling.enabled=true \
  --override profiling.wait=1 \
  --override profiling.warmup=1 \
  --override profiling.active=2

Output Locations

Chrome Traces

ROCm rocprofv3 Capture

Use the wrapper script to profile an entire ROCm run:

bash scripts/rocprof_capture.sh config/default.yaml --override training.max_steps=50

Output Location

Outputs land under rocprof_traces/run_<timestamp>/.

Environment Variables

Override location or extra flags with environment variables:

The script mirrors launch_rocm.sh but executes through rocprofv3, so you can merge traces with the JSONL metrics using the shared iteration timestamps.

Generating Reports

Run the analyser to build summaries and plots from one or more log directories:

python analysis/overlap_report.py \
  --log-dir artifacts_rocm --label rocm \
  --log-dir artifacts_cuda --label cuda \
  --output reports/2024-roc-vs-cuda \
  --reference cuda --candidate rocm

Report Outputs

Use these artefacts to pinpoint scheduling or synchronisation regressions between hardware backends.

Diagnostic Insights

Overlap Breakdown

Overlap Ratio

Key Metrics

Metric Interpretation
Overlap Ratio (overlap_ratio) Values close to 1 indicate strong overlap; values near 0 imply communications block compute
Compute All-Reduce (compute_allreduce_ms) Time spent in all-reduce operations
Compute Reduce-Scatter (compute_reducescatter_ms) Time spent in reduce-scatter operations

Analysis Tips

Advanced Profiling

For deeper inspection, combine these scripts with nsys, rocprof, or PyTorch profiler traces using the iteration timestamps documented in the JSON traces.

Next Steps