Benchmarks
Benchmark results for the HIP EP v0.5.1 release.
| Release | v0.5.1 |
|---|---|
| OS | Windows |
| Hardware | Ryzen AI Max ("Strix Halo", gfx1151) |
These results were collected on a single machine. Results from a different GPU, driver, power configuration, or thermal state are not directly comparable.
What this page covers
This page reports a single benchmark snapshot for v0.5.1. It is not a release history or a comparison with other runtimes.
The results apply to v0.5.1 on the hardware listed above. If you are running a different release or configuration, treat these numbers as reference data rather than measurements of your setup.
Metrics
| Metric | Unit | Direction | Description |
|---|---|---|---|
| TTFT | seconds | Lower is better | Time from submitting the prompt to the first generated token. |
| TPS | tokens/second | Higher is better | Steady-state generation rate after the first token. |
TTFT is primarily affected by prefill and therefore increases with prompt length. TPS measures decode throughput after the first token.
Compile time is excluded from both metrics. HIP EP compiles the graph on first use; subsequent calls use the compiled graph. In v0.5.1, compiled artifacts are kept for the lifetime of the session and are not cached on disk, so a new process recompiles the graph.
See the Overview for details on the compilation flow.
Results
The benchmarks use prompt lengths of 128, 2048, 16384 tokens.
TPS at a 2048-token prompt on Ryzen AI Max (“Strix Halo”, gfx1151), using HIP EP v0.5.1. Higher is better. The chart shows one prompt length; the table below contains results for all three.
| Model | Parameters | TTFT 128 (s) | TTFT 2048 (s) | TTFT 16384 (s) | TPS 128 | TPS 2048 | TPS 16384 |
|---|---|---|---|---|---|---|---|
| GPT-OSS 20B | 20B total, 3.6B active | 0.17 | 1.21 | 11.27 | 77.62 | 77.67 | 68.20 |
| GPT-OSS 120B | 120B total, 5.1B active | 0.74 | 5.32 | 44.20 | 44.53 | 43.37 | 39.02 |
| Qwen2.5-Coder 14B Instruct | 14B | 0.16 | 2.35 | 27.51 | 26.77 | 24.49 | 17.56 |
| Qwen2.5 14B Instruct | 14B | 0.16 | 2.34 | 27.31 | 26.84 | 24.49 | 17.60 |
| Llama 3.1 8B | 8B | 0.12 | 1.50 | 16.48 | 46.75 | 43.14 | 30.59 |
| Mistral 7B Instruct v0.3 | 7B | 0.11 | 1.45 | 16.30 | 51.56 | 48.02 | 33.10 |
| Phi-4 14B | 14B | 0.17 | 2.52 | 26.52 | 25.70 | 24.01 | 18.05 |
| DeepSeek-R1-Distill-Llama 70B | 70B | 1.05 | 12.03 | 116.60 | 6.09 | 5.93 | 5.20 |
| Qwen3.6 35B-A3B | 35B total, 3B active | 2.59 | 3.95 | 15.37 | 58.66 | 57.77 | 52.13 |
| Qwen3.5 9B | 9B | 2.71 | 4.45 | 20.82 | 36.94 | 36.21 | 33.38 |
| Qwen3.5 35B-A3B | 35B total, 3B active | 2.66 | 3.74 | 15.73 | 58.40 | 57.67 | 51.96 |
| Qwen3.6 27B | 27B | 6.20 | 11.55 | 54.49 | 13.00 | 12.90 | 11.78 |
| Qwen3.8 27B | 27B | 5.61 | 10.28 | 53.44 | 13.78 | 13.60 | 12.34 |
| Gemma 3 4B IT | 4B | 0.79 | 1.41 | 7.12 | 65.96 | 62.98 | 57.42 |
| Gemma 4 12B IT | 12B | 0.52 | 2.46 | 21.54 | 23.99 | 21.85 | 19.40 |
| Gemma 4 26B-A4B IT | 26B total, 4B active | 0.84 | 1.80 | 11.45 | 54.46 | 47.21 | 18.86 |
TTFT generally increases with prompt length because prefill processes the prompt tokens. TPS generally decreases as the prompt grows because each generated token attends over a larger KV cache.
For sparse MoE models, the parameter count is shown as total / active parameters. Decode performance is more closely related to the active parameter count than the total parameter count.
For vision-language models, TTFT includes the vision encoder. This contributes to the higher TTFT often seen at short prompt lengths.
Scope
This page covers generative model benchmarks using TTFT and TPS. Other workload classes are validated separately and are not included here.
The Procyon AI Inference Benchmark is also part of the test suite. For v0.5.1, its workloads are functionality-verified, with performance optimization still in progress. Results are therefore not published here.
Measurement conditions
The benchmark harness uses the following conditions:
- Serial execution. Only one benchmark runs on the GPU at a time.
- Debug and tracing disabled.
HIPDNN_EP_PERF=1andHIPDNN_EP_DEBUG=1are not set. - Autotune caches primed. Cold autotune runs are excluded.
- First inference excluded. The first inference includes one-time graph compilation.
Thermal state is not controlled. Sustained workloads can reduce clock or memory performance as the system heats up, particularly on thin systems.
Different drivers, power profiles, memory configurations, and thermal conditions can also affect absolute results. Reproduced results should therefore be compared only under similar conditions.
Reproducing the results
The release package includes model_benchmark for Windows. It runs models
through OGA and reports TTFT and TPS directly. See the
binary package page
for usage and benchmark options.
Before comparing results, verify that the model is actually running on the GPU. ONNX Runtime can fall back to CPU execution. Set:
$env:HIPDNN_EP_STRICT=1
to make unsupported GPU execution fail instead of silently falling back to CPU.