Benchmarks

Benchmark results for the HIP EP v0.5.1 release.

Releasev0.5.1
OSWindows
HardwareRyzen AI Max ("Strix Halo", gfx1151)

These results were collected on a single machine. Results from a different GPU, driver, power configuration, or thermal state are not directly comparable.

What this page covers

This page reports a single benchmark snapshot for v0.5.1. It is not a release history or a comparison with other runtimes.

The results apply to v0.5.1 on the hardware listed above. If you are running a different release or configuration, treat these numbers as reference data rather than measurements of your setup.

Metrics

Metric Unit Direction Description
TTFT seconds Lower is better Time from submitting the prompt to the first generated token.
TPS tokens/second Higher is better Steady-state generation rate after the first token.

TTFT is primarily affected by prefill and therefore increases with prompt length. TPS measures decode throughput after the first token.

Compile time is excluded from both metrics. HIP EP compiles the graph on first use; subsequent calls use the compiled graph. In v0.5.1, compiled artifacts are kept for the lifetime of the session and are not cached on disk, so a new process recompiles the graph.

See the Overview for details on the compilation flow.

Results

The benchmarks use prompt lengths of 128, 2048, 16384 tokens.

Horizontal bar chart. Every model in the table below, ordered by tokens generated per second at a 2048-token prompt on Ryzen AI Max ("Strix Halo", gfx1151). The table carries the same figures at all 3 prompt lengths. 0 20 40 60 80GPT-OSS 20B 77.67Gemma 3 4B IT 62.98Qwen3.6 35B-A3B 57.77Qwen3.5 35B-A3B 57.67Mistral 7B Instruct v0.3 48.02Gemma 4 26B-A4B IT 47.21GPT-OSS 120B 43.37Llama 3.1 8B 43.14Qwen3.5 9B 36.21Qwen2.5-Coder 14B Instruct 24.49Qwen2.5 14B Instruct 24.49Phi-4 14B 24.01Gemma 4 12B IT 21.85Qwen3.8 27B 13.60Qwen3.6 27B 12.90DeepSeek-R1-Distill-Llama 70B 5.93
Tokens generated per second at a 2048-token prompt, on Ryzen AI Max ("Strix Halo", gfx1151), in v0.5.1. Higher is better. The order is the measured rate at this one prompt length and nothing else — these models differ in size, task and modality, so it is not a ranking of them against each other, and a model that leads here need not lead at 16384 tokens.

TPS at a 2048-token prompt on Ryzen AI Max (“Strix Halo”, gfx1151), using HIP EP v0.5.1. Higher is better. The chart shows one prompt length; the table below contains results for all three.

Model Parameters TTFT 128 (s) TTFT 2048 (s) TTFT 16384 (s) TPS 128 TPS 2048 TPS 16384
GPT-OSS 20B 20B total, 3.6B active 0.17 1.21 11.27 77.62 77.67 68.20
GPT-OSS 120B 120B total, 5.1B active 0.74 5.32 44.20 44.53 43.37 39.02
Qwen2.5-Coder 14B Instruct 14B 0.16 2.35 27.51 26.77 24.49 17.56
Qwen2.5 14B Instruct 14B 0.16 2.34 27.31 26.84 24.49 17.60
Llama 3.1 8B 8B 0.12 1.50 16.48 46.75 43.14 30.59
Mistral 7B Instruct v0.3 7B 0.11 1.45 16.30 51.56 48.02 33.10
Phi-4 14B 14B 0.17 2.52 26.52 25.70 24.01 18.05
DeepSeek-R1-Distill-Llama 70B 70B 1.05 12.03 116.60 6.09 5.93 5.20
Qwen3.6 35B-A3B 35B total, 3B active 2.59 3.95 15.37 58.66 57.77 52.13
Qwen3.5 9B 9B 2.71 4.45 20.82 36.94 36.21 33.38
Qwen3.5 35B-A3B 35B total, 3B active 2.66 3.74 15.73 58.40 57.67 51.96
Qwen3.6 27B 27B 6.20 11.55 54.49 13.00 12.90 11.78
Qwen3.8 27B 27B 5.61 10.28 53.44 13.78 13.60 12.34
Gemma 3 4B IT 4B 0.79 1.41 7.12 65.96 62.98 57.42
Gemma 4 12B IT 12B 0.52 2.46 21.54 23.99 21.85 19.40
Gemma 4 26B-A4B IT 26B total, 4B active 0.84 1.80 11.45 54.46 47.21 18.86

TTFT generally increases with prompt length because prefill processes the prompt tokens. TPS generally decreases as the prompt grows because each generated token attends over a larger KV cache.

For sparse MoE models, the parameter count is shown as total / active parameters. Decode performance is more closely related to the active parameter count than the total parameter count.

For vision-language models, TTFT includes the vision encoder. This contributes to the higher TTFT often seen at short prompt lengths.

Scope

This page covers generative model benchmarks using TTFT and TPS. Other workload classes are validated separately and are not included here.

The Procyon AI Inference Benchmark is also part of the test suite. For v0.5.1, its workloads are functionality-verified, with performance optimization still in progress. Results are therefore not published here.

Measurement conditions

The benchmark harness uses the following conditions:

  • Serial execution. Only one benchmark runs on the GPU at a time.
  • Debug and tracing disabled. HIPDNN_EP_PERF=1 and HIPDNN_EP_DEBUG=1 are not set.
  • Autotune caches primed. Cold autotune runs are excluded.
  • First inference excluded. The first inference includes one-time graph compilation.

Thermal state is not controlled. Sustained workloads can reduce clock or memory performance as the system heats up, particularly on thin systems.

Different drivers, power profiles, memory configurations, and thermal conditions can also affect absolute results. Reproduced results should therefore be compared only under similar conditions.

Reproducing the results

The release package includes model_benchmark for Windows. It runs models through OGA and reports TTFT and TPS directly. See the binary package page for usage and benchmark options.

Before comparing results, verify that the model is actually running on the GPU. ONNX Runtime can fall back to CPU execution. Set:

$env:HIPDNN_EP_STRICT=1

to make unsupported GPU execution fail instead of silently falling back to CPU.