Overview
This page introduces HIP EP, explains how it executes an ONNX graph on an AMD GPU, and points you to the next steps.
HIP EP is an ONNX Runtime Execution Provider (EP) for AMD GPUs. It runs as part of ONNX Runtime rather than as a standalone inference server or API. Once registered, ONNX Runtime delegates supported parts of the graph to HIP EP for GPU execution.
This site documents v0.5.1. Release-specific links on
this page point to the v0.5.1 tag rather than main, so
the documentation remains tied to that release.
If your application already uses ONNX Runtime, no code changes are required beyond registering HIP EP. The same application can then run the supported graph operations on the GPU.
How a graph gets executed
HIP EP is an LLM inference backend rather than a standalone operator library. When ONNX Runtime assigns a subgraph to HIP EP, the subgraph goes through an MLIR-based compilation pipeline:
- ONNX → HIP dialect. ONNX operations are converted to a custom MLIR
hipdialect that explicitly represents GPU memory, kernels, and library calls. - HIP dialect → LLVM IR. Shape inference, memory planning and buffer pooling are performed before the dialect is lowered to LLVM IR.
- Execution. The compiled result schedules the workload, using hipBLASLt for matrix operations and custom HIP kernels for other operations.
The compiler output drives GPU execution but does not contain the GPU kernels
themselves. The kernels are built ahead of time and shipped in the package as
custom_kernels_*. As a result, a single package can support multiple
architectures without requiring a GPU compiler on the target machine.
By default, the per-model artifact is OS-portable LLVM bitcode. HIP EP
JIT-loads this bitcode in-process together with the embedded runtime bitcode.
Native .dll/.so model artifacts are also available as an opt-in mode.
The first inference is slower because the model is compiled at runtime. This compilation happens once per process. Benchmarks should therefore include a warm-up phase; otherwise, the results include compilation overhead.
Supported hardware
The v0.5.1 Windows package covers the following RDNA 3.5
GPUs. The performance data on this site was collected on Strix Halo
(gfx1151); the other entries indicate package coverage and should not be
interpreted as equivalent benchmark results.
| GPU | Architecture |
|---|---|
| Ryzen AI Max (“Strix Halo”) | gfx1151 |
| Ryzen AI (“Strix Point”) | gfx1150 |
| Ryzen AI (“Krackan Point”) | gfx1152 |
The Windows release provides a single package containing code objects for all
three parts, with no architecture-specific variants. For other AMD GPU
families, build HIP EP from source and specify the target with --hip_arch.
Version pinning
HIP EP uses specific versions of its upstream dependencies. Using a different ONNX Runtime version is not supported because the EP is loaded as a plugin against a specific ABI.
| Component | Version |
|---|---|
| HIP EP | v0.5.1 |
| ONNX Runtime | 1.30.0 |
| ONNX Runtime GenAI (OGA) | 0.14.0 + AMDGPU integration PR 2194 |
The full dependency set — including LLVM/MLIR/LLD, protobuf, flatbuffers, ONNX Runtime,
TheRock ROCm — is pinned in
cmake/deps.txt.
OGA packaging is defined in
windows-deps.yml,
including the OGA_VERSION and OGA_PR_PATCHES settings.
These files are the source of truth if the table above differs from the repository.
Official sources
| Source | What to use it for |
|---|---|
| v0.5.1 release | Download the release and Python packages documented on this site |
cmake/deps.txt |
Check pinned dependency versions |
windows-build-real.yml |
See which binaries, libraries, and wheels are staged into the Windows packages |
| ONNX Runtime GenAI | Learn about the tokenizer, KV-cache and decode loop used for LLM inference |
| Supported operations | See which ONNX operations are supported by the compile pipeline |
Where to go next
Run a model
Extract the release archive, add
bin to PATH, and run an inference on the GPU. No
compiler required.
Use the Python package
Use the Python package to serve an LLM, verify that it is running on the GPU, run benchmarks, and try your own model.
The model matrix lists the models included in the v0.5.1 snapshot, while the benchmarks page explains how to interpret the results. For implementation details, see the design documentation which covers pass ordering, the compiler/runtime ABI and memory planning.