Overview

This page introduces HIP EP, explains how it executes an ONNX graph on an AMD GPU, and points you to the next steps.

HIP EP is an ONNX Runtime Execution Provider (EP) for AMD GPUs. It runs as part of ONNX Runtime rather than as a standalone inference server or API. Once registered, ONNX Runtime delegates supported parts of the graph to HIP EP for GPU execution.

This site documents v0.5.1. Release-specific links on this page point to the v0.5.1 tag rather than main, so the documentation remains tied to that release.

If your application already uses ONNX Runtime, no code changes are required beyond registering HIP EP. The same application can then run the supported graph operations on the GPU.

How a graph gets executed

HIP EP is an LLM inference backend rather than a standalone operator library. When ONNX Runtime assigns a subgraph to HIP EP, the subgraph goes through an MLIR-based compilation pipeline:

  1. ONNX → HIP dialect. ONNX operations are converted to a custom MLIR hip dialect that explicitly represents GPU memory, kernels, and library calls.
  2. HIP dialect → LLVM IR. Shape inference, memory planning and buffer pooling are performed before the dialect is lowered to LLVM IR.
  3. Execution. The compiled result schedules the workload, using hipBLASLt for matrix operations and custom HIP kernels for other operations.

The compiler output drives GPU execution but does not contain the GPU kernels themselves. The kernels are built ahead of time and shipped in the package as custom_kernels_*. As a result, a single package can support multiple architectures without requiring a GPU compiler on the target machine.

By default, the per-model artifact is OS-portable LLVM bitcode. HIP EP JIT-loads this bitcode in-process together with the embedded runtime bitcode. Native .dll/.so model artifacts are also available as an opt-in mode.

The first inference is slower because the model is compiled at runtime. This compilation happens once per process. Benchmarks should therefore include a warm-up phase; otherwise, the results include compilation overhead.

Supported hardware

The v0.5.1 Windows package covers the following RDNA 3.5 GPUs. The performance data on this site was collected on Strix Halo (gfx1151); the other entries indicate package coverage and should not be interpreted as equivalent benchmark results.

GPU Architecture
Ryzen AI Max (“Strix Halo”) gfx1151
Ryzen AI (“Strix Point”) gfx1150
Ryzen AI (“Krackan Point”) gfx1152

The Windows release provides a single package containing code objects for all three parts, with no architecture-specific variants. For other AMD GPU families, build HIP EP from source and specify the target with --hip_arch.

Version pinning

HIP EP uses specific versions of its upstream dependencies. Using a different ONNX Runtime version is not supported because the EP is loaded as a plugin against a specific ABI.

Component Version
HIP EP v0.5.1
ONNX Runtime 1.30.0
ONNX Runtime GenAI (OGA) 0.14.0 + AMDGPU integration PR 2194

The full dependency set — including LLVM/MLIR/LLD, protobuf, flatbuffers, ONNX Runtime, TheRock ROCm — is pinned in cmake/deps.txt.

OGA packaging is defined in windows-deps.yml, including the OGA_VERSION and OGA_PR_PATCHES settings. These files are the source of truth if the table above differs from the repository.

Official sources

Source What to use it for
v0.5.1 release Download the release and Python packages documented on this site
cmake/deps.txt Check pinned dependency versions
windows-build-real.yml See which binaries, libraries, and wheels are staged into the Windows packages
ONNX Runtime GenAI Learn about the tokenizer, KV-cache and decode loop used for LLM inference
Supported operations See which ONNX operations are supported by the compile pipeline

Where to go next

The model matrix lists the models included in the v0.5.1 snapshot, while the benchmarks page explains how to interpret the results. For implementation details, see the design documentation which covers pass ordering, the compiler/runtime ABI and memory planning.