ONNX Runtime Execution Provider

The fastest, most efficient LLM inference backend on AMD iGPU

HIP EP compiles your ONNX graph through an MLIR pipeline and executes it on AMD GPUs with hipBLASLt and custom HIP kernels.

Model showcase

Models that already run on your GPU Measured on release v0.5.1

Qwen

LLM · Dense

Qwen2.5 14B Instruct

Prefill speed, at a prompt of Prefill 2K
873 tok/s
Decode throughput, Decode
26.84 tok/s

LLM · Dense

Qwen2.5-Coder 14B Instruct

Prefill speed, at a prompt of Prefill 2K
869 tok/s
Decode throughput, Decode
26.77 tok/s

VLM · Dense

Qwen3.5 9B

Prefill speed, at a prompt of Prefill 16K
901 tok/s
Decode throughput, Decode
36.94 tok/s

VLM · MoE

Qwen3.5 35B-A3B

Prefill speed, at a prompt of Prefill 16K
1192 tok/s
Decode throughput, Decode
58.40 tok/s

VLM · Dense

Qwen3.6 27B

Prefill speed, at a prompt of Prefill 16K
344 tok/s
Decode throughput, Decode
13.00 tok/s

VLM · MoE

Qwen3.6 35B-A3B

Prefill speed, at a prompt of Prefill 16K
1220 tok/s
Decode throughput, Decode
58.66 tok/s

VLM · Dense

Qwen3.8 27B

Prefill speed, at a prompt of Prefill 2K
389 tok/s
Decode throughput, Decode
13.78 tok/s

Gemma

VLM · Dense

Gemma 3 4B IT

Prefill speed, at a prompt of Prefill 16K
2300 tok/s
Decode throughput, Decode
65.96 tok/s

VLM · Dense

Gemma 4 12B IT

Prefill speed, at a prompt of Prefill 2K
928 tok/s
Decode throughput, Decode
23.99 tok/s

VLM · MoE

Gemma 4 26B-A4B IT

Prefill speed, at a prompt of Prefill 16K
1430 tok/s
Decode throughput, Decode
54.46 tok/s

GPT-OSS

LLM · MoE

GPT-OSS 20B

Prefill speed, at a prompt of Prefill 2K
1640 tok/s
Decode throughput, Decode
77.62 tok/s

LLM · MoE

GPT-OSS 120B

Prefill speed, at a prompt of Prefill 2K
373 tok/s
Decode throughput, Decode
44.53 tok/s

Llama

LLM · Dense

Llama 3.1 8B

Prefill speed, at a prompt of Prefill 2K
1338 tok/s
Decode throughput, Decode
46.75 tok/s

DeepSeek

Mistral

LLM · Dense

Mistral 7B Instruct v0.3

Prefill speed, at a prompt of Prefill 2K
1389 tok/s
Decode throughput, Decode
51.56 tok/s

Phi

LLM · Dense

Phi-4 14B

Prefill speed, at a prompt of Prefill 2K
796 tok/s
Decode throughput, Decode
25.70 tok/s