LLM · Dense
Qwen2.5 14B Instruct
- Prefill speed, at a prompt of Prefill 2K
- 873 tok/s
- Decode throughput, Decode
- 26.84 tok/s
ONNX Runtime Execution Provider
HIP EP compiles your ONNX graph through an MLIR pipeline and executes it on AMD GPUs with hipBLASLt and custom HIP kernels.
Model showcase
LLM · Dense
LLM · Dense
VLM · Dense
VLM · MoE
VLM · Dense
VLM · MoE
VLM · Dense
VLM · Dense
VLM · Dense
VLM · MoE
LLM · MoE
LLM · MoE
LLM · Dense
LLM · Dense
LLM · Dense
LLM · Dense
Get Started
From an empty machine to a model on the GPU, one page per install path, each ending in a check that catches the silent CPU fallback.
Overview
What HIP EP is, how a graph reaches the GPU, which hardware is covered, and what it is pinned to.
Models
The language and vision-language models validated on every release, and what each one is built out of.
Benchmarks
One snapshot per release — throughput and latency, with the machine, the commands and the warm-up rules behind them.
Something here wrong, missing, or contradicted by your own machine? The compiler, the runtime, the kernels and this site are all in one open-source repository — open an issue, including the case where the documentation is what is broken.