Model Matrix
The models included in the HIP EP v0.5.1 benchmark snapshot.
This page lists the models with published quantized exports and benchmark results in this release. It is not a complete list of models supported by HIP EP. More than 50 models are validated across a release cycle, while this matrix covers only the models included in this benchmark snapshot.
A model not listed here may still run successfully. Use the same ONNX Runtime or OGA path, verify GPU execution, and benchmark it on your target hardware.
Model matrix
There are 16 models across 7 families: 8 text models and 8 vision-language models. The table can be sorted or filtered by model, family, architecture, or quantization.
| Model | Family | Type | Parameters | Architecture | Quantization | Prefill @2Ktok/s | Decodetok/s |
|---|---|---|---|---|---|---|---|
| GPT-OSS 20B | GPT-OSS | Text (LLM) | 20B total, 3.6B active | Sparse MoE | int4 RTN | 1640 | 77.62 |
| GPT-OSS 120B | GPT-OSS | Text (LLM) | 120B total, 5.1B active | Sparse MoE | uint4 AWQ | 373 | 44.53 |
| Qwen2.5 14B Instruct | Qwen | Text (LLM) | 14B | Dense | int4 RTN | 873 | 26.84 |
| Qwen2.5-Coder 14B Instruct | Qwen | Text (LLM) | 14B | Dense | int4 RTN | 869 | 26.77 |
| Llama 3.1 8B | Llama | Text (LLM) | 8B | Dense | int4 AWQ | 1338 | 46.75 |
| Mistral 7B Instruct v0.3 | Mistral | Text (LLM) | 7B | Dense | int4 AWQ | 1389 | 51.56 |
| Phi-4 14B | Phi | Text (LLM) | 14B | Dense | int4 RTN | 796 | 25.70 |
| DeepSeek-R1-Distill-Llama 70B | DeepSeek | Text (LLM) | 70B | Dense | int4 AWQ | 167 | 6.09 |
| Qwen3.5 9B | Qwen | Vision-language (VLM) | 9B | Dense | int4 | 888 | 36.94 |
| Qwen3.5 35B-A3B | Qwen | Vision-language (VLM) | 35B total, 3B active | Sparse MoE | int4 | 1057 | 58.40 |
| Qwen3.6 27B | Qwen | Vision-language (VLM) | 27B | Dense | int4 RTN | 342 | 13.00 |
| Qwen3.6 35B-A3B | Qwen | Vision-language (VLM) | 35B total, 3B active | Gated DeltaNet + MoE | int4 | 1001 | 58.66 |
| Qwen3.8 27B | Qwen | Vision-language (VLM) | 27B | Dense | int4 k-quant | 389 | 13.78 |
| Gemma 3 4B IT | Gemma | Vision-language (VLM) | 4B | Dense | int4 RTN | 1619 | 65.96 |
| Gemma 4 12B IT | Gemma | Vision-language (VLM) | 12B | Dense | int4 k-quant | 928 | 23.99 |
| Gemma 4 26B-A4B IT | Gemma | Vision-language (VLM) | 26B total, 4B active | Sparse MoE | int4 QMoE | 1268 | 54.46 |
No model matches that filter.
Prefill and decode rates use fixed input and output lengths across the matrix: 2K input tokens for prefill and 128 generated tokens for decode. Prefill is derived from prompt length and time-to-first-token.
Prefill results are not directly comparable between text and vision-language models because VLM prefill also includes image processing through the vision encoder. Filter by model type before comparing results. Results for all three prompt lengths, together with the underlying TTFT measurements, are available on the Benchmarks page.
By family
| Family | Organization | Models |
|---|---|---|
| Qwen | Alibaba | 7 |
| Gemma | 3 | |
| GPT-OSS | OpenAI | 2 |
| Llama | Meta | 1 |
| DeepSeek | DeepSeek | 1 |
| Mistral | Mistral AI | 1 |
| Phi | Microsoft | 1 |
See the family pages for model-specific architecture, quantization, and runtime details.
The same models are also listed on the landing page as filterable cards.
What the snapshot records
| Checked | What it means | What to check locally |
|---|---|---|
| Function | The model compiles and produces output in the release run | Missing operators, unresolved shapes, or crashes on your exact export |
| Performance | Time-to-first-token and tokens-per-second for the published prompt lengths | Driver version, power profile, and thermal conditions can affect absolute numbers |
| Accuracy | Perplexity and task scores do not degrade against the reference implementation | Validate with your own prompts, tasks, or reference outputs |
Performance and accuracy are separate measurements. A throughput result describes the performance of the published quantized export on the reference system; it does not establish whether the model is suitable for a particular application.
Dense and sparse
The matrix distinguishes between dense models and sparse mixture-of-experts (MoE) models.
A dense model uses all of its parameters for each token. A sparse MoE model
activates only a subset of its parameters for each token. For models with an
A3B or A4B suffix, the suffix indicates the approximate number of active
parameters per token.
For sparse models, total parameters are primarily relevant to memory requirements, while active parameters are more relevant to compute and throughput. Neither value should be used on its own to compare model performance.
Qwen3.6 also combines sparse routing with Gated DeltaNet rather than standard attention. It uses the same HIP EP compilation pipeline.
What “int4” means here
All models in the matrix use quantized weights. The weights are represented as 4-bit integers grouped along the input dimension, with a scale for each group and, for asymmetric schemes, a zero point. Activations remain in fp16.
HIP EP consumes the standard ONNX MatMulNBits representation used by the
broader ONNX ecosystem; the format is not specific to HIP EP.
For vision-language models, the vision encoder remains in fp16 while the text decoder is quantized. The encoder runs once per image, while the decoder runs for each generated token.
The matrix includes several quantization schemes, including RTN, AWQ, k-quant, symmetric and asymmetric variants, with group sizes from 32 to 128. HIP EP can consume these different formats.
Quantization is part of the model artifact rather than a runtime option. HIP EP does not quantize models at runtime. An fp16 model therefore runs in fp16 and retains the corresponding memory footprint.
Running a model
These models use the same ONNX Runtime execution path as other ONNX models. See Get Started for installation and setup.
For generative models, ONNX Runtime GenAI (OGA) provides the surrounding inference flow, including tokenization, KV-cache management, and the decode loop. HIP EP executes the ONNX graph used by that flow.
Calling session.run() directly on a decoder graph produces the logits for a
single token; it does not provide the tokenizer, KV-cache, or decode loop needed
for autoregressive generation.
Where the numbers are
Family pages show a representative prompt length for each model. The Benchmarks page contains the full benchmark snapshot, including all three prompt lengths and the test conditions.
The benchmark data represents a single release snapshot and is intended for comparison within the conditions documented on that page.