Model Matrix

The models included in the HIP EP v0.5.1 benchmark snapshot.

This page lists the models with published quantized exports and benchmark results in this release. It is not a complete list of models supported by HIP EP. More than 50 models are validated across a release cycle, while this matrix covers only the models included in this benchmark snapshot.

A model not listed here may still run successfully. Use the same ONNX Runtime or OGA path, verify GPU execution, and benchmark it on your target hardware.

Model matrix

There are 16 models across 7 families: 8 text models and 8 vision-language models. The table can be sorted or filtered by model, family, architecture, or quantization.

Model Family Type Parameters Architecture Quantization Prefill @2Ktok/s Decodetok/s
GPT-OSS 20B GPT-OSS Text (LLM) 20B total, 3.6B active Sparse MoE int4 RTN 1640 77.62
GPT-OSS 120B GPT-OSS Text (LLM) 120B total, 5.1B active Sparse MoE uint4 AWQ 373 44.53
Qwen2.5 14B Instruct Qwen Text (LLM) 14B Dense int4 RTN 873 26.84
Qwen2.5-Coder 14B Instruct Qwen Text (LLM) 14B Dense int4 RTN 869 26.77
Llama 3.1 8B Llama Text (LLM) 8B Dense int4 AWQ 1338 46.75
Mistral 7B Instruct v0.3 Mistral Text (LLM) 7B Dense int4 AWQ 1389 51.56
Phi-4 14B Phi Text (LLM) 14B Dense int4 RTN 796 25.70
DeepSeek-R1-Distill-Llama 70B DeepSeek Text (LLM) 70B Dense int4 AWQ 167 6.09
Qwen3.5 9B Qwen Vision-language (VLM) 9B Dense int4 888 36.94
Qwen3.5 35B-A3B Qwen Vision-language (VLM) 35B total, 3B active Sparse MoE int4 1057 58.40
Qwen3.6 27B Qwen Vision-language (VLM) 27B Dense int4 RTN 342 13.00
Qwen3.6 35B-A3B Qwen Vision-language (VLM) 35B total, 3B active Gated DeltaNet + MoE int4 1001 58.66
Qwen3.8 27B Qwen Vision-language (VLM) 27B Dense int4 k-quant 389 13.78
Gemma 3 4B IT Gemma Vision-language (VLM) 4B Dense int4 RTN 1619 65.96
Gemma 4 12B IT Gemma Vision-language (VLM) 12B Dense int4 k-quant 928 23.99
Gemma 4 26B-A4B IT Gemma Vision-language (VLM) 26B total, 4B active Sparse MoE int4 QMoE 1268 54.46

Prefill and decode rates use fixed input and output lengths across the matrix: 2K input tokens for prefill and 128 generated tokens for decode. Prefill is derived from prompt length and time-to-first-token.

Prefill results are not directly comparable between text and vision-language models because VLM prefill also includes image processing through the vision encoder. Filter by model type before comparing results. Results for all three prompt lengths, together with the underlying TTFT measurements, are available on the Benchmarks page.

By family

Family Organization Models
Qwen Alibaba 7
Gemma Google 3
GPT-OSS OpenAI 2
Llama Meta 1
DeepSeek DeepSeek 1
Mistral Mistral AI 1
Phi Microsoft 1

See the family pages for model-specific architecture, quantization, and runtime details.

The same models are also listed on the landing page as filterable cards.

What the snapshot records

Checked What it means What to check locally
Function The model compiles and produces output in the release run Missing operators, unresolved shapes, or crashes on your exact export
Performance Time-to-first-token and tokens-per-second for the published prompt lengths Driver version, power profile, and thermal conditions can affect absolute numbers
Accuracy Perplexity and task scores do not degrade against the reference implementation Validate with your own prompts, tasks, or reference outputs

Performance and accuracy are separate measurements. A throughput result describes the performance of the published quantized export on the reference system; it does not establish whether the model is suitable for a particular application.

Dense and sparse

The matrix distinguishes between dense models and sparse mixture-of-experts (MoE) models.

A dense model uses all of its parameters for each token. A sparse MoE model activates only a subset of its parameters for each token. For models with an A3B or A4B suffix, the suffix indicates the approximate number of active parameters per token.

For sparse models, total parameters are primarily relevant to memory requirements, while active parameters are more relevant to compute and throughput. Neither value should be used on its own to compare model performance.

Qwen3.6 also combines sparse routing with Gated DeltaNet rather than standard attention. It uses the same HIP EP compilation pipeline.

What “int4” means here

All models in the matrix use quantized weights. The weights are represented as 4-bit integers grouped along the input dimension, with a scale for each group and, for asymmetric schemes, a zero point. Activations remain in fp16.

HIP EP consumes the standard ONNX MatMulNBits representation used by the broader ONNX ecosystem; the format is not specific to HIP EP.

For vision-language models, the vision encoder remains in fp16 while the text decoder is quantized. The encoder runs once per image, while the decoder runs for each generated token.

The matrix includes several quantization schemes, including RTN, AWQ, k-quant, symmetric and asymmetric variants, with group sizes from 32 to 128. HIP EP can consume these different formats.

Quantization is part of the model artifact rather than a runtime option. HIP EP does not quantize models at runtime. An fp16 model therefore runs in fp16 and retains the corresponding memory footprint.

Running a model

These models use the same ONNX Runtime execution path as other ONNX models. See Get Started for installation and setup.

For generative models, ONNX Runtime GenAI (OGA) provides the surrounding inference flow, including tokenization, KV-cache management, and the decode loop. HIP EP executes the ONNX graph used by that flow.

Calling session.run() directly on a decoder graph produces the logits for a single token; it does not provide the tokenizer, KV-cache, or decode loop needed for autoregressive generation.

Where the numbers are

Family pages show a representative prompt length for each model. The Benchmarks page contains the full benchmark snapshot, including all three prompt lengths and the test conditions.

The benchmark data represents a single release snapshot and is intended for comparison within the conditions documented on that page.