Run models with the binary package
Download the HIP EP v0.5.1 Windows test package, select AMDGPU in an existing model directory, and run its text or vision-language benchmark.
The Windows GPU test package is a portable HIP EP installation. It contains the EP and OGA/ONNX Runtime binaries, the user-mode ROCm runtime, and the MSVC/UCRT import libraries needed by the JIT linker. Extracting the archive is the only installation step; it does not register system components.
This guide covers release v0.5.1 on Strix Halo, Strix Point, Krackan Point,
and GPT2. It assumes that Windows and the graphics driver are already installed.
The v0.5.1 procedure uses GPT2 as its fourth target label. Because that is not
a retail adapter name, confirm the matching product with the release owner
before using that target.
Before starting, obtain an OGA-ready ONNX model directory. It must contain
genai_config.json and the model and tokenizer files referenced by that
configuration. A base-model repository by itself is not enough. The
Model Matrix identifies validated model
names and configurations, but it does not distribute the ready-to-run directories.
1. Prerequisites
| Item | Requirement |
|---|---|
| OS | Windows (amd64) |
| Shell | Windows PowerShell. The commands below use PowerShell syntax; HIP EP itself does not require a particular shell |
| GPU | Strix Halo, Strix Point, Krackan Point, or GPT2, with a current graphics driver. The archive supplies the user-mode HIP runtime DLLs; the installed driver supplies the kernel-mode GPU support |
| Python | 3.14 (cp314) for the VLM benchmark only. The wheels in wheels/ use this ABI; model_benchmark.exe does not use Python |
| Model | An existing OGA-ready ONNX directory containing genai_config.json and the files it references |
| Network | Required to download the package and model. The VLM dependency install also uses pip’s configured package index unless those packages are already cached |
2. Download and extract the test package
Download
hipep_0.5.1_windows.zip.
The archive contains the portable bin/, lib/, and wheels/ directories.
Models are distributed separately.
tar on Windows 10 and later extracts .zip files:
mkdir gpu-test-package
tar -xf hipep_0.5.1_windows.zip -C gpu-test-package
Use an existing model directory and edit only its genai_config.json before
running. Text models use model_benchmark.exe. Vision-language models use
vlm_benchmark.py and also require a local image.
<working-directory>\
hipep_0.5.1_windows.zip
gpu-test-package\
bin\
lib\
wheels\
<model-directory>\
genai_config.json
<image>
3. Set the EP in genai_config.json
OGA reads execution-provider settings from each model directory. Before the
first run, open <model-directory>\genai_config.json. Under model → decoder
→ session_options, replace the existing provider_options block with:
"provider_options": [
{
"AMDGPU": {
"profile": "hip"
}
}
]
Leave the rest of the file unchanged. Both model_benchmark.exe and
vlm_benchmark.py read this block. This model-level setting is an OGA contract:
the benchmark command does not independently override the model’s provider.
4. Run
The commands below assume section 3 is complete. Replace <model-directory>
with the directory already present in your working folder.
LLM benchmark
Run model_benchmark.exe by its full path so it remains next to the DLLs in
bin.
.\gpu-test-package\bin\model_benchmark.exe -i <model-directory> -l 128 -g 128 -r 5 -w 1 -b 1 -v
| Flag | Purpose |
|---|---|
-i <dir> |
OGA-ready model directory |
-l <n> |
Prompt length |
-g <n> |
Tokens to generate |
-r <n> |
Benchmark repetitions |
-w <n> |
Warm-up runs |
-b <n> |
Batch size |
-v |
Verbose output |
Example output:
Running warmup iterations (1)...
[PROMPT BEGIN]
[OUTPUT BEGIN]
Running iterations (5)...
Batch size: 1, prompt tokens: 128, tokens to generate: 128
VLM Benchmark (multimodal models)
bin/vlm_benchmark.py reports image-preprocessing time, time to first token,
and token-generation rate. The commands below create a Python 3.14 environment
and run the benchmark. Replace <model-directory> and <image> with local
paths.
py -3.14 --version
py -3.14 -m venv .venv
.venv\Scripts\activate
pip install (Get-Item gpu-test-package\wheels\*.whl) # bash: pip install gpu-test-package/wheels/*.whl
pip install psutil pillow numpy
cd gpu-test-package\bin
python vlm_benchmark.py `
--model_path ..\..\<model-directory> `
--image_path ..\..\<image> `
--max_tokens 128 --max_length 4096 `
--num_iterations 5 --warmup_iterations 1 `
--output_json ..\..\vlm-results.json --verbose
| Flag | Purpose |
|---|---|
--model_path <dir> |
OGA-ready model directory |
--image_path <path> |
Input image |
--max_tokens <n> |
Maximum tokens to generate |
--max_length <n> |
Maximum sequence length |
--num_iterations <n> |
Benchmark runs |
--warmup_iterations <n> |
Warm-up runs |
--output_json <path> |
Results file |
--verbose |
Verbose output |
In the output below, Device: AMDGPU is the benchmark’s direct confirmation
that the AMDGPU provider is active:
Loading model...
Device: AMDGPU
Prompt: <|im_start|>user
Running 1 warmup iteration(s)...
Running 5 benchmark iteration(s)...
============================================================
VLM BENCHMARK RESULTS
============================================================