GPT-OSS

Two sparse mixture-of-experts text models, including the largest model in the matrix.

Both GPT-OSS entries are sparse mixtures of experts, which is why the matrix has a 120B model in it at all. A sparse model holds far more parameters than it reads: only a few billion of them take part in any one token, so the rate it generates at tracks that active figure rather than its total size. The total still has to fit in memory — which is what makes these a workload for unified memory rather than a workload for a small discrete card.

Sparse routing is also the case that breaks naive per-operator dispatch, since which experts run is decided at runtime. Both models compile through the standard pipeline regardless; see the overview for what the release checks.

GPT-OSS 20B

The fastest model in the matrix, and a sparse MoE — the architecture that breaks naive per-operator dispatch.

  • Kind: Language model
  • Task: Text generation, reasoning
  • Parameters: 20B total, 3.6B active
  • Architecture: Sparse mixture of experts
  • Quantization: int4 RTN, block 32
  • Base model: openai/gpt-oss-20b
  • Validation status: Validated
  • Measured in v0.5.1: 77.67 tokens/s, 1.21 s to first token, at a 2048-token prompt — all three lengths

GPT-OSS 120B

120B parameters on an integrated GPU, at interactive speed. This is what 128 GB of unified memory is for.

  • Kind: Language model
  • Task: Text generation, reasoning
  • Parameters: 120B total, 5.1B active
  • Architecture: Sparse mixture of experts
  • Quantization: uint4 AWQ, per-group asymmetric
  • Base model: openai/gpt-oss-120b
  • Validation status: Validated
  • Measured in v0.5.1: 43.37 tokens/s, 5.32 s to first token, at a 2048-token prompt — all three lengths