Qwen

Two Qwen2.5 text models and five Qwen3.x vision-language models in the release matrix.

Qwen is the widest family in the matrix and the one that covers the most architectural ground. Qwen2.5 is a pair of conventional dense text decoders, one of them specialised for code. Qwen3.x is vision-language, and between its five entries it spans dense decoders, sparse mixtures of experts, and — in Qwen3.6 35B-A3B — Gated DeltaNet in place of plain attention.

That range is why this page is worth reading rather than skimming: none of those four shapes needed an operator library of its own. They compile through the same pipeline as everything else on the overview, and the benchmark snapshot keeps their differences in separate rows rather than turning them into separate install paths.

Qwen2.5 14B Instruct

  • Kind: Language model
  • Task: Text generation
  • Parameters: 14B
  • Architecture: Dense transformer
  • Quantization: int4 RTN, group 128
  • Base model: Qwen/Qwen2.5-14B-Instruct
  • Validation status: Validated
  • Measured in v0.5.1: 24.49 tokens/s, 2.34 s to first token, at a 2048-token prompt — all three lengths

Qwen2.5-Coder 14B Instruct

A coding assistant that fits in local memory and holds a 16K-token file in its context without falling over.

  • Kind: Language model
  • Task: Code generation
  • Parameters: 14B
  • Architecture: Dense transformer
  • Quantization: int4 RTN, group 128
  • Base model: Qwen/Qwen2.5-Coder-14B-Instruct
  • Validation status: Validated
  • Measured in v0.5.1: 24.49 tokens/s, 2.35 s to first token, at a 2048-token prompt — all three lengths

Qwen3.5 9B

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 9B
  • Architecture: Dense transformer + vision encoder
  • Quantization: int4 text decoder, group 32; fp16 vision encoder
  • Base model: Qwen/Qwen3.5-9B
  • Validation status: Validated
  • Measured in v0.5.1: 36.21 tokens/s, 4.45 s to first token, at a 2048-token prompt — all three lengths

Qwen3.5 35B-A3B

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 35B total, 3B active
  • Architecture: Sparse mixture of experts + vision encoder
  • Quantization: int4 text decoder, group 32; fp16 vision encoder
  • Base model: Qwen/Qwen3.5-35B-A3B
  • Validation status: Validated
  • Measured in v0.5.1: 57.67 tokens/s, 3.74 s to first token, at a 2048-token prompt — all three lengths

Qwen3.6 27B

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 27B
  • Architecture: Dense transformer + vision encoder
  • Quantization: int4 RTN text decoder, group 32; fp16 vision encoder
  • Base model: Qwen/Qwen3.6-27B
  • Validation status: Validated
  • Measured in v0.5.1: 12.90 tokens/s, 11.55 s to first token, at a 2048-token prompt — all three lengths

Qwen3.6 35B-A3B

Gated DeltaNet and a sparse MoE in the same graph. Nothing about this model is a stock transformer, and it still compiles to one artifact.

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 35B total, 3B active
  • Architecture: Gated DeltaNet + sparse mixture of experts
  • Quantization: int4 text decoder, group 32; fp16 vision encoder
  • Base model: Qwen/Qwen3.6-35B-A3B
  • Validation status: Validated
  • Measured in v0.5.1: 57.77 tokens/s, 3.95 s to first token, at a 2048-token prompt — all three lengths

Qwen3.8 27B

Added in v0.4.0. A model that did not exist when the compiler was designed, brought up without a new operator library.

  • Kind: Vision-language model
  • Task: Image understanding, text generation
  • Parameters: 27B
  • Architecture: Dense transformer + vision encoder
  • Quantization: int4 k-quant text decoder, group 128; fp16 vision encoder
  • Base model: Qwen/Qwen3.8-27B
  • Validation status: Validated
  • Measured in v0.5.1: 13.60 tokens/s, 10.28 s to first token, at a 2048-token prompt — all three lengths