Llama

The dense decoder most of the ecosystem's tooling assumes.

Llama is the baseline shape. It is a plain dense transformer decoder, it is what most exporters, quantizers and inference harnesses were written against first, and it is the model to try when you want to know whether a problem is your model or your setup.

The v0.5.1 matrix also includes a Llama-shaped decoder under a different name: DeepSeek-R1-Distill-Llama 70B is a reasoning distillation onto this architecture at roughly nine times the size.

Llama 3.1 8B

  • Kind: Language model
  • Task: Text generation
  • Parameters: 8B
  • Architecture: Dense transformer
  • Quantization: int4 AWQ, group 128, asymmetric
  • Base model: meta-llama/Llama-3.1-8B
  • Validation status: Validated
  • Measured in v0.5.1: 43.14 tokens/s, 1.50 s to first token, at a 2048-token prompt — all three lengths