Llama
The dense decoder most of the ecosystem's tooling assumes.
Llama is the baseline shape. It is a plain dense transformer decoder, it is what most exporters, quantizers and inference harnesses were written against first, and it is the model to try when you want to know whether a problem is your model or your setup.
The v0.5.1 matrix also includes a Llama-shaped decoder under a different name: DeepSeek-R1-Distill-Llama 70B is a reasoning distillation onto this architecture at roughly nine times the size.
Llama 3.1 8B
- Kind: Language model
- Task: Text generation
- Parameters: 8B
- Architecture: Dense transformer
- Quantization: int4 AWQ, group 128, asymmetric
- Base model: meta-llama/Llama-3.1-8B
- Validation status: Validated
- Measured in v0.5.1: 43.14 tokens/s, 1.50 s to first token, at a 2048-token prompt — all three lengths