DeepSeek
A reasoning distillation onto a 70B Llama-shaped decoder — the largest dense graph in the matrix.
DeepSeek-R1-Distill-Llama 70B is a reasoning model distilled onto the Llama architecture, so it brings nothing architecturally new to the matrix. What it brings is size: it is the largest dense graph in the v0.5.1 matrix, and a dense model reads every parameter it holds on every token.
That makes it the memory-planning case. Where a sparse model of similar total size gets away with touching a fraction of itself per token, this one does not, and it is included for exactly that reason.
DeepSeek-R1-Distill-Llama 70B
The largest dense model in the matrix. Included because a 70B dense decoder is where memory planning stops being an implementation detail.
- Kind: Language model
- Task: Text generation, reasoning
- Parameters: 70B
- Architecture: Dense transformer
- Quantization: int4 AWQ, block 128
- Base model: deepseek-ai/DeepSeek-R1-Distill-Llama-70B
- Validation status: Validated
- Measured in v0.5.1: 5.93 tokens/s, 12.03 s to first token, at a 2048-token prompt — all three lengths