DeepSeek

A reasoning distillation onto a 70B Llama-shaped decoder — the largest dense graph in the matrix.

DeepSeek-R1-Distill-Llama 70B is a reasoning model distilled onto the Llama architecture, so it brings nothing architecturally new to the matrix. What it brings is size: it is the largest dense graph in the v0.5.1 matrix, and a dense model reads every parameter it holds on every token.

That makes it the memory-planning case. Where a sparse model of similar total size gets away with touching a fraction of itself per token, this one does not, and it is included for exactly that reason.

DeepSeek-R1-Distill-Llama 70B

The largest dense model in the matrix. Included because a 70B dense decoder is where memory planning stops being an implementation detail.

  • Kind: Language model
  • Task: Text generation, reasoning
  • Parameters: 70B
  • Architecture: Dense transformer
  • Quantization: int4 AWQ, block 128
  • Base model: deepseek-ai/DeepSeek-R1-Distill-Llama-70B
  • Validation status: Validated
  • Measured in v0.5.1: 5.93 tokens/s, 12.03 s to first token, at a 2048-token prompt — all three lengths