ATOM Logo

Getting Started

  • Installation
    • Requirements
    • Installation methods
      • From source
      • Docker
    • Environment variables
    • Verify the installation
    • Troubleshooting
  • Quick start
    • Offline inference
      • Single prompt
      • Batch inference
    • Distributed inference
    • API server
    • Performance tips
    • Next steps
  • Model run guide
    • Quick start
    • Supported models
      • vLLM plugin backend
    • Nightly CI benchmark configurations
    • Live dashboard

User Guides

  • ATOM architecture guide
    • Quick reference
    • System overview
    • Component architecture
    • Request lifecycle
    • Forward context pattern
    • Multi-process architecture
    • Sequence lifecycle
    • Speculative decoding (MTP / EAGLE)
      • Supported MTP architectures
      • Weight sharing
      • Propose loop: MHA vs MLA branching
      • prepare_mtp_decode()
    • Source files
  • ATOM configuration guide
    • Quick reference
    • Master configuration (Config)
    • Compilation configuration (CompilationConfig)
      • Compilation levels (CompilationLevel)
      • CompilationConfig fields
      • CUDA graph mode (CUDAGraphMode)
    • Quantization configuration (QuantizationConfig & LayerQuantConfig)
      • LayerQuantConfig fields
      • QuantizationConfig attributes
      • QuantType values (from AITER)
      • Supported quantization dtypes
      • Auto-detection from HuggingFace
      • Layer-level quantization dispatch
      • Online quantization at load time
    • Parallel configuration (ParallelConfig)
    • Speculative decoding configuration (SpeculativeConfig)
      • Table-driven MTP config
      • Post-init behaviour (hf_config_override)
    • Sampling parameters (SamplingParams)
    • CLI arguments (EngineArgs)
    • Environment variables
      • Variables registered in atom/utils/envs.py
      • Additional environment variables (used outside envs.py)
    • Decision tree — choosing a compilation level
    • Source files
  • ATOM Model Support Guide
    • Quick Reference
    • Supported model architectures
    • Model architecture details
      • Qwen3 (Qwen3ForCausalLM)
      • Qwen3-MoE (Qwen3MoeForCausalLM)
      • Llama (LlamaForCausalLM)
      • Mixtral (MixtralForCausalLM)
      • DeepSeek V2/V3 (DeepseekV2ForCausalLM)
      • DeepSeek MTP (DeepSeekMTP)
      • GPT-OSS (GptOssForCausalLM)
      • GLM-5 / GlmMoeDsa (GlmMoeDsaForCausalLM)
      • GLM4-MoE (Glm4MoeForCausalLM)
      • Qwen3-Next (Qwen3NextForCausalLM)
      • Qwen3.5 (Qwen3_5ForConditionalGeneration and Qwen3_5MoeForConditionalGeneration)
      • Qwen3.5 MTP (Qwen3_5MTP)
    • Weight loading
      • Function signature
      • Loading flow
      • Batched expert staging
      • Layers beyond num_hidden_layers
    • Adding a new model
      • Step 1: Create the model file
      • Step 2: Implement layer classes
      • Step 3: Implement the model and CausalLM classes
      • Step 4: Register the model
      • Step 5: Handle weight loading
    • Model-specific optimizations
      • Llama: fused RMSNorm+Quant and SiLU+Mul+Quant
      • DeepSeek V2/V3: MLA + fused input norm + QK norm fusion
      • Qwen3-MoE: QK norm + RoPE + cache + quant fusion
      • MTP: Multi-token prediction (speculative decoding)
    • Source files
  • ATOM model operations guide
    • Quick reference
    • AITER integration overview
      • AITER kernel mapping table
    • Linear operations
      • Class hierarchy
      • Quantization dispatch
      • Tensor parallel sharding
      • Weight processing
    • Attention operations
      • Base: Attention (base_attention.py)
      • Multi-head attention (attention_mha.py)
      • Multi-head latent attention (attention_mla.py)
      • Backend abstraction (attentions/backends.py)
      • KV cache operations
    • Mixture of experts (MoE)
      • FusedMoE class (moe.py)
      • Quantization methods
      • TopK routing (topK.py)
      • FusedMoEParallelConfig
      • MORI integration (fused_moe/mori_prepare_finalize.py)
      • MoE quantization config (fused_moe/config.py)
      • Triton MoE fallback (fused_moe_triton.py)
    • Normalization
      • RMSNorm (layernorm.py)
      • LayerNorm (layernorm.py)
    • Activation functions
      • SiluAndMul (activation.py)
    • Embedding and output head
      • VocabParallelEmbedding (embed_head.py)
      • ParallelLMHead (embed_head.py)
    • Rotary position embedding (RoPE)
      • RotaryEmbedding (rotary_embedding.py)
      • get_rope() factory
      • Integration in attention
    • Sampling
      • Sampler (sampler.py)
      • RejectionSampler (rejection_sampler.py)
    • Fused kernel chains
    • Source files
      • atom/model_ops/
      • atom/model_ops/attentions/
      • atom/model_ops/fused_moe/
      • atom/utils/
  • ATOM scheduling & KV cache guide
    • Quick reference
    • Scheduling algorithm
      • Scheduler initialization
      • Schedule flow
      • Delay factor
      • Preemption
    • ScheduledBatch structure
      • Constructor signature
      • Fields
      • ScheduledBatchOutput
    • Block manager
      • Block class
      • BlockManager initialization
      • Allocation (allocate)
      • Deallocation (deallocate)
      • Can-allocate and can-append checks
      • May-append (decode extension)
    • Prefix caching
      • Hash function
      • Hash chaining
      • Cache lookup during allocation
      • Reference counting
      • Enabling prefix caching
    • Postprocessing
      • Signature
      • Token appending
      • Stop condition checking
      • Stream output
      • Sequence cleanup
      • Placeholder insertion
    • Speculative decoding integration
      • Scheduler tracking
      • Draft tokens in scheduling
      • Acceptance statistics
      • Draft token storage on sequences
    • Sequence management
      • Constructor
      • Core fields
      • Timing fields
      • Computed properties
      • num_tokens setter
      • Lifecycle
      • SequenceStatus enum
      • SequenceType enum
    • Source files
  • ATOM Distributed Inference Guide
    • Quick Reference
    • Tensor Parallelism (TP)
      • Weight Sharding
      • Process Group Initialization
      • AllReduce
      • Configuration
    • Data Parallelism (DP)
      • Architecture
      • DP Process Group Initialization
      • Synchronized busy loop
      • Dummy batch execution
      • Device assignment
      • DPMetadata
      • CoreManager (DP orchestration)
      • DP request load balancing
      • Configuration
    • Expert Parallelism (EP)
      • FusedMoEParallelConfig
      • Expert distribution
      • MORI communication
      • Configuration
    • Environment variables
    • Multi-GPU deployment examples
      • DeepSeek-R1 on 8 GPUs (TP8)
      • Qwen3-235B-A22B on 8 GPUs (TP8 + EP)
      • Kimi-K2-Thinking on 4 GPUs (TP4)
    • Combined parallelism strategies
      • TP only (dense models)
      • TP + EP (MoE models)
      • TP + DP (dense throughput)
      • TP + DP + EP (MoE throughput)
      • DP attention mode
    • Source files
  • Prefill Context Parallel (PCP) Guide
    • When to use PCP
    • Quick Reference
    • CLI usage
    • Launching server
      • DeepSeek-V4: TP4 + PCP2 (8 GPUs)
      • DeepSeek-V4: TP1 + PCP8 + prefill TBO overlap (8 GPUs)
    • Tuning for long sequences
    • Performance baseline
      • PCP + TBO (prefill overlap)
    • Constraints & Compatibility
    • How it works
    • Source Files
  • ATOM compilation & CUDA graphs guide
    • Compilation levels
      • Level 0 — NO_COMPILATION
      • Level 1 — DYNAMO_AS_IS
      • Level 2 — DYNAMO_ONCE
      • Level 3 — PIECEWISE (production default)
    • CUDA graph modes
      • NONE (value: 0)
      • PIECEWISE (value: 1)
      • FULL (value: 2)
      • FULL_DECODE_ONLY (value: (FULL, NONE))
      • FULL_AND_PIECEWISE (value: (FULL, PIECEWISE))
      • Helper methods
    • CUDA graph capture
      • Capture flow
      • Graph keying
      • Graph pool sharing
      • Default capture sizes
      • Graph replay in run_model()
    • Piecewise compilation
      • Splitting operations
      • Compilation pipeline
      • Cache management
    • Forward context & stateless dispatch
      • ForwardContext fields
      • Lifecycle
      • Context dataclass
      • Integration with CUDA graphs
    • Compiler backend
      • CompilerManager
      • CompilerInterface
      • InductorAdaptor
      • InductorStandaloneAdaptor
      • VllmBackend
      • @support_torch_compile decorator
      • Custom op registration
    • Configuration options
    • Decision tree
      • Common configurations
    • Source files
  • ATOM serving & benchmarking guide
    • Quick reference
    • OpenAI-compatible server
      • Endpoints
      • Request models
      • Response models
      • Server startup
      • Example: curl
    • Programmatic API (LLMEngine)
      • Initialization
      • SamplingParams
      • Core methods
      • Synchronous generation example
      • Asynchronous / streaming usage
    • Simple inference
      • Usage
      • What it does
    • Benchmarking
      • Metrics
      • Key CLI arguments
      • Backend request functions
      • Full benchmark example
    • Profiling
      • Configuration
      • Online profiling (HTTP)
      • Programmatic profiling
      • Offline profiling script
      • Profiling during benchmarks
    • Speculative decoding (MTP)
      • Architecture
      • Configuration
      • MTP statistics
      • How rejection sampling works
    • Deployment examples
      • Single-GPU
      • Multi-GPU with tensor parallelism
      • Docker deployment
      • Engine CLI arguments (EngineArgs)
    • Accuracy validation
      • Setup
      • Run evaluation
    • Source files

Framework Integrations

  • vLLM-ATOM backend
    • Architecture
      • Design overview
      • How it works
      • Plugin lifecycle
      • Key modules
      • Component diagram
    • Configuration translation
      • PluginConfig fields
      • vLLM config mapping
    • Attention integration
      • How the backend is selected
    • Supported models
    • Installation and quick start
      • Prerequisites
      • Set up the environment
      • Launch vLLM with ATOM plugin
      • Benchmark serving
      • Enable profiling
      • Disable ATOM plugin
    • Environment variables
      • Online quantization

API Reference

  • Serving API
    • LLMEngine class
      • Methods
        • generate()
    • SamplingParams
    • Return values
    • Example
  • Supported models
    • Llama models
    • GPT models
    • Mixtral
    • Other architectures
    • Model configuration
    • Performance by model size
    • Quantization
ATOM
  • Search


© Copyright 2026, AMD.

Built with Sphinx using a theme provided by Read the Docs.