MORI

Getting Started

  • Install MORI
    • Installation
      • Prerequisites
      • Validate the environment
      • Install
        • From PyPI (stable release)
        • Nightly (pre-built, tested daily)
        • From source
      • Verify installation
    • Optional FlyDSL integration
    • Next steps
  • Quickstart
    • Runnable checks
    • API sketches
    • MORI-EP: Dispatch and Combine
    • MORI-IO: Point-to-Point Transfers
    • MORI Shmem: Symmetric Memory
    • MORI-IR: Device Bitcode for GPU Kernels
    • Profiling with MORI-VIZ
    • Next Steps
  • Choosing an EP API and backend
    • EPv2 configuration and lifecycle
    • Validation before tuning

User Guides

  • MORI-EP User Guide
    • Table of Contents
    • Quick Reference
    • 1. Kernel Types
    • 2. Configuration
      • EpDispatchCombineConfig
    • 3. Operator API
      • EpDispatchCombineOp
      • dispatch()
      • combine()
      • Split dispatch/combine (send + recv)
      • reset()
      • get_dispatch_src_token_pos()
      • get_registered_combine_input_buffer()
    • 4. Standard MoE Compatibility (DeepEP)
      • dispatch_standard_moe()
      • combine_standard_moe()
      • convert_dispatch_output()
      • convert_combine_input()
    • 5. Initialization
      • Method 1: PyTorch Process Group (Recommended)
      • Method 2: Unique ID (No PyTorch Distributed)
    • 6. Launch Configuration
      • Manual Mode (Default)
      • Auto Mode
    • 7. Complete Example
    • 8. Environment Variables
    • 9. Benchmarking
      • Intra-node
      • Inter-node
      • Reference Performance (DeepSeek V3 config)
    • 10. Tuning
      • Overview
      • Tuning Config Files
      • Intra-node Tuning
      • Inter-node Tuning
      • Tuning on a New GPU Platform
    • 11. Profiling with MORI-VIZ
    • 12. Framework Integration
    • Build Options
    • Source Files
  • MORI Shmem Guide
    • Table of Contents
    • Quick Reference
    • 1. Concepts
      • Symmetric Memory
      • Processing Element (PE)
      • Symmetric Heap
    • 2. Initialization
      • Method 1: PyTorch Process Group (Recommended)
      • Method 2: Unique ID (No PyTorch Distributed)
      • Method 3: MPI Communicator (C++ / MPI environments)
      • Initialization Comparison
      • Finalization
    • 3. Query APIs
    • 4. Memory Management
      • Allocating Symmetric Memory
      • Freeing Symmetric Memory
      • Registering Existing Buffers
    • 5. P2P Address Translation
    • 6. Synchronization
    • 7. HIP Module Init (Triton Integration)
    • 8. Initialization Flags
    • Environment Variables
    • Rail-only connections
    • Source Files
  • MORI CCO Guide
    • Table of Contents
    • Quick Reference
    • 1. Concepts
      • Communicator (ccoComm)
      • Symmetric memory and windows
      • Device communicator (ccoDevComm)
      • Teams
    • 2. Initialization
    • 3. Memory and Windows
    • 4. Device Communicator
    • 5. Host Barrier
    • 6. Device-Side Programming
    • 7. Examples
    • 8. C++ API
    • 9. Environment Variables
  • MORI-IR Guide
    • Table of Contents
    • Architecture
    • 1. Host-Side Python API
      • mori.ir (framework-agnostic)
      • mori.ir.triton (Triton-specific reference backend)
    • 2. Device Bitcode
    • 3. Device Functions
      • Query
      • Point-to-Point
      • PutNbi (Thread / Warp / Block)
      • PutNbi with Signal
      • Immediate Put
      • Atomics
      • Wait
      • Synchronization
    • 4. Integration Example: Triton
    • 5. Integration Example: Raw Bitcode (no framework)
    • 6. Bitcode JIT Compilation
    • 7. Examples
    • 8. Testing
    • Known Limitations
  • MORI-IO Introduction
    • Table of Contents
    • Design & Concepts
    • Workflow
    • Architecture
    • Python API Quick Reference
    • Example: Basic Read/Write
    • Environment & Configuration
    • Profiling MORI-IO with ROCTX Markers
    • Source Files
  • MORI UMBP Single-Node Smoke Test
    • Overview
    • Prerequisites
    • Configuration Inputs
      • Script constants
      • Runtime environment variables
    • Launch Steps
    • Core Dumps
    • Troubleshooting
  • MORI-VIZ (Visualizer)
    • Table of Contents
    • Overview
    • Instrumentation
      • 1. Include Headers
      • 2. Define Profiler Context
      • 3. Add Trace Points
    • Build & Code Generation
      • Enabling the Profiler
      • Code Generation
    • Python Analysis
      • Device-side Elapsed Time Measurement
      • Viewing the Trace
    • Best Practices

Benchmarks

  • MORI-EP Benchmark
    • Table of Contents
    • Intra-node
      • Profiling intranode kernels
    • Inter-node
      • NIC Selection
      • Bandwidth Computation
      • RDMA bandwidth
        • Algo BW vs Physical BW
      • NVL / XGMI bandwidth
      • Why DeepEP and Mori RDMA token counts are equivalent
      • Combine
      • Reference: print_bw_values.py
      • Profiling Internode Kernels
      • Analyzing traces with analyze_ep_kernel_trace.py
  • MORI GEMM + All-Reduce Benchmark
    • Table of Contents
    • Running the benchmark
    • Headline
    • Where the time goes
    • The fp8 wire
      • Who moves the gather
      • The pull grid
      • What fp8 costs, numerically
    • Model-level evaluation
    • Negative results
    • Reproducing the end-to-end numbers
  • MORI-IO Benchmark
    • Table of Contents
    • Which benchmark to use
    • C++ Benchmark (default, nixlbench-matching)
      • Build
      • Run (two nodes)
      • Sweep modes
      • Differences from the Python flags
    • Python Benchmark
      • Fabric (cross-node scale-up UALink super-node)
    • Benchmark Arguments
    • Results: Thor2 RDMA Read
    • Results: Thor2 RDMA Write
      • Message Size Sweep
      • Batch Size Sweep
    • Results: CX7 RDMA (Batch Size = 1)
      • Write
      • Read
    • Benchmark Tuning Guide
      • Recommended starting config
      • Parameters that help
      • The merged-WR size limit
      • Parameters that usually do not help (or can hurt)
      • Host (CPU) vs GPU memory
  • MORI UMBP PD Disaggregation Benchmark
    • Table of Contents
    • Prerequisites
    • Topology
    • Configuration
    • Step 1 — Start Docker on all nodes
    • Step 2 — Start Grafana and Prometheus
    • Step 3 — Build mori on all nodes
    • Step 4 — Create launcher scripts
    • Step 5 — Kill stale processes and launch prefill
    • Step 6 — Wait for ALL prefills ready
    • Step 7 — Launch decode and benchmark
    • Step 8 — Monitor decode and benchmark
    • Step 9 — Show results
    • Teardown
    • Environment Variable Reference
      • Topology (this guide’s wrapper variables)
      • UMBP / mori (consumed by run_pd_disagg_bench_dp8ep8.sh)
    • Troubleshooting
  • UMBP Distributed Tier-Management Benchmark
    • Build
    • Run, generate, and replay
    • JSON backend and tier policy
    • Synthetic profiles
    • Trace contract
    • Results
  • RDMA Bandwidth Utilization in Dispatch/Combine Kernels
    • Overview
    • Kernel Execution Path
    • What RDMA Actually Transfers
      • Dispatch send (DispatchInterNodeSend / DispatchInterNodeLLSend)
      • Combine send (CombineInterNodeTyped / CombineInterNodeLLTyped)
    • Theoretical Peak vs. Observed Throughput
      • Computing theoretical peak
      • Understanding the receive-side timing
    • How to Check RDMA Bandwidth from the Trace
      • Step 1: Collect a profiler trace
      • Step 2: Read key durations from the text summary
      • Step 3: Compute observed RDMA bandwidth
      • Step 4: Interpret results
    • Key Tuning Parameters That Affect RDMA Utilization
      • v1 vs. v1_ll (low-latency) tradeoff
    • Sanity Checks from Hardware Counters (optional)
    • Summary

Developer Guides

  • Mori IR Integration
  • MORI JIT Compilation Framework
    • Quick Start
    • Architecture Overview
    • Packaging: How Wheel Installs Support JIT
    • Compilation Split: Host CXX vs Device JIT
      • What compiles as CXX (at pip install time)
      • What compiles as HIP (runtime JIT only)
      • Template Args vs Raw Args
    • What Gets JIT-Compiled
    • Kernel Launch Flow
    • HIP Driver API (Python ctypes)
    • Triton Integration
    • Cache Structure
    • Environment Variables
    • Testing
      • Docker Environment
      • 1. Non-Editable Install (Wheel)
      • 2. Editable Install (Development)
      • 3. Pre-compile All Kernels
      • 4. Verify JIT Configuration
      • 5. Verify Compilation Separation
      • 6. Dispatch/Combine Correctness
      • 7. Dispatch/Combine Benchmark
      • 8. Shmem API
      • 9. Triton Integration
      • 10. Full Clean-Slate Test
    • Kernel Source Files
    • Adding a New Kernel
    • Host/Device NIC Macro Separation
    • ROCm Version Compatibility
  • MORI JIT v2
    • 1. 核心不变量
    • 2. 分层
    • 3. 一个 kernel 的三件套
      • 3.1 Cfg 作为单一模板参数(C++20 structural NTTP)
      • 3.2 字段列表的完整性是承重的
      • 3.3 生成的源码
      • 3.4 缓存
      • 3.5 host/device 常量只有一份,且 host 不需要 hipcc
      • 3.6 环境变量
      • 3.7 dtype 是传输属性,不是算术属性
    • 4. op 层
    • 5. 两个后端共用一个 op
      • 5.1 几何调优表:每个后端一份,dispatch 和 combine 再分开
    • 6. Python 绑定
    • 7. 几个有证据的决定
    • 8. 性能
    • 9. 移植来源与覆盖范围
    • 10. 测试

API Reference

  • Communication APIs
    • MORI-EP (Expert Parallelism)
    • MORI Shmem (Symmetric Memory)
    • MORI-IO (Point-to-Point I/O)
    • MORI-IR (Device Bitcode Integration)
  • MORI-VIZ Kernel Profiler
    • Requirements
    • Usage
    • export_to_perfetto
    • C++ Profiler APIs
  • UMBP Master Client
    • Starting the Master Server
    • UMBPTierType
    • UMBPExternalKvNodeMatch
    • UMBPMasterClient
    • Cache Hit Tracking
      • Finding hot KV blocks (the typical workflow)
      • Lifetime cumulative — not a rate
    • UMBPExternalKvHitCountEntry
    • Usage Examples
    • End-to-End Example
MORI
  • Search


© Copyright 2026, AMD. Last updated on 2026-09-18.

Built with Sphinx using a theme provided by Read the Docs.