Home
Preface
What is MLSys?
An intuitive understanding
Comparison to traditional systems
Formal definition
Matrix multiplication
Tensor
Ops
Matrix multiplication (math)
Matmul (a naive implementation)
Parallel computing
Multi-threading
Race condition
Sync call
Async call
Atomic operations
Barrier
Memory
Memory wall
Host memory vs. device memory
Memory hierarchy
Bandwidth
Latency
GPU architecture
The GPU
The streaming multiprocessor
GPU programming model
GPU memory hierarchy
Performance basics
Pipelining
Memory-bound
Compute-bound
Op fusion
Arithmetic intensity
Throughput
Roofline model
Optimizing matmul
Naive tiling
Optimized tiling
Shared memory utilization
Tiling memory optimization
Matmul GPU Kernel
Tensor Cores
To be continued ...
MLSYS ENGINEERING
To be continued ...