Home
Preface
1. What is MLSys?
An intuitive understanding
Comparison to traditional systems
Formal definition
2. Matrix multiplication
Tensor
Ops
Matrix multiplication (math)
Matmul (a naive implementation)
3. Parallel computing
Multi-threading
Race condition
Sync call
Async call
Atomic operations
Barrier
4. Memory
Memory wall
Host memory vs. device memory
Memory hierarchy
Bandwidth
Latency
5. GPU architecture
The GPU
The streaming multiprocessor
GPU programming model
GPU memory hierarchy
6. Performance basics
Pipelining
Memory-bound
Compute-bound
Op fusion
7. Performance metrics
Throughput
Arithmetic intensity
Arithmetic intensity (math)
Roofline model
8. Optimizing matmul
Naive tiling
Optimized tiling
Shared memory utilization
Tiling memory optimization
Matmul GPU Kernel
Tensor Cores
To be continued ...
MLSYS ENGINEERING
To be continued ...