This project implements memory-efficient attention and fused CUDA kernels for GPT-2 inference, with optimization decisions driven by Nsight Compute and Nsight Systems profiles.
Implementation
- Avoided materializing the full attention matrix using tiled computation and online softmax.
- Reduced HBM traffic by roughly 10× and raised arithmetic intensity from 15 to 157 FLOP/byte.
- Fused attention-related kernels to cut launch overhead by roughly 50%.
- Achieved up to 9% end-to-end speedup through profiling-driven tuning.
Read the final report or explore the implementation.