← Papers & projects

Systems project Final report

FlashAttention-style CUDA Optimization

Memory-efficient kernels and fusion for GPT-2 inference

FlashAttention-style CUDA Optimization
10×
Lower HBM traffic
157
FLOP per byte
+9%
End-to-end speedup

This project implements memory-efficient attention and fused CUDA kernels for GPT-2 inference, with optimization decisions driven by Nsight Compute and Nsight Systems profiles.

Implementation

  • Avoided materializing the full attention matrix using tiled computation and online softmax.
  • Reduced HBM traffic by roughly 10× and raised arithmetic intensity from 15 to 157 FLOP/byte.
  • Fused attention-related kernels to cut launch overhead by roughly 50%.
  • Achieved up to 9% end-to-end speedup through profiling-driven tuning.

Read the final report or explore the implementation.