FlashAttention: Mathematical Foundations, Online Softmax Tiling, IO-Awareness, and Exact Attention Scaling
Standard multi-head self-attention in the Transformer architecture exhibits quadratic time and memory complexity with respect to sequence length $N$. While the $O(N^2)$ computational complexity is widely cited, the primary performance bottleneck in production hardware is not arithmetic throughput, but memory access overhead. On modern GPU architectures such as NVIDIA A100 and H100, tensor processing cores execute matrix multiplications at teraflop and petaflop scales, but memory bandwidth betwee
1 min
