flashattention2-230tflops-a100

IN premiseentries/2026/06/21/wiki-Transformer_deep_learning_architecture-chunk-5.md

Created 2026-06-21T09:55:56+00:00

FlashAttention-2 achieves up to 230 TFLOPS/s on A100, 2x faster than v1 and 9x over standard PyTorch attention.