+50.8%
FlashAttention-2 on A100
Best case from an automated run; 25 to 51% at sequence length 1,024 and up with batch size 4 and up.
Published results
Every number on this site comes from a public report. Each is listed here with the conditions it was measured under.
+50.8%
Best case from an automated run; 25 to 51% at sequence length 1,024 and up with batch size 4 and up.
+13.3%
Over the open-source kernel, after an agent studied traces of the closed-source cuDNN kernel. 5.2 to 13.3% across shapes.
+4.2%
At best, on a kernel already above 90% of peak throughput. Typically 1 to 3%. Submitted upstream.
+1.3%
Average across attention modes, precisions and sequence lengths from 1k to 32k, with no accuracy regression.
We gave a coding agent one task four times: raise the throughput of FlashAttention-3 on an H100. Each session read a different profile. Everything else stayed the same.
All four sessions ended near the same plateau, about 566 TFLOPS. The chart shows the first iteration at which each came within 2.5 TFLOPS of it. The existing traces reported a wait that their own probes had created, which sent the agent in the wrong direction.
FlashAttention-3 on NVIDIA H100. Source: G-Watch, September 2026.
Anatomy of one optimization
One real run, on the FlashAttention-4 forward kernel for NVIDIA Blackwell. Step through it.
The starting kernel runs at about 1,541 TFLOPS at batch 8 and sequence length 8,192. It is already a heavily tuned kernel.
Range profiling shows the tensor pipe, which does the matrix multiplies, is only 76.88% utilized. Something is making it wait.
Tracing inside the kernel finds it. The matrix-multiply warp finishes an iteration in about 1,712 ns, then waits for a softmax stage that takes about 3,091 ns. Binary analysis shows why: every softmax warp is queuing for the same hardware exp2 unit.
The fix moves more of the exponentials off that unit and onto fused multiply-add units that were underused, by retuning two parameters of the kernel's software emulation. No C++ rebuild.
Throughput rises by 1.3% on average across every tested configuration, with no accuracy regression. Four other ideas were tried and rejected: two hung the kernel and two made it slower.
Nanoseconds, approximate. FlashAttention-4 forward on Blackwell. Source: engineering report, March 2026.
On B300, cuDNN's attention kernel ran faster than open-source FlashAttention-4, and cuDNN ships no source. We traced both on the same inputs and asked an agent to close the gap using only the two traces.
| What the trace showed | What the agent changed | Gain |
|---|---|---|
| The kernel ran its thread blocks in four waves, with idle time between them. | Made the grid persistent, so each block fetches its next tile in place. | +9.4% |
| Persistent blocks finished up to 1.49× apart. | Sorted tiles across all groups, not only within each group. | +2.0% |
| Long tiles were paired with long tiles. | Used a fixed schedule on those shapes, pairing long with short. | +11.1% |
| After the last tile, the kernel still waited 544 ns to free a buffer nothing would use. | Dropped the final wait. | +1.2% |
Gains as reported for each change; not every change applies to every shape. Together they lift FlashAttention-4 by 5.2 to 13.3% across four shapes. Source: G-Watch, September 2026.
| Kernel | Hardware | Result | Conditions | Source |
|---|---|---|---|---|
| FlashAttention-2 forward | NVIDIA A100 | up to +50.8% | 25 to 51% at sequence length 1,024 and up with batch size 4 and up. Two small configurations were 7 to 8% slower. | March 2026 |
| FlashAttention-3 forward | NVIDIA H200 | up to +4.19% | Typically 1 to 3%, on a baseline already above 90% of peak throughput. Submitted to the FlashAttention-3 repository. | March 2026 |
| FlashAttention-4 forward | NVIDIA Blackwell | +1.3% average | Across attention modes, two precisions and sequence lengths from 1k to 32k. No accuracy regression. | March 2026 |
| FlashAttention-4 | NVIDIA B300 | +5.2 to +13.3% | Four shapes, from traces of cuDNN's closed-source attention kernel. | September 2026 |
| FlashAttention-3, agent run | NVIDIA H100 | 3.9× sooner | Iteration 14 with Xtrace, against 54 with the best existing tracer. | September 2026 |
Talk to the optimization team, or start with the open-source tools.