Search papers, labs, and topics across Lattice.
This paper introduces FVAttn, an adaptive sparse attention system designed to enhance the efficiency of video generation using Video Diffusion Transformers by addressing the bottleneck of self-attention in high-resolution sequences. By implementing runtime load balancing and a combination of Top-$p$ routing and video-aware block organization, FVAttn significantly reduces workload imbalance and improves distributed execution under multi-GPU settings. The key result shows a remarkable $4.41\times$ speedup in attention processing compared to FlashAttention, while maintaining competitive video quality during inference.
Runtime load balancing in FVAttn slashes attention processing time by over four times, transforming video generation efficiency.
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creates uneven per-head workloads under multi-GPU sequence parallelism. The resulting workload heterogeneity turns sparse attention into a rank-level straggler problem. We present \method{}, a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism. \method{} uses Top-$p$ routing, a Top-$k$ safety floor, and video-aware block organization as the sparse-routing frontend, then repairs the materialized mask at runtime. Runtime Load Balancing migrates a small number of heavy heads via P2P communication to shorten the current critical path. Slack-Aware Sparse Augmentation fills residual non-critical-rank slack with additional high-value blocks, while overlap hides scheduling and migration overhead behind existing computation. On step-distilled Wan2.2 I2V, \method{} reduces average load imbalance from 1.34 to 1.08 and delivers a $4.41\times$ attention speedup over FlashAttention, while achieving a $2.02$--$2.11\times$ DiT inference speedup with competitive video quality.