Episode
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
- Podcast
- Daily Paper Cast
- Published
- Jul 29, 2026
- Duration seconds
- 1321
- Processing state
not_requested- Canonical source
- https://share.transistor.fm/s/46907ae6
Actions
POST https://stenobird.com/v1/public/podcasts/daily-paper-cast-7079649/episodes/sol-attn-accelerating-video-generation-inference-via-on-the-fly-attention-sparsification/transcription-requests
Idempotently request low-priority transcript generation for this episode.GET https://stenobird.com/podcast/daily-paper-cast-7079649/sol-attn-accelerating-video-generation-inference-via-on-the-fly-attention-sparsification.md
Read the agent-friendly Markdown representation of this episode resource.
Summary
🤗 Upvotes: 27 | cs.CV Authors: Haopeng Li, Yitong Li, Junsong Chen, Tian Ye, Haozhe Liu, Jincheng Yu, Duomin Wang, Ruihua Zhang, Zeke Xie, Enze Xie, Song Han Title: Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Arxiv: http://arxiv.org/abs/2607.24027v1 Abstract: Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas retaining blocks to reach a target cumulative proxy probability mass yields dynamic but potentially imbalanced budgets; both incur non-negligible overhead from computing and materializing proxy scores. (2) Lossy keep-or-drop sparsification: unselected blocks are discarded entirely, degrading accuracy under aggressive sparsity. These limitations motivate cheaper dynamic-budget routing while limiting accuracy degradation. In this paper, we introduce training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention. The core of Sol-Attn is on-the-fly block thresholding with proxy-score reuse, which selects critical blocks by comparing block proxy scores against a threshold during online softmax. This design enables dynamic yet controllable block budgets without materializing the proxy map, while directly reusing the…