NVIDIA Research · Efficient AI Team & Singapore Lab

Sol-AttnAccelerating Video Generation Inference via On-the-Fly Attention Sparsification

Haopeng Li, Yitong Li, Junsong Chen, Tian Ye, Haozhe Liu, Jincheng Yu, Duomin Wang, Ruihua Zhang, Zeke Xie, Enze Xie, Song Han

NVIDIA

Wan 2.1
Dense
2.02×
Sol-Attn
Hunyuan 1.5
Dense
2.12×
Sol-Attn
LTX 2.3
Dense
1.90×
Sol-Attn
Bernini
Dense
2.34×
Sol-Attn

Paper overview

Abstract

Training-free on-the-fly attention sparsification for efficient visual generation

Training-free dynamic sparse attention accelerates pretrained visual generators, but existing methods rely on costly offline routing and discard every unselected block. Sol-Attn unifies query-dependent threshold routing, sparse computation, and approximate correction within a single online-softmax pass. It selects blocks on chip without materializing a proxy map, then reuses below-threshold score columns to retain the long-tail contribution of skipped blocks.

Across image and video generation, Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering up to 2.1× speedup for video generation and 2.3× for video editing while preserving visual quality. Integrated with complementary acceleration techniques in Sol-Engine, it reaches up to 5× end-to-end speedup.

Dense vs. Sol-Attn

Visual comparisons.

Dense-attention and Sol-Attn outputs, aligned case by case.

Paper

Sol-Attn

Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

BibTeX
@misc{li2026solattnacceleratingvideogeneration,
      title={Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification},
      author={Haopeng Li and Yitong Li and Junsong Chen and Tian Ye and Haozhe Liu and Jincheng Yu and Duomin Wang and Ruihua Zhang and Zeke Xie and Enze Xie and Song Han},
      year={2026},
      eprint={2607.24027},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.24027},
}