LTX 2.3
1080p two-stage generation
Paper overview
Training-free on-the-fly attention sparsification for efficient visual generation
Training-free dynamic sparse attention accelerates pretrained visual generators, but existing methods rely on costly offline routing and discard every unselected block. Sol-Attn unifies query-dependent threshold routing, sparse computation, and approximate correction within a single online-softmax pass. It selects blocks on chip without materializing a proxy map, then reuses below-threshold score columns to retain the long-tail contribution of skipped blocks.
Across image and video generation, Sol-Attn advances the quality-efficiency frontier of training-free sparse attention, delivering up to 2.1× speedup for video generation and 2.3× for video editing while preserving visual quality. Integrated with complementary acceleration techniques in Sol-Engine, it reaches up to 5× end-to-end speedup.
Dense vs. Sol-Attn
Dense-attention and Sol-Attn outputs, aligned case by case.
1080p two-stage generation
720p video editing
720p text-to-video
720p text-to-video
Outputs use identical prompts and seeds; speedups are end-to-end over the corresponding dense-attention pipeline.
Paper
Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
@misc{li2026solattnacceleratingvideogeneration,
title={Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification},
author={Haopeng Li and Yitong Li and Junsong Chen and Tian Ye and Haozhe Liu and Jincheng Yu and Duomin Wang and Ruihua Zhang and Zeke Xie and Enze Xie and Song Han},
year={2026},
eprint={2607.24027},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.24027},
}