01 · Creature detail
A feathered dinosaur standing in a sunlit forest.
- Feather
- Backlight
- Motion
One-step high-resolution refinement
Generate a draft. Refine once. Deliver high-resolution video across different base generators.
NVIDIA Research, Efficient AI Team & Singapore Lab. · * Equal contribution
Abstract
High-resolution video generation is costly because inference scales rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck.
SoL-Refiner transforms low-resolution outputs from diverse base generators into 4K videos with a single target-resolution denoising step. Its training recipe combines high-resolution continual training, frame-based RL post-training, and one-step distribution-matching distillation with bidirectional and streaming inference.
H3 deployment acceleration
27.03× wall-clock speedup
Protocol: a five-second, 1344×768 video at 24 fps on NVIDIA GB200 GPUs. Full-resolution MiniMax H3 takes 152.3s on one GPU; our two-stage pipeline uses one GPU per stage, with Stage 1 at 896×512 taking 4.0629s and one-step refinement taking 1.5567s. Measured end-to-end latency is 5.6354s (27.03×). The animation is time-compressed.
Where the refinement shows up
This representative H3 portrait shows how SoL-Refiner refines the video in one target-resolution step, with magnified views highlighting the refined details.
Quality and efficiency
Refiner effect across generators
Drag each divider to compare the same frame with SoL-Refiner OFF and ON. H3 previews are 1080p; all other previews are 720p.
Refiner effect · 8 comparisons
Deployment pipeline: 152.3s full-res H3 vs. 5.64s with Stage 1 + SoL-Refiner (27.03×). Protocol: five-second 1344×768 video at 24 fps using GB200 GPUs.
01 · Creature detail
02 · Portrait
03 · Fine material
04 · Structure + smoke
05 · Human motion
06 · Landscape
07 · Architecture
08 · Nature
Refiner effect · 4 comparisons
Compute pipeline: 20.04s direct at 1248×704 vs. 12.54s with 832×480 generation + one-step refinement (1.60×). Protocol: 1× H100, 81 frames at 16 fps.
01 · Human motion
02 · Costume detail
03 · Portrait
04 · Landscape
Refiner effect · 4 comparisons
A 35-step Cosmos-Nano generator with one-step refinement achieves a 2.81× speedup. Protocol: 1× H100, 189 frames at 24 fps.
01 · Animal macro
02 · Portrait
03 · Fast motion
04 · Nature
Refiner effect · 4 comparisons
Direct generation takes 272.69s at 1280×720, while low-res generation + one-step refinement takes 78.71s at 1280×704 (3.46×). Protocol: 1× H100, 81 frames at 24 fps.
01 · Portrait
02 · Animal motion
03 · Architecture
04 · Vehicle
BibTeX
@techreport{liu2026solrefiner,
title = {SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video},
author = {Haozhe Liu and Tian Ye and Shuchen Xue and Yitong Li and
Junsong Chen and Haopeng Li and Jincheng Yu and Duomin Wang and
Ruihua Zhang and Lei Zhu and Song Han and Enze Xie},
institution = {NVIDIA Research, Efficient AI Team \& Singapore Lab.},
year = {2026},
note = {Technical Report}
}The public paper link will be added after release. Code currently points to the SoL-Engine branch.