One-step high-resolution refinement

SoL-RefinerSpeed-of-Light One-Step Refinement for High-Resolution Video

Generate a draft. Refine once. Deliver high-resolution video across different base generators.

  • Haozhe Liu*
  • Tian Ye*
  • Shuchen Xue*
  • Yitong Li
  • Junsong Chen
  • Haopeng Li
  • Jincheng Yu
  • Duomin Wang
  • Ruihua Zhang
  • Lei Zhu
  • Song HanEnze Xie

NVIDIA Research, Efficient AI Team & Singapore Lab. · * Equal contribution

Code ↗AbstractSamplesPaper Coming soon

Abstract

High-resolution video generation is costly because inference scales rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck.

SoL-Refiner transforms low-resolution outputs from diverse base generators into 4K videos with a single target-resolution denoising step. Its training recipe combines high-resolution continual training, frame-based RL post-training, and one-step distribution-matching distillation with bidirectional and streaming inference.

H3 deployment acceleration

152.3 seconds → 5.64 seconds

27.03× wall-clock speedup

Full-res H3152.3s
Stage 14.0629s
Refiner service1.5567s
Total latency5.6354s

Protocol: a five-second, 1344×768 video at 24 fps on NVIDIA GB200 GPUs. Full-resolution MiniMax H3 takes 152.3s on one GPU; our two-stage pipeline uses one GPU per stage, with Stage 1 at 896×512 taking 4.0629s and one-step refinement taking 1.5567s. Measured end-to-end latency is 5.6354s (27.03×). The animation is time-compressed.

A closer look at H3 refinement.

This representative H3 portrait shows how SoL-Refiner refines the video in one target-resolution step, with magnified views highlighting the refined details.

Quality and efficiency

Cross-generator support. Two-stage generation lowers latency while improving average quality for WAN and Cosmos-Nano. The measurements use one H100 GPU. Open the chart for a full-size view.
High-resolution quality. One-step refinement remains competitive at 2K and improves both reported metrics over LTX-2.3 at 4K. Open the chart for a full-size view.

Refiner effect across generators

One-step refinement. Four generators.

Drag each divider to compare the same frame with SoL-Refiner OFF and ON. H3 previews are 1080p; all other previews are 720p.

BibTeX

@techreport{liu2026solrefiner,
  title       = {SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video},
  author      = {Haozhe Liu and Tian Ye and Shuchen Xue and Yitong Li and
                 Junsong Chen and Haopeng Li and Jincheng Yu and Duomin Wang and
                 Ruihua Zhang and Lei Zhu and Song Han and Enze Xie},
  institution = {NVIDIA Research, Efficient AI Team \& Singapore Lab.},
  year        = {2026},
  note        = {Technical Report}
}

The public paper link will be added after release. Code currently points to the SoL-Engine branch.