LongLive-Plug: Once-for-All Distillation for Video Generation
NVIDIA
* Equal contribution
Distill once. Unlock every task.
Developers build specialized video models on top of a shared backbone. Why should every downstream task need its own distillation?
Meet LongLive-Plug ↓Task-specific distillation
0 / 3 readyTrain every task
LongLive-Plug
0 / 6 readyNo downstream training
02 / TRANSFER COVERAGE
A little plug.
A much wider world.
Diverse downstream tasks across backbone families.
One adapter set per backbone, reused across its family’s tasks.
Each pair preserves the source task’s settings. Expand a family to explore more downstream tasks. Open a card for its prompt, sampling settings and notes. Native schedules vary; MiniMax-H3 uses its published four-forward configuration.
03 / THE METHOD
One backbone.
Three reusable capabilities.
Distill guidance, sampling and long-context behavior into separate LoRAs. Compose the capabilities you need.
Fewer sampling steps.
Guidance you can adjust.
Quality that lasts longer.
One adapter. Three strengths.
λcfg is the CFG LoRA weight. Guidance is distilled into one conditional forward per sampling step. The scales below show adapter weights, not the model’s native CFG setting.
Compose the two adapters.
See the measured downstream outputs below.
04 / DOWNSTREAM TRANSFER
Same adapters.
New worlds to control.
Move the few-step and CFG LoRAs into a downstream model in the same family. Keep sampling fast; adjust guidance independently.
No downstream retraining. These adapters are reused from the base model. All three results are shown together: only the distilled CFG LoRA weight changes, while the few-step weight stays fixed at 1.
05 / LONG CONTEXT
Keep the detail.
Long after the first frame.
For causal autoregressive models, transfer the long-context LoRA to reduce quality degradation over extended rollouts.
Selected late-rollout excerpts from the paper. Scores average seven VBench dimensions across all evaluation cases; this is not the official VBench overall score. Long-context transfer uses models that already support causal autoregressive generation.
06 / COMPARISON
Task-specific quality.
Without task-specific training.
Compare recorded outputs and benchmark results against per-task distillation. LongLive-Plug transfers directly from the backbone.
Train once.
Let your models do more.
07 / CITATION
BibTeX