MiniMax H3 Super Acceleration fast draft generation and high-resolution refinement, powered by Sol Engine 6.85 s for a 5-second 768p video · 14.93 s for a 10-second video

H3 Super Acceleration first uses H3 with a LoRA to generate a four-step draft at 896×512. It then upsamples the draft and performs three LTX refinement steps at the target resolution with Sol-Attn. Combining the measured stages on one NVIDIA GB200 gives 22.2× speedup for a 5-second 1344×768 video and 27.7× speedup for a 10-second video over the published SGLang baseline.

Production at a glance

Up to 378K videos every 30 days

One NVIDIA GB200 running H3 Super Acceleration continuously can turn the measured end-to-end latency into production-scale output.

768p · 5-second videos6.852 s measured E2E
12.6Kvideos / day17.5 hours of finished video
378Kvideos / 30 days525 hours of finished video
768p · 10-second videos14.931 s measured E2E
5.79Kvideos / day16.1 hours of finished video
174Kvideos / 30 days482 hours of finished video

Ideal full-utilization estimate. Counts assume one GB200, batch size 1, 24 × 7 serial inference, no idle time or failed jobs, and a 30-day month. Finished-video hours express media volume; encoded storage in GB or TB depends on the delivery codec and bitrate.

01 — Motivation

A practical speed–quality tradeoff for production video inference.

Our target is to reduce end-to-end inference time while keeping the visual and audio differences small enough for practical use. Instead of running many H3 steps at full resolution, H3 Super Acceleration performs most generation work in a short low-resolution draft and uses a lightweight high-resolution refinement pass.

This is not lossless acceleration. H3 Super Acceleration changes the sampling path, so it is not bit-exact with the SGLang baseline. Differences can appear in detail, texture, motion, or audio. The paired videos below let readers assess that tradeoff directly.

02 — Two-Stage Generation Pipeline

A low-resolution H3 draft followed by a short LTX refinement pass.

Each stage has one job. H3 creates the initial video at 896×512 in four denoising steps. LTX then upsamples and refines that draft at the requested output resolution in three steps with Sol-Attn. The two stages run serially on one GB200.

Hardware
1× GB200
Single-GPU inference
Stage 1 · 896×512
4 steps
H3 + LoRA · 24 FPS
Stage 2 · 768p / 1080p / 2K
3 steps
LTX · Sol-Attn · 24 FPS
Total
7 steps
Draft generation + refinement
Stage 1 · 896×512 · 24 FPS
Fast draft generation
4 steps
Prompt and conditioningGeneration inputs enter the H3 pipeline.
H3 + LoRAFour-step denoising at 896×512 uses the task-specific LoRA to generate the draft efficiently.
Draft videoComposition and motion are ready for refinement.
Stage 2 · 768p / 1080p / 2K · 24 FPS
High-resolution refinement
3 steps
Spatial upsamplingThe draft video is resized to the target resolution.
LTX refinementA three-step Sol-Attn pass restores high-resolution detail and consistency.
Final high-resolution videoThe refined result is decoded and returned as the final output.

H3 generates a 24 FPS draft at 896×512; after upsampling, LTX refines the same video at 24 FPS and 768p, 1080p, or 2K output resolution.

StageModelResolutionFPSDenoising stepsOutput
Stage 1MiniMax-H3 + LoRA896×512244Draft video
Stage 2LTX · Sol-Attn768p / 1080p / 2K243Final refined video

03 — Measured Latency

End-to-end 768p latency on one NVIDIA GB200.

The latest Sol-Super results use Sol-Attn for the complete Stage 2 service. Adding the measured Stage 1 and Stage 2 blocks gives 6.852 seconds for a 5-second 1344×768 video and 14.931 seconds for a 10-second video, corresponding to 22.2× and 27.7× speedup over the published SGLang baseline.

ResolutionDurationDiffusersSGLangSol EngineSol-Super
E2E
Sol-Super
vs. SGLang
1344×7685 s167.8 s152.3 s45.4 s6.852 s22.2×
1344×76810 s468.2 s414.1 s132.5 s14.931 s27.7×

Benchmark basis. Diffusers, SGLang, and Sol Engine are author-supplied warm end-to-end measurements. For the 5-second and 10-second settings, Sol-Super combines separately measured Stage 1 blocks of 4.1728 and 9.8204 seconds, respectively, with the latest Stage 2 means from 10 hot complete-service requests on one GB200; model loading and warmup are excluded. Sol-Super speedup is SGLang E2E ÷ Sol-Super E2E.


04 — Visual Comparison

Paired visual comparisons plus standalone 1440p examples.

Each 768p and 1080p pair uses matching generation inputs to compare the SGLang baseline with H3 Super Acceleration. No matched 1440p baseline was recorded, so the last panel contains two standalone H3 Super Acceleration results and makes no speed claim.

Resolution 1344×768 · Duration 5 s

Scene — A man and woman exchange a tense remark on a crowded Japanese street.

SGLang BaselineReference result
H3 Super AccelerationAccelerated result

Resolution 1344×768 · Duration 10 s

Scene — A man speaks to a handheld camera on a quiet suburban street.

SGLang BaselineReference result
H3 Super AccelerationAccelerated result

Resolution 1080p · Duration 5 s

Scene — Two martial artists face one another in a bamboo forest.

SGLang BaselineReference result
H3 Super AccelerationAccelerated result

Resolution 1080p · Duration 10 s

Scene — A woman speaks to a handheld camera after a gym workout.

SGLang BaselineReference result
H3 Super AccelerationAccelerated result

Resolution 2560×1440 · Duration 10 s

Two standalone H3 Super Acceleration results; no matched baseline or speed claim is shown.

City Crowd ReactionH3 Super Acceleration
Futuristic WarriorH3 Super Acceleration

05 — Tokenomics

Throughput, revenue, and compute-adjusted margin for one fully utilized NVIDIA GB200.

MiniMax's official API pricing lists 768P H3 output at $0.08 per second: $0.40 for a 5-second video and $0.80 for a 10-second video. Under the clearly stated assumptions below, the combined Sol-Super latency corresponds to a 97.1%–97.4% GPU-only gross margin.

Setting / priceSGLang
videos / hour
Sol-Super
videos / hour
SGLang
revenue / hour
Sol-Super
revenue / hour
Speedup
vs. SGLang
Sol-Super
gross margin
768p · 5 s$0.40 / video23.64525.41$9.46$210.1622.2×97.4%
768p · 10 s$0.80 / video8.69241.10$6.95$192.8827.7×97.1%
Ideal full-utilization scenario. We assume one GB200 runs continuously at 100% utilization and costs $5.50 per GPU-hour under a one-year Long-Term Agreement. Videos per hour are 3,600 ÷ measured E2E latency; hourly revenue is throughput × price per video; gross margin is (revenue − $5.50) ÷ revenue. This is a GPU-only margin: it excludes idle time, failed jobs, storage, networking, staffing, software, and other operating costs, so it is not net profit.