MiniMax H3 Super Acceleration fast draft generation and high-resolution refinement, powered by Sol Engine 6.85 s for a 5-second 768p video · 14.93 s for a 10-second video
H3 Super Acceleration first uses H3 with a LoRA to generate a four-step draft at 896×512. It then upsamples the draft and performs three LTX refinement steps at the target resolution with Sol-Attn. Combining the measured stages on one NVIDIA GB200 gives 22.2× speedup for a 5-second 1344×768 video and 27.7× speedup for a 10-second video over the published SGLang baseline.
Up to 378K videos every 30 days
One NVIDIA GB200 running H3 Super Acceleration continuously can turn the measured end-to-end latency into production-scale output.
Ideal full-utilization estimate. Counts assume one GB200, batch size 1, 24 × 7 serial inference, no idle time or failed jobs, and a 30-day month. Finished-video hours express media volume; encoded storage in GB or TB depends on the delivery codec and bitrate.
A practical speed–quality tradeoff for production video inference.
Our target is to reduce end-to-end inference time while keeping the visual and audio differences small enough for practical use. Instead of running many H3 steps at full resolution, H3 Super Acceleration performs most generation work in a short low-resolution draft and uses a lightweight high-resolution refinement pass.
A low-resolution H3 draft followed by a short LTX refinement pass.
Each stage has one job. H3 creates the initial video at 896×512 in four denoising steps. LTX then upsamples and refines that draft at the requested output resolution in three steps with Sol-Attn. The two stages run serially on one GB200.
H3 generates a 24 FPS draft at 896×512; after upsampling, LTX refines the same video at 24 FPS and 768p, 1080p, or 2K output resolution.
| Stage | Model | Resolution | FPS | Denoising steps | Output |
|---|---|---|---|---|---|
| Stage 1 | MiniMax-H3 + LoRA | 896×512 | 24 | 4 | Draft video |
| Stage 2 | LTX · Sol-Attn | 768p / 1080p / 2K | 24 | 3 | Final refined video |
End-to-end 768p latency on one NVIDIA GB200.
The latest Sol-Super results use Sol-Attn for the complete Stage 2 service. Adding the measured Stage 1 and Stage 2 blocks gives 6.852 seconds for a 5-second 1344×768 video and 14.931 seconds for a 10-second video, corresponding to 22.2× and 27.7× speedup over the published SGLang baseline.
| Resolution | Duration | Diffusers | SGLang | Sol Engine | Sol-Super E2E | Sol-Super vs. SGLang |
|---|---|---|---|---|---|---|
| 1344×768 | 5 s | 167.8 s | 152.3 s | 45.4 s | 6.852 s | 22.2× |
| 1344×768 | 10 s | 468.2 s | 414.1 s | 132.5 s | 14.931 s | 27.7× |
Benchmark basis. Diffusers, SGLang, and Sol Engine are author-supplied warm end-to-end measurements. For the 5-second and 10-second settings, Sol-Super combines separately measured Stage 1 blocks of 4.1728 and 9.8204 seconds, respectively, with the latest Stage 2 means from 10 hot complete-service requests on one GB200; model loading and warmup are excluded. Sol-Super speedup is SGLang E2E ÷ Sol-Super E2E.
The SGLang baseline fills the track in each setting. Sol-Super bars are enlarged only to keep the Stage 1 and Stage 2 labels readable; their internal proportions, printed latency values, and speedups are the measured results. Each red arrow ends at the right edge of its corresponding SGLang bar.
All segments use one linear scale. Labels are placed inside the bar wherever they fit; the compact keys below each bar list only the narrow segments. The red arrow spans from the current latency to the baseline endpoint. Totals exclude the shared mux and cleanup stage. In our baseline, Stage 1 combines 3.772 seconds of H3 DiT and shared work with 3.427 seconds of official H3 VAE decode, for 7.199 seconds total. Sol-Super Stage 1 combines the same 3.772 seconds of H3 work with 0.028 seconds of TAEH3 decode and 0.3728 seconds of text/image encoding, for 4.1728 seconds total. The decoder values are averages from two matched 1344×768 cases. Refiner blocks report complete Refiner latency; the Sol-Attn DiT portion takes 0.985 seconds. The original LTX-2.5 Video VAE decoder takes 6.404 seconds in the matched 121-frame decode benchmark.
Paired visual comparisons plus standalone 1440p examples.
Each 768p and 1080p pair uses matching generation inputs to compare the SGLang baseline with H3 Super Acceleration. No matched 1440p baseline was recorded, so the last panel contains two standalone H3 Super Acceleration results and makes no speed claim.
Resolution 1344×768 · Duration 5 s
Scene — A man and woman exchange a tense remark on a crowded Japanese street.
Resolution 1344×768 · Duration 10 s
Scene — A man speaks to a handheld camera on a quiet suburban street.
Resolution 1080p · Duration 5 s
Scene — Two martial artists face one another in a bamboo forest.
Resolution 1080p · Duration 10 s
Scene — A woman speaks to a handheld camera after a gym workout.
Resolution 2560×1440 · Duration 10 s
Two standalone H3 Super Acceleration results; no matched baseline or speed claim is shown.
Throughput, revenue, and compute-adjusted margin for one fully utilized NVIDIA GB200.
MiniMax's official API pricing lists 768P H3 output at $0.08 per second: $0.40 for a 5-second video and $0.80 for a 10-second video. Under the clearly stated assumptions below, the combined Sol-Super latency corresponds to a 97.1%–97.4% GPU-only gross margin.
| Setting / price | SGLang videos / hour | Sol-Super videos / hour | SGLang revenue / hour | Sol-Super revenue / hour | Speedup vs. SGLang | Sol-Super gross margin |
|---|---|---|---|---|---|---|
| 768p · 5 s$0.40 / video | 23.64 | 525.41 | $9.46 | $210.16 | 22.2× | 97.4% |
| 768p · 10 s$0.80 / video | 8.69 | 241.10 | $6.95 | $192.88 | 27.7× | 97.1% |