Scroll to the paper

World-action models · real-time robot control

Long-WAM Scaling the Context of World-Action Models

Wei Huang*Bohan Zhang*Chenzhi LiuIsabella LiuShuai YangWeian MaoLuozhou WangYicheng XiaoWeifeng LinQixin HuBryan ChuSifei LiuJim (Linxi) FanXiaojuan QiSong HanYukang Chen

NVIDIAMITHKUUCSD

* Equal contribution

TL;DRLong-WAM scales the visual history of a causal world-action model and keeps it real-time. Autoregressive video pretraining turns longer context into better control, and a co-designed runtime runs the full predict-then-act model at 107.4 ms per action chunk on an RTX 5090, with deployment on DGX Spark and Jetson AGX Thor.

78.7%RoboCasa GR-1with 19.2 s of context
99.5%LIBERO-Longsimulation benchmark
94.4%RoboTwin 2.0average, clean + randomized
107.4msper action chunkRTX 5090 · incl. future prediction
95%dynamic cup stackingUnitree G1 · 19 / 20
54.4%RoboCasa365with GPT-6 Astra · 50 tasks

Abstract

Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model–system framework for scaling the context of causal world-action models under real-time control constraints. Access to history does not by itself ensure effective context use: in our comparisons, longer histories amplify the advantage of AR video pretraining. We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 2.4 to 19.2 seconds raises success from 66.3% to 78.7%, while bidirectional initialization shows no net gain; robot-domain AR pretraining further improves peak success on GR-1 and LIBERO-Long. Long-WAM also achieves leading results on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-customized acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor while retaining future prediction. Inference takes 107.4 ms per action chunk on RTX 5090, including future-video latent prediction. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.

01 · Method

Remember the past. Imagine the future. Act.

Remember
Observed video historythe context we scale
Video expertinitialized from LongLive2.0-Robot
Imagine
Future latentspartially denoised · no pixel decoding
Action expertreads history + predicted future
ActAction chunkexecuted asynchronously

Inverse-dynamics (IDM) inference: the video expert predicts future latents from the observed history; the action expert denoises an action chunk conditioned on that history and the predicted future. Video queries never read action tokens.

1

Autoregressive video pretraining

LongLive2.0-Robot continues LongLive-2.0’s AR checkpoint on ≈10,000 window-equivalent hours of robot and egocentric video (RoVid-X, AgiBot World, EgoDex, EgoVerse, VITRA). It learns teacher-forced next-chunk prediction on sequences up to 30 s, without action labels.

2

Causal-to-causal world-action adaptation

Video and action experts are coupled through an asymmetric attention interface (a Mixture of Transformers). Video queries never read action tokens. Action queries read observed history, partially denoised future latents and the action chunk, so Long-WAM predicts the future first and then acts, without decoding pixels.

3

Context scaling

We scale the duration of real observations in the causal prefix and keep the forecast and action horizons fixed. Each window is trained separately and evaluated at its own length, up to 38.4 s on RoboCasa GR-1.

4

Real-time infrastructure

Pure asynchronous execution, streaming VAE encoding, NVFP4 quantization, KV reuse, CUDA Graph and device-specific kernel tuning bring the full model to RTX 5090, DGX Spark and Jetson AGX Thor.

02 · Context scaling

The value of context depends on AR pretraining.

Both autoregressive initializations turn longer history into higher success. The bidirectional initialization shows no net gain on RoboCasa GR-1 between 2.4 s and 19.2 s. On LIBERO-Long, 2.4 s of history lifts success from 94.5% to 99.5%.

Video expert initialized fromLongLive2.0-Robot (AR)LongLive2.0 (AR)Wan2.2 (bidirectional)

RoboCasa GR-1

63.3% → 78.7%

LIBERO-Long

94.5% → 99.5%

Context windows are equally spaced on the x-axis. At 38.4 s, 80.4% of sampled history frames are padding (training trajectories average 12.1 s). The drop there reflects limited effective history, not an established memory limit.

Context explorer

drag to change the observed history
now
0 s2.4 s4.8 s9.6 s19.2 s38.4 s
History19.2 s
RTX 5090 latency341.0 msper action chunk

RoboCasa GR-1 comparison

success rate (%)

Latency vs. context

RTX 5090 · per action chunk

03 · Simulation benchmarks

Leading results across tasks and environments.

Long-WAM (IDM) reaches 99.5% average on LIBERO and 94.4% on RoboTwin 2.0. After dynamic-data fine-tuning it reaches 34.9% success on DOMINO, the highest among the compared methods.

04 · Real world

Catch what moves. Stay on task.

Unitree G1 grasps from a moving conveyor at up to 7.5 cm/s; both baselines fail every trial at 6.0 and 7.5 cm/s. Dynamic cup stacking succeeds in 19 of 20 trials versus 0 of 20 for both baselines. On YAM, tasks lasting over 40 s average 81.7% success.

π0.5Fast-WAMLong-WAMHuman (teleoperation)

Pick up the moving cup

success · 20 trials per policy and speed

Dynamic cup stacking

3 cm/s

Long-horizon manipulation on YAM

mean 81.7%

05 · Real time on the edge

107.4 ms per action chunk, with future prediction retained.

Shared optimizations and device-specific kernel tuning speed up the full video-then-action pipeline by 3.3×, 4.1× and 3.2× over BF16 eager on RTX 5090, DGX Spark and Jetson AGX Thor.

Acceleration explorer

step through the optimization stages
BaselineBF16 eager

Long-WAM (IDM) · V4/A4 · full observation VAE included · each stage keeps the previous optimizations.

End-to-end latency on RTX 5090

ms per action chunk

Pure asynchronous execution

RoboTwin 2.0

Long-WAM keeps 94.2% success asynchronously (94.4% synchronously) with the lowest overlap RMSE and jerk. The AR video foundation conditions actions on temporally continuous predictions.

Cumulative acceleration

latency (ms) and speedup vs. BF16 eager

06 · GPT-6 Astra + Long-WAM

Reasoning needs a capable executor.

Coupled with GPT-6 Astra, the unchanged Long-WAM checkpoint rises from 31.4% to 54.4% overall on RoboCasa365. On unseen compositions it rises from 6.1% to 35.0%.

GPT-6 Astraagentic reasoning over the whole task
Plandecomposes the goal into sub-instructions
Recoverhandles failures and recovery
Correctadjusts actions when needed
Long-WAMexecutes memory-informed action chunks

Success by split

RoboCasa365 · 50 tasks
GPT-6 Astraπ0.5 + GPT-6 AstraLong-WAMLong-WAM + GPT-6 Astra

RoboCasa365 · all methods

success rate (%)

07 · Demos

See it in action.

Every clip is labelled with its source and playback speed.

Imagination · LongLive2.0-Robot predictions

Each clip is generated from a single image and a language instruction. They are video predictions, not executed rollouts.

Place the brown shoe into the blue containergenerated · 1×
Put the bun into the steamergenerated · 1×
Cover the frying pan with a glass lidgenerated · 1×
Pour water into a bowlgenerated · 1×
Fold the clothinggenerated · 2×
Empty the water from the measuring cupgenerated · 1×

Dynamic manipulation · Unitree G1

Dynamic cup stacking at 3 cm/s: intercept, align, stack · 19/20 successes in evaluationreal robot · 1×
Picking up the cup at 3.0 / 4.5 / 6.0 / 7.5 cm/s · 100 / 100 / 95 / 90% over 20 trials eachreal robot · 1×
π0.5Fast-WAM
Same stacking task: both baselines miss the moving green cup · 0/20 each in evaluationreal robot · 1×
Synchronous execution vs. our real-time infrastructure on the same dynamic graspreal robot · 1×
Synchronous (top) vs. asynchronous (bottom) execution at four conveyor speeds: the robot keeps moving while the next chunk is computedreal robot · 1×

Long-horizon manipulation · YAM

Stack bowls · 85%Sort bricks by color · 80%
Selected rollouts from tasks that average over 40 s · three-task mean 81.7% (20 trials per task)real robot · 2×

GPT-6 Astra + Long-WAM · RoboCasa365

Left: Long-WAM alone. Right: GPT-6 Astra plans, handles recovery and corrects actions when needed, while Long-WAM executes. Both sides use the same checkpoint. The SUCCESS and FAILED badges come from the original recordings.

Long-WAM aloneGPT-6 Astra + Long-WAM
Pick the kettle and place it on the tray, then pick the mug from the cabinet and place it on the tray, then close the cabinet doorssimulation · 4×
Long-WAM aloneGPT-6 Astra + Long-WAM
Place the appropriate cutting tool for cutting the cucumber skin on the cutting boardsimulation · 4×

The full story · narrated overview (4:39)

Method, results and demos with narration and burned-in captionsvoice: synthetic (Kokoro)

Citation

BibTeX

@article{huang2026longwam,
  title   = {Long-WAM: Scaling the Context of World-Action Models},
  author  = {Huang, Wei and Zhang, Bohan and Liu, Chenzhi and Liu, Isabella and
             Yang, Shuai and Mao, Weian and Wang, Luozhou and Xiao, Yicheng and
             Lin, Weifeng and Hu, Qixin and Chu, Bryan and Liu, Sifei and
             Fan, Linxi and Qi, Xiaojuan and Han, Song and Chen, Yukang},
  journal = {arXiv preprint},
  year    = {2026}
}