SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Co-First Authors     * Core Contributors     Equal Advising
Science Robotics, Vol. 11, Issue 117 (2026)
All models shown in the videos will be released.

Trailers

One Day at NVIDIA

Skill Montage

Abstract

Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through: (1) a real-time kinematic planner bridging motion tracking to tasks such as navigation, enabling natural and interactive control, and (2) a unified token space supporting VR teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control.

All results are generated using a single unified control policy.

Connection to VLA Foundation Model

We connect a VLA foundation model (GR00T N1.5) through the same universal control interface, combining high-level reasoning with fast, reactive whole-body control. All policies are fully autonomous.

Drill

Dropping Soda into Trash Can


Carrot

Sponge


Trash Can

Apple


Teleoperation

Video Teleoperation

Using video as input and GEM for pose estimation, the humanoid tracks and reproduces complex motions from human demonstrations in real-time.

Kung Fu

Crawling



VR Teleoperation with Keypoints

A hybrid control mode using only three VR tracking points (head and hands) as upper-body humanoid motion and a kinematic planner to generate lower-body motion, enabling intuitive manipulation tasks.

Lawn Mowing

Object Manipulation



VR Teleoperation with Whole Body

Full-body VR tracking captures the operator's complete body motion, enabling precise and natural humanoid control for complex whole-body manipulation tasks.

Whole-Body Tracking

Manipulation with Whole Body


Multi-Modal Control

Music Control

Leveraging our universal control interface, the humanoid can perform expressive, human-like dance motions synchronized to music. The choreography is generated by GEM.



Text Control

Natural language commands are translated into human motions by GEM and directly followed by the humanoid, enabling intuitive text-based control.

Backward Walking

Monkey Movement


Interactive Kinematic Planner

Stylized Locomotion

Our real-time kinematic planner enables interactive gamepad control with diverse locomotion styles, allowing the humanoid to navigate while maintaining distinct movement characteristics.

Happy Walking

Running


Stealth Walking

Injured Walking



Squatting, Kneeling, and Crawling

The kinematic planner supports diverse body configurations beyond standing locomotion, enabling low-posture movements essential for navigating constrained environments.

Squatting

Kneeling


Hand Crawling

Elbow Crawling



Boxing

Athletic motions produced by the planner demonstrate the policy's ability to track and execute dynamic, coordinated movements that require precise timing and balance.

Boxing

Boxing with Movement


Tracking Robustness

The policy demonstrates robust motion tracking under challenging conditions, maintaining stable whole-body control despite external perturbations.


Method

GEAR-SONIC employs a universal control policy that seamlessly handles robot motion, human motion, and hybrid motion through a shared latent representation. Specialized encoders process diverse motion commands into a universal token space, enabling diverse applications including interactive gamepad control, VR teleoperation, video teleoperation, and multi-modal control from text and music.

Method Overview

BibTeX


@article{
doi:10.1126/scirobotics.aed4592,
author = {Zhengyi Luo  and Ye Yuan  and Tingwu Wang  and Chenran Li  and Fernando Castañeda  and Sirui Chen  and Zi-Ang Cao  and Jiefeng Li  and David Minor  and Qingwei Ben  and Jinhyung Park  and David Sami  and Zi Wang  and Xingye Da  and Runyu Ding  and Cyrus Hogg  and Lina Song  and Edy Lim  and Eugene Jeong  and Tairan He  and Haoru Xue  and Wenli Xiao  and Simon Yuen  and Jan Kautz  and Yan Chang  and Umar Iqbal  and Linxi “Jim” Fan  and Yuke Zhu },
title = {SONIC: Supersizing motion tracking for natural humanoid whole-body control},
journal = {Science Robotics},
volume = {11},
number = {117},
pages = {eaed4592},
year = {2026},
doi = {10.1126/scirobotics.aed4592},
URL = {https://www.science.org/doi/abs/10.1126/scirobotics.aed4592},
eprint = {https://www.science.org/doi/pdf/10.1126/scirobotics.aed4592},
abstract = {Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2 to 42 million parameters), dataset volume (more than 100 million frames from 700 hours of motion capture), and compute (21,000 GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through a real-time kinematic planner that bridges motion tracking to tasks such as navigation, enabling natural and interactive control, as well as a unified token space that supports virtual reality (VR) teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body locomanipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: Performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control. Super-scaling motion tracking yields a single, versatile behavior foundation model for humanoid whole-body control.}}