Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2M to 42M parameters), dataset volume (100M+ frames from 700 hours of motion capture), and compute (21k GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through: (1) a real-time kinematic planner bridging motion tracking to tasks such as navigation, enabling natural and interactive control, and (2) a unified token space supporting VR teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body loco-manipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control.
We connect a VLA foundation model (GR00T N1.5) through the same universal control interface, combining high-level reasoning with fast, reactive whole-body control. All policies are fully autonomous.
Using video as input and GEM for pose estimation, the humanoid tracks and reproduces complex motions from human demonstrations in real-time.
A hybrid control mode using only three VR tracking points (head and hands) as upper-body humanoid motion and a kinematic planner to generate lower-body motion, enabling intuitive manipulation tasks.
Full-body VR tracking captures the operator's complete body motion, enabling precise and natural humanoid control for complex whole-body manipulation tasks.
Leveraging our universal control interface, the humanoid can perform expressive, human-like dance motions synchronized to music. The choreography is generated by GEM.
Natural language commands are translated into human motions by GEM and directly followed by the humanoid, enabling intuitive text-based control.
Our real-time kinematic planner enables interactive gamepad control with diverse locomotion styles, allowing the humanoid to navigate while maintaining distinct movement characteristics.
The kinematic planner supports diverse body configurations beyond standing locomotion, enabling low-posture movements essential for navigating constrained environments.
Athletic motions produced by the planner demonstrate the policy's ability to track and execute dynamic, coordinated movements that require precise timing and balance.
The policy demonstrates robust motion tracking under challenging conditions, maintaining stable whole-body control despite external perturbations.
GEAR-SONIC employs a universal control policy that seamlessly handles robot motion, human motion, and hybrid motion through a shared latent representation. Specialized encoders process diverse motion commands into a universal token space, enabling diverse applications including interactive gamepad control, VR teleoperation, video teleoperation, and multi-modal control from text and music.
@article{
doi:10.1126/scirobotics.aed4592,
author = {Zhengyi Luo and Ye Yuan and Tingwu Wang and Chenran Li and Fernando Castañeda and Sirui Chen and Zi-Ang Cao and Jiefeng Li and David Minor and Qingwei Ben and Jinhyung Park and David Sami and Zi Wang and Xingye Da and Runyu Ding and Cyrus Hogg and Lina Song and Edy Lim and Eugene Jeong and Tairan He and Haoru Xue and Wenli Xiao and Simon Yuen and Jan Kautz and Yan Chang and Umar Iqbal and Linxi “Jim” Fan and Yuke Zhu },
title = {SONIC: Supersizing motion tracking for natural humanoid whole-body control},
journal = {Science Robotics},
volume = {11},
number = {117},
pages = {eaed4592},
year = {2026},
doi = {10.1126/scirobotics.aed4592},
URL = {https://www.science.org/doi/abs/10.1126/scirobotics.aed4592},
eprint = {https://www.science.org/doi/pdf/10.1126/scirobotics.aed4592},
abstract = {Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gains have not been shown for humanoid control. Current neural controllers for humanoids remain modest in size, target a limited set of behaviors, and are trained on a handful of GPUs. We show that scaling model capacity, data, and compute yields a generalist humanoid controller capable of natural, robust whole-body movements. We position motion tracking as a scalable task for humanoid control, leveraging dense supervision from diverse motion-capture data to acquire human motion priors without manual reward engineering. We build a foundation model for motion tracking by scaling along three axes: network size (1.2 to 42 million parameters), dataset volume (more than 100 million frames from 700 hours of motion capture), and compute (21,000 GPU hours). Beyond demonstrating the benefits of scale, we further show downstream utility through a real-time kinematic planner that bridges motion tracking to tasks such as navigation, enabling natural and interactive control, as well as a unified token space that supports virtual reality (VR) teleoperation and vision-language-action (VLA) models with a single policy. Through this interface, we demonstrate autonomous VLA-driven whole-body locomanipulation requiring coordinated hand and foot placement. Scaling motion tracking exhibits favorable properties: Performance improves steadily with compute and data diversity, and learned policies generalize to unseen motions, establishing motion tracking at scale as a practical foundation for humanoid control. Super-scaling motion tracking yields a single, versatile behavior foundation model for humanoid whole-body control.}}