Scalable Training with SLURM#

This guide covers running ProtoMotions training jobs on SLURM-managed HPC clusters.

What is SLURM?#

SLURM (Simple Linux Utility for Resource Management) is a widely-used job scheduler for high-performance computing clusters. It manages compute resources, queues jobs, and handles multi-node distributed workloads. Most academic and enterprise GPU clusters use SLURM for job scheduling.

Overview#

ProtoMotions provides train_slurm.py, a launcher script that:

  1. Syncs your code to the cluster (via rsync over SSH)

  2. Generates a SLURM batch script with the correct job parameters

  3. Submits the job to the cluster queue

  4. Handles auto-resume via SLURM job arrays (jobs continue after timeouts)

The script is a template designed to be customized for your specific cluster setup.

Configuring for Your Cluster#

Describe your cluster in a site YAML rather than editing train_slurm.py. The launcher looks for one at --site-config PATH, then $PROTOMOTIONS_SLURM_SITE, then slurm_site.yaml in the repository root:

login_node: login.mycluster.edu
base_dir: /scratch/{account}/experiments   # {account} is filled from --account
account: my_allocation
partition: gpu
container_mounts: /scratch:/scratch:rw
container_images:
  isaacgym: /containers/isaacgym.sqsh
  isaaclab: /containers/isaaclab.sqsh
  newton: /containers/newton.sqsh
python_executables:
  isaacgym: python
  isaaclab: /workspace/isaaclab/isaaclab.sh -p
  newton: python

Check what resolved before submitting anything:

python protomotions/train_slurm.py --print-site

Submitting with an unfilled placeholder is refused, naming each unset setting, before any code is synced.

Choosing Between Multiple Clusters#

If the same allocation runs on more than one cluster, list them under clusters: and select one with --cluster NAME. Keys at the top level are shared by every cluster – keep account there, since the allocation is the axis that does not change when you switch cluster – and each cluster overrides only what differs (its login node, partition, container images):

account: my_allocation          # shared: same on every cluster
base_dir: /scratch/{account}/experiments
container_mounts: /scratch:/scratch:rw
default_cluster: a100           # used when --cluster is omitted
clusters:
  a100:
    login_node: a100-login.mycluster.edu
    partition: gpu
    container_images: {isaaclab: /containers/isaaclab.sqsh}
  l40:
    login_node: l40-login.mycluster.edu
    partition: batch
    container_images: {isaaclab: /containers/isaaclab.sqsh}

Then --cluster l40 submits the same job to the L40 cluster with the same account. --print-site --cluster l40 shows exactly what that resolves to.

Container Setup#

You’ll need containerized environments with ProtoMotions dependencies. Convert your Docker images to Singularity (.sif) or Enroot (.sqsh) format as required by your cluster.

Auto-Resume with Job Arrays#

Long training runs often exceed cluster time limits (e.g., 4-hour walltime). ProtoMotions handles this automatically using two mechanisms:

1. SLURM Job Arrays

The launcher submits jobs as arrays (--array=0-5%1), meaning up to 5 sequential jobs will run. When a job times out, the next array task starts and resumes from the last checkpoint.

2. AutoResume Callback

When --use-slurm is enabled, training registers the AutoResumeCallbackSrun callback. This callback:

  • Tracks elapsed training time

  • Saves a checkpoint before the SLURM time limit (default: after 3.5 hours)

  • Gracefully stops training so the next array job can resume

# From protomotions/agents/callbacks/slurm_autoresume_srun.py
class AutoResumeCallbackSrun(Callback):
    def __init__(self, autoresume_after=12600):  # 3.5 hours in seconds
        self.autoresume_after = autoresume_after

    def _check_autoresume(self, agent):
        if time.time() - self.start_time >= self.autoresume_after:
            agent.save()           # Save checkpoint
            agent._should_stop = True  # Signal graceful stop

The default autoresume_after=12600 (3.5 hours) works well with 4-hour job limits, providing buffer time for checkpoint saving.

Understanding Scaling Parameters#

The --num-envs and --batch-size parameters are specified per GPU. With multi-GPU and multi-node training, the effective totals scale accordingly:

Total GPUs = ngpu × nodes
Effective num-envs = num-envs × Total GPUs
Effective batch-size = batch-size × Total GPUs

Example:

With --ngpu=4 --nodes=2 --num-envs=4096 --batch-size=16384:

  • Total GPUs: 4 × 2 = 8 GPUs

  • Effective environments: 4,096 × 8 = 32,768 parallel environments

  • Effective batch size: 16,384 × 8 = 131,072 samples per update

This scaling is automatic—you specify per-GPU values and the distributed training handles aggregation across all processes.

Running a Training Job#

Once configured, launch training from your local machine:

python protomotions/train_slurm.py \
    --robot-name=g1 \
    --simulator=isaaclab \
    --num-envs=4096 \
    --batch-size=16384 \
    --motion-file=/cluster/path/to/motions.pt \
    --experiment-path=examples/experiments/mimic/mlp_bm_l2c2.py \
    --experiment-name=g1_motion_tracker \
    --user=myusername \
    --ngpu=4 \
    --nodes=1 \
    --slurm-time=4:00:00 \
    --use-wandb

Key arguments:

Argument

Description

--robot-name

Robot to train (e.g., g1, smpl, h1_2)

--simulator

Physics backend (isaacgym, isaaclab, newton)

--num-envs

Parallel environments (scale with GPU memory)

--batch-size

PPO batch size (typically 2-4x num-envs)

--motion-file

Path to motion data on the cluster

--experiment-path

Experiment config file (relative to repo root)

--experiment-name

Unique name for this experiment

--user

Your cluster username

--cluster

Which cluster in the site file’s clusters: to submit to (default: default_cluster)

--ngpu

GPUs per node

--nodes

Number of compute nodes

--slurm-time

Job time limit (HH:MM:SS)

--array-size

Number of auto-resume attempts (default: 5)

--training-max-iterations

Maximum complete rollout and optimization iterations

--use-wandb

Enable Weights & Biases logging

--wandb-project

Weights & Biases project name (default: physical_animation)

Multi-Node Training#

For large-scale training across multiple nodes:

python protomotions/train_slurm.py \
    --robot-name=smpl \
    --simulator=isaacgym \
    --num-envs=8192 \
    --batch-size=16384 \
    --motion-file=/cluster/path/to/amass_train.pt \
    --experiment-path=examples/experiments/mimic/mlp.py \
    --experiment-name=smpl_motion_tracker_4node \
    --user=myusername \
    --ngpu=8 \
    --nodes=4 \
    --slurm-time=4:00:00 \
    --use-wandb

ProtoMotions uses PyTorch Fabric for distributed training. Each node runs --ngpu processes, and gradients are synchronized across all nodes.

Monitoring Jobs#

After submission, the script prints monitoring commands:

# Monitor live output
ssh myusername@cluster 'tail -f /path/to/exp/slurm_output.log'

# Check job status
ssh myusername@cluster 'squeue -u myusername'

# Cancel a job
ssh myusername@cluster 'scancel <job_id>'

Next Steps#