Scalable Training with SLURM#
This guide covers running ProtoMotions training jobs on SLURM-managed HPC clusters.
What is SLURM?#
SLURM (Simple Linux Utility for Resource Management) is a widely-used job scheduler for high-performance computing clusters. It manages compute resources, queues jobs, and handles multi-node distributed workloads. Most academic and enterprise GPU clusters use SLURM for job scheduling.
Overview#
ProtoMotions provides train_slurm.py, a launcher script that:
Syncs your code to the cluster (via rsync over SSH)
Generates a SLURM batch script with the correct job parameters
Submits the job to the cluster queue
Handles auto-resume via SLURM job arrays (jobs continue after timeouts)
The script is a template designed to be customized for your specific cluster setup.
Configuring for Your Cluster#
Describe your cluster in a site YAML rather than editing train_slurm.py.
The launcher looks for one at --site-config PATH, then
$PROTOMOTIONS_SLURM_SITE, then slurm_site.yaml in the repository root:
login_node: login.mycluster.edu
base_dir: /scratch/{account}/experiments # {account} is filled from --account
account: my_allocation
partition: gpu
container_mounts: /scratch:/scratch:rw
container_images:
isaacgym: /containers/isaacgym.sqsh
isaaclab: /containers/isaaclab.sqsh
newton: /containers/newton.sqsh
python_executables:
isaacgym: python
isaaclab: /workspace/isaaclab/isaaclab.sh -p
newton: python
Check what resolved before submitting anything:
python protomotions/train_slurm.py --print-site
Submitting with an unfilled placeholder is refused, naming each unset setting, before any code is synced.
Choosing Between Multiple Clusters#
If the same allocation runs on more than one cluster, list them under
clusters: and select one with --cluster NAME. Keys at the top level are
shared by every cluster – keep account there, since the allocation is the
axis that does not change when you switch cluster – and each cluster overrides
only what differs (its login node, partition, container images):
account: my_allocation # shared: same on every cluster
base_dir: /scratch/{account}/experiments
container_mounts: /scratch:/scratch:rw
default_cluster: a100 # used when --cluster is omitted
clusters:
a100:
login_node: a100-login.mycluster.edu
partition: gpu
container_images: {isaaclab: /containers/isaaclab.sqsh}
l40:
login_node: l40-login.mycluster.edu
partition: batch
container_images: {isaaclab: /containers/isaaclab.sqsh}
Then --cluster l40 submits the same job to the L40 cluster with the same
account. --print-site --cluster l40 shows exactly what that resolves to.
Container Setup#
You’ll need containerized environments with ProtoMotions dependencies. Convert your
Docker images to Singularity (.sif) or Enroot (.sqsh) format as required
by your cluster.
Auto-Resume with Job Arrays#
Long training runs often exceed cluster time limits (e.g., 4-hour walltime). ProtoMotions handles this automatically using two mechanisms:
1. SLURM Job Arrays
The launcher submits jobs as arrays (--array=0-5%1), meaning up to 5 sequential
jobs will run. When a job times out, the next array task starts and resumes from
the last checkpoint.
2. AutoResume Callback
When --use-slurm is enabled, training registers the AutoResumeCallbackSrun
callback. This callback:
Tracks elapsed training time
Saves a checkpoint before the SLURM time limit (default: after 3.5 hours)
Gracefully stops training so the next array job can resume
# From protomotions/agents/callbacks/slurm_autoresume_srun.py
class AutoResumeCallbackSrun(Callback):
def __init__(self, autoresume_after=12600): # 3.5 hours in seconds
self.autoresume_after = autoresume_after
def _check_autoresume(self, agent):
if time.time() - self.start_time >= self.autoresume_after:
agent.save() # Save checkpoint
agent._should_stop = True # Signal graceful stop
The default autoresume_after=12600 (3.5 hours) works well with 4-hour job limits,
providing buffer time for checkpoint saving.
Understanding Scaling Parameters#
The --num-envs and --batch-size parameters are specified per GPU. With
multi-GPU and multi-node training, the effective totals scale accordingly:
Total GPUs = ngpu × nodes
Effective num-envs = num-envs × Total GPUs
Effective batch-size = batch-size × Total GPUs
Example:
With --ngpu=4 --nodes=2 --num-envs=4096 --batch-size=16384:
Total GPUs: 4 × 2 = 8 GPUs
Effective environments: 4,096 × 8 = 32,768 parallel environments
Effective batch size: 16,384 × 8 = 131,072 samples per update
This scaling is automatic—you specify per-GPU values and the distributed training handles aggregation across all processes.
Running a Training Job#
Once configured, launch training from your local machine:
python protomotions/train_slurm.py \
--robot-name=g1 \
--simulator=isaaclab \
--num-envs=4096 \
--batch-size=16384 \
--motion-file=/cluster/path/to/motions.pt \
--experiment-path=examples/experiments/mimic/mlp_bm_l2c2.py \
--experiment-name=g1_motion_tracker \
--user=myusername \
--ngpu=4 \
--nodes=1 \
--slurm-time=4:00:00 \
--use-wandb
Key arguments:
Argument |
Description |
|---|---|
|
Robot to train (e.g., |
|
Physics backend ( |
|
Parallel environments (scale with GPU memory) |
|
PPO batch size (typically 2-4x num-envs) |
|
Path to motion data on the cluster |
|
Experiment config file (relative to repo root) |
|
Unique name for this experiment |
|
Your cluster username |
|
Which cluster in the site file’s |
|
GPUs per node |
|
Number of compute nodes |
|
Job time limit (HH:MM:SS) |
|
Number of auto-resume attempts (default: 5) |
|
Maximum complete rollout and optimization iterations |
|
Enable Weights & Biases logging |
|
Weights & Biases project name (default: |
Multi-Node Training#
For large-scale training across multiple nodes:
python protomotions/train_slurm.py \
--robot-name=smpl \
--simulator=isaacgym \
--num-envs=8192 \
--batch-size=16384 \
--motion-file=/cluster/path/to/amass_train.pt \
--experiment-path=examples/experiments/mimic/mlp.py \
--experiment-name=smpl_motion_tracker_4node \
--user=myusername \
--ngpu=8 \
--nodes=4 \
--slurm-time=4:00:00 \
--use-wandb
ProtoMotions uses PyTorch Fabric for distributed training. Each node runs
--ngpu processes, and gradients are synchronized across all nodes.
Monitoring Jobs#
After submission, the script prints monitoring commands:
# Monitor live output
ssh myusername@cluster 'tail -f /path/to/exp/slurm_output.log'
# Check job status
ssh myusername@cluster 'squeue -u myusername'
# Cancel a job
ssh myusername@cluster 'scancel <job_id>'
Next Steps#
Configuration System - Configuration system details
Experiments - Creating custom experiments
Domain Randomization & Sim2Sim - Domain randomization for robust policies