How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

2026-09-23 · Hugging Face

Accelerating Robotics Simulation with NVIDIA Warp and MJWarp

As robotics learning workloads grow, the critical question shifts from how fast a single simulation world can run to how many worlds can run simultaneously. GPU acceleration enables advancing these worlds in large batches while keeping simulation and learning data close to the device. This article, the second in the "State of Simulation for Physical AI" series, focuses on preparing and scaling simulation environments rather than training policies.

Technology Stack and Decision Guide

The integration of these technologies forms a robust stack for robotics simulation:

  • NVIDIA Warp: A Python kernel language supporting Single Instruction, Multiple Threads (SIMT), automatic differentiation, and PyTorch/JAX interoperability.
  • MJWarp: Implements MuJoCo physics on Warp, maintaining the same MJCF model format while delivering batched GPU throughput.
  • Your Scene (e.g., SO-101): Utilizes familiar Menagerie or Robot Studio assets combined with task geometry.
  • Next Integrations (Newton / Isaac Lab): Provides multi-solver APIs, USD support, sensors, managers, and training loops.

Developers can choose their tools based on specific needs:

  • For single-robot MPC or teleoperation, use MuJoCo CPU.
  • For maximum throughput on raw MuJoCo physics, use MJWarp.
  • For JAX training recipes, use MuJoCo Playground or MJX.
  • For multi-solver and Isaac Lab integration, look forward to Newton.

Core Advantages of NVIDIA Warp

NVIDIA Warp is a Python framework for authoring high-performance, GPU-accelerated kernels. Developers write statically typed kernels in Python, which are then compiled for CPU or CUDA execution. The first launch builds and caches a native module, which is reused in subsequent launches. Warp offers three main value propositions:

1. Performance: Achieves native-CUDA speed through JIT compilation, kernel fusion, and CUDA Graphs.

2. Ease of Use: Enables pure Python authoring with built-in vectors, matrices, quaternions, BVHs, hash grids, sparse matrices, and tile primitives.

3. Capability: Features differentiable kernels and DLPack-style interop, allowing simulation to be integrated inside an ML training loop.

For robotics, three properties make Warp particularly useful:

  • Explicit Parallel Work: `wp.tid()` identifies the specific point, contact, body, or world owned by the current logical thread.
  • Explicit Device Arrays: Arrays live on the selected device. Calling `.numpy()` on a CUDA array synchronizes and copies it to CPU memory.
  • Composable Kernel Launches: Programs can launch sequences of focused kernels and capture supported CUDA work into a graph to reduce dispatch overhead.

Additionally, Warp kernels are differentiable, using a `wp.Tape` to record forward launches and replay their adjoints in reverse. From version 1.15, Warp also supports deterministic execution modes, trading some performance for reproducible ordering in validation and regression tests.

What is MuJoCo Warp (MJWarp)?

A robot simulator repeatedly computes the next state of a scene based on current joint positions, velocities, controls, and contacts. In this context, a "world" refers to one independent copy of that scene and its state. MJWarp takes compatible MuJoCo models into the GPU-scale regime. By leveraging NVIDIA Warp, MJWarp allows developers to scale familiar MuJoCo workflows—like the SO-101 follower arm example—up to 2,048 parallel environments, significantly accelerating the development, testing, and learning workflows for physical AI.

Source