UMI-WM: Interactive World Modeling for Bimanual Manipulation from Wrist Cameras

Jingxiang Guo1Haoyu Zhao1Xidong Zhang2Yuhang Zheng1Hao Wu1
Chen Gao1,†Jianshu Zhou1,†Shuicheng Yan1,†Weihao Yuan3,†
1National University of Singapore2Harbin Institute of Technology3Nanjing University

jingxiangguo@u.nus.edu · †Corresponding authors

UMI-WM uses wrist observation history, camera poses, and gripper widths to generate paired wrist videos for teleoperation and action recovery.
UMI-WM jointly predicts both wrist views from per-arm camera poses and gripper widths.
An inverse dynamics model recovers actions for evaluation and robot execution.

Abstract

Universal Manipulation Interface (UMI) pairs wrist-camera videos with camera trajectories and gripper openings, providing scalable, robot-free data for learning how bimanual actions change a shared scene. Existing interactive world models, however, typically rely on text or single-camera control and struggle to follow each arm's actions while maintaining consistency between the two wrist views. We introduce UMI-WM, an action-conditioned bimanual world model that jointly generates both wrist views from observation history, per-arm relative camera poses, and gripper widths, with cross-view consistency, action following, and visual fidelity. UMI-WM injects each arm's actions into its own wrist stream, while sparse hub tokens exchange scene information between the two streams. Camera-aware geometry conditioning and camera-guided noise further support per-view motion control and cross-view geometric exchange. We adapt the model for block-causal generation and distill the 30-step Teacher into a Student with four denoising steps. We train UMI-WM on self-collected open-world and in-studio UMI recordings combined with public manipulation data. To evaluate action following, we use an inverse dynamics model (IDM) to recover relative camera poses and gripper widths from generated videos, and introduce Bimanual Action Fidelity (BAF) to measure how closely these recovered actions match the conditioning actions. UMI-WM achieves the best results on all five reference-video fidelity metrics and the best or competitive performance on VBench and WorldArena. The Teacher and Student outperform all compared baselines in action following, while the four-step Student generates paired wrist videos 6.9× faster than the Teacher.

Paired-Wrist Video Predictions

Synchronized paired-wrist predictions across five tasks.

Unzipping a backpack
Three timestamps from the paper
Ground truth, Cosmos, MiniMax-H3, WorldPlay, UMI-WM Teacher and Student compared at 1, 3, and 5 seconds, with paired left and right wrist views.
Each row shows one method, with paired-wrist observations at 1, 3, and 5 seconds.

Quantitative Results

Evaluation uses 320 held-out UMI episodes with 1,008 paired-wrist windows. The Teacher leads PSNR, LPIPS, FID, and FVD, while the Student leads SSIM. Both variants outperform all compared baselines in Bimanual Action Fidelity (BAF), which measures agreement between conditioning actions and actions recovered from generated videos.

Reference-video fidelity and efficiency from Table 1, with Bimanual Action Fidelity from Table 3. Bold / underline mark the best / second-best generated-model results shown.
ModelPSNR ↑SSIM ↑LPIPS ↓FID ↓FVD ↓BAF ↑Paired FPS ↑
Wan2.2 TI2V-5B8.3450.3120.71599.763477.0618.640.18
Cosmos Predict2.57.7590.3080.68598.62967.4017.540.10
MiniMax-H39.4770.3710.57758.91687.9415.200.70
Gamma-World ◇12.2640.5480.39740.30331.2552.930.55
HY-WorldPlay 5B8.5990.3380.60668.15638.0836.530.81
LingBot-World-Fast7.6600.2700.752128.162190.3624.192.32
UMI-WM Teacher12.5130.5760.33233.52232.5376.360.26
UMI-WM Student12.3340.5800.33637.89284.6871.551.81

◇ Gamma-World is fine-tuned on the UMI training split. FID and FVD use one selected window per episode. Paired FPS counts synchronized wrist-frame pairs on one A100-80GB GPU. The four-step Student reaches 1.81 paired FPS versus 0.26 for the 30-step Teacher, a 6.9× offline speedup. Scores use two generation seeds and equal episode weights; BAF is an action-agreement score, not a task-success percentage.

VBench and WorldArena results

All reported VBench and WorldArena metrics from Table 1. Higher is better. GT is an unranked reference; bold / underline indicate the best / second-best generated-model results, including ties.
ModelVBench
Subject consistency ↑
VBench
Background consistency ↑
VBench
Motion smoothness ↑
VBench
Imaging quality ↑
WorldArena
JEPA similarity ↑
WorldArena
Photometric smoothness ↑
WorldArena
Perspectivity ↑
GT video0.8600.9260.9840.7161.0003.3773.750
Wan2.2 TI2V-5B0.8240.9030.9880.6040.0591.8273.875
Cosmos Predict2.50.6880.8500.9740.6310.2841.7473.500
MiniMax-H30.7680.8740.9820.6090.0801.7723.375
Gamma-World ◇0.8740.9330.9840.6550.3824.2363.375
HY-WorldPlay 5B0.8250.8990.9680.7050.1352.2963.250
LingBot-World-Fast0.6270.8550.9580.4640.1791.6693.250
UMI-WM Teacher0.8850.9380.9840.7070.3864.3393.500
UMI-WM Student0.8750.9330.9840.6550.4084.3003.125

Method Overview

UMI-WM jointly predicts both wrist videos from observation history and per-arm actions, represented by relative translation, a 6D rotation representation, and gripper width in 10 dimensions. Each arm’s actions are injected into its own wrist stream. SRAE encodes view, role, and actor identities, while Sparse Hub Attention exchanges scene information through learned hub tokens. CameraNoise supplies motion-aware noise, while CamRay and PRoPE provide camera-aware geometry conditioning for per-view motion control and cross-view geometric exchange. We adapt the model for block-causal generation and distill the 30-step Teacher into a Student with four denoising steps.

UMI-WM Teacher and Student architecture, per-arm temporal action encoder, cross-view modules, camera conditioning, and inverse dynamics model.
The world model maps bimanual actions to paired wrist videos. The IDM separates gripper foreground and scene background to recover camera poses and gripper widths for BAF evaluation and robot execution via inverse kinematics.

Data Sources and Processing Pipeline

A three-tier data pyramid combines public manipulation data, self-collected open-world UMI recordings, and calibrated in-studio recordings. The pipeline synchronizes wrist videos, camera poses, and gripper states at 30 Hz, expresses each trajectory relative to its window’s first frame, and applies calibration and quality checks. Open-world wrist trajectories remain in separate local frames; in-studio recordings register both wrists in a shared world frame. UMI-WM uses the two wrist views as visual observations.

Self-collected in-studio and open-world UMI demonstrations and public robot data are synchronized, canonicalized, and quality checked.

Appendix: Physical Robot Execution

After interactive approval of a generated rollout, the IDM recovers wrist-camera poses and gripper widths. A calibrated mapping converts the poses to end-effector targets for inverse kinematics, while the recovered widths drive the grippers. The left panel shows physical robot execution; the right panel shows the corresponding wrist-view preview.

Physical executionModel preview
Left: physical robot execution. Right: the corresponding world-model preview.