PointWAM

Point World Action Model

3D World Action Modeling for Dexterous Robotic Manipulation

scene trajectories hand trajectories hand keypoints

*Co-first authors  ·  †Equal contribution

Forecast the scene and the hands as 3D points,
then retarget the hands to robot actions.

One 3D frame.

Scene points and hand keypoints share one space-time coordinate frame.

Pre-trained on human videos.

1.15M EgoDex and VITRA episodes, with no robot actions: +56.9 points.

State of the art on DexJoCo.

69.0% over ten dexterous tasks, 11.7 points above GR00T N1.6.

On a real robot.

Ahead of GR00T N1.6 and π0.5 on an OpenArm humanoid.

Overview

Large-scale human demonstration pre-training

PointWAM forecasts the scene and the hands as 3D point trajectories. Human and robot demonstrations share this form, so 1.15M human episodes from EgoDex [1] and VITRA [2] pre-train it.

1.15M human episodes · pre-training
Robot demonstrations · fine-tuning
Abstract

World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.

Method

Forecast in 3D, then retarget the hands

Drawn from real data (DexJoCo microwave /B). ❄ frozen module.

Trajectory forecaster

One transformer reads scene, hand and instruction tokens, and two heads forecast every point. No task-specific object is selected.

Action retargeter

A decoder maps the forecast hand trajectories and the robot state to an action chunk. Scene forecasts supervise the trunk only.

Human first, robot second

Pre-train the forecaster on 1.15M human episodes from EgoDex and VITRA, then fine-tune on robot demonstrations.

Interactive

Explore the trajectories in 3D

Scene and hand trajectories in 3D, from the robot and human demonstrations PointWAM learns from.

loading    
t = 0

Robot

Human

Drag to orbit · scroll to zoom · right-drag to pan

Results

One policy, ten dexterous tasks

DexJoCo [3]: Franka arms with 16-DoF Allegro hands, six single-arm and four bimanual tasks. Each task is run for 50 episodes on each of three seeds.

Baselines: DP-T [4], π0.5 [5], GR00T N1.6 [6], Fast-WAM [7], PointACT [8].

Per-task success rates
TaskDP-Tπ0.5GR00T N1.6Fast-WAMPointACTPointWAM

SR (%), subscripts = std over three seeds, bold = row best, /B = bimanual. DP-T and π0.5 are from the DexJoCo paper.

Human pre-training transfers, and it scales

Mean DexJoCo SR over three seeds; whiskers and bands are ±1 std. Baselines build on large pre-trained models; PointWAM trains its trunk from scratch.

What matters

Volt-L12 reference, one change at a time.

(a) vision encoder
DINOv3 [11] (lifted 2D)59.0
Mosaic3D [9] (3D)64.4
(b) world modeling
without53.5
2D frame latents60.1
3D points64.4
(c) hand representation
keypoints, 1064.4
surface points, 1043.6
surface points, 51244.6
(d) voxel size
0.5 cm59.9
1 cm64.4
2 cm52.2

Real robot

OpenArm, two Inspire hands, one ZED 2i; 24 trials per task.

The OpenArm platform with two arms, two Inspire hands and a ZED 2i camera above the table
Rollouts

On a real robot and in DexJoCo

Real robot

DexJoCo

One multi-task policy per method across the ten tasks, each rollout from the same randomized object placement. Rollouts are chosen in proportion to PointWAM's success rate on the task; played at 2× speed.

Citation

BibTeX

@article{park2026pointwam,
  title   = {{PointWAM}: {3D} World Action Modeling for Dexterous Robotic Manipulation},
  author  = {Park, Chunghyun and Kim, Beomjun and Park, Seungcheol and Kwon, Heeseung and
             Shukla, Yashu and Sim, Seunghoon and Shin, Jinwoo and Cho, Minsu},
  journal = {arXiv preprint},
  year    = {2026}
}
References
  1. Hoque et al. EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video. ICLR 2026.
  2. Li et al. Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos. ICRA 2026.
  3. Wang et al. DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo. arXiv 2026.
  4. Chi et al. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. RSS 2023.
  5. Black et al. π0.5: a Vision-Language-Action Model with Open-World Generalization. CoRL 2025.
  6. NVIDIA et al. GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv 2025.
  7. Yuan et al. Fast-WAM: Do World Action Models Need Test-time Future Imagination? arXiv 2026.
  8. Chen et al. PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction. RSS 2026.
  9. Lee et al. Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation. CVPR 2025.
  10. Yilmaz et al. Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding. ECCV 2026.
  11. Siméoni et al. DINOv3. TMLR 2026.