One 3D frame.
Scene points and hand keypoints share one space-time coordinate frame.
3D World Action Modeling for Dexterous Robotic Manipulation
*Co-first authors · †Equal contribution
Forecast the scene and the hands as 3D points,
then retarget the hands to robot actions.
Scene points and hand keypoints share one space-time coordinate frame.
1.15M EgoDex and VITRA episodes, with no robot actions: +56.9 points.
69.0% over ten dexterous tasks, 11.7 points above GR00T N1.6.
Ahead of GR00T N1.6 and π0.5 on an OpenArm humanoid.
PointWAM forecasts the scene and the hands as 3D point trajectories. Human and robot demonstrations share this form, so 1.15M human episodes from EgoDex [1] and VITRA [2] pre-train it.
World action models jointly learn to forecast world dynamics and predict robot actions, such that the learned internal world dynamics guide accurate actions. Existing approaches typically represent the world as RGB frames or latent counterparts while predicting actions as end-effector poses or joint angles, but they often struggle to capture the 3D spatial structure and contact geometry central to dexterous manipulation. We introduce Point World Action Model (PointWAM), a 3D world action model that decomposes the world into a scene (i.e., environment) and hands (i.e., actor), and jointly forecasts both as 3D point trajectories within a shared space-time coordinate frame. This explicit, disentangled representation enables effective pre-training on large-scale human demonstration videos without requiring any task-specific object or keypoint selection. Given a colored point cloud and a language instruction, PointWAM predicts how the scene and hands co-evolve in 3D space over time, then retargets the forecast hand motion to robot actions. Pre-training on human videos improves average DexJoCo success by 56.9 percentage points, and scene-trajectory supervision adds 10.9 points over forecasting the hands alone. With both, PointWAM surpasses the prior state of the art on ten DexJoCo tasks by 11.7 points and outperforms strong VLAs on a real robot.
One transformer reads scene, hand and instruction tokens, and two heads forecast every point. No task-specific object is selected.
A decoder maps the forecast hand trajectories and the robot state to an action chunk. Scene forecasts supervise the trunk only.
Pre-train the forecaster on 1.15M human episodes from EgoDex and VITRA, then fine-tune on robot demonstrations.
Scene and hand trajectories in 3D, from the robot and human demonstrations PointWAM learns from.
Drag to orbit · scroll to zoom · right-drag to pan
DexJoCo [3]: Franka arms with 16-DoF Allegro hands, six single-arm and four bimanual tasks. Each task is run for 50 episodes on each of three seeds.
Baselines: DP-T [4], π0.5 [5], GR00T N1.6 [6], Fast-WAM [7], PointACT [8].
| Task | DP-T | π0.5 | GR00T N1.6 | Fast-WAM | PointACT | PointWAM |
|---|
SR (%), subscripts = std over three seeds, bold = row best, /B = bimanual. DP-T and π0.5 are from the DexJoCo paper.
Mean DexJoCo SR over three seeds; whiskers and bands are ±1 std. Baselines build on large pre-trained models; PointWAM trains its trunk from scratch.
Volt-L12 reference, one change at a time.
OpenArm, two Inspire hands, one ZED 2i; 24 trials per task.

One multi-task policy per method across the ten tasks, each rollout from the same randomized object placement. Rollouts are chosen in proportion to PointWAM's success rate on the task; played at 2× speed.
@article{park2026pointwam,
title = {{PointWAM}: {3D} World Action Modeling for Dexterous Robotic Manipulation},
author = {Park, Chunghyun and Kim, Beomjun and Park, Seungcheol and Kwon, Heeseung and
Shukla, Yashu and Sim, Seunghoon and Shin, Jinwoo and Cho, Minsu},
journal = {arXiv preprint},
year = {2026}
}