UniWAM

UniWAM: Unified World-Action Model

Core Contributors

Jiayi Chen1,2 Jingbo Wang1,2 Shuai Zhou3 Xicheng Gong4

Contributors

Ziyang Zhou1,2 Zehua Fan5 Junwu E1 Haodong Yan1 Fuhao Li1,2 Qize Yu4 Xu Huang2 Pengwei Wang6 Wen Chen2 Haoang Li1,2

Project Lead

Wenxuan Song1

Project PI

Shunbo Zhou2

1The Hong Kong University of Science and Technology (Guangzhou) 2OLA Dimensions 3Carnegie Mellon University 4Peking University 5Shanghai Jiao Tong University 6Beijing Academy of Artificial Intelligence

Real-World Demonstrations

Pick-Anything

Identify and grasp the target object in a multi-object scene.

Diverse-Interaction

Ground both the target object and the requested interaction.

Place-Relative

Place an object in the instructed spatial relation to another object.

Drawer-Storage

Coordinate both arms to open a drawer and store the target object.

Long-Horizon Organization

Follow “Tidy up the desk.” through cup placement, drawer storage, and marker placement.

Abstract

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.

Contributions

  1. 1

    Three-source pre-training

    Robot demonstrations, human egocentric data, and VQA provide complementary supervision; source-level ablations evaluate their contributions.

  2. 2

    History-conditioned post-training

    Recent action history and perturbed future visual latents are incorporated into joint flow matching.

  3. 3

    Broad evaluation

    Simulation and real-world evaluations assess in-distribution performance, robustness, generalization, instruction following, and long-horizon execution.

  4. 4

    Human–robot co-training at scale

    The report studies a log-linear scaling relationship for unified human–robot co-training and the gains from mixed-data pre-training.

Key Results

99.2% LIBERO · Average success
92.6% LIBERO-Plus · Overall success
68.32% RoboTwin 2.0 · C2R success

UniWAM also achieves 75.14% success in RoboTwin 2.0 Clean2Clean (C2C). Clean2Rand (C2R) evaluates policies trained on clean data under domain randomization.

Comparisons below refer only to the methods listed in each table. Bold and underlined values reproduce the report’s table annotations; † preserves a source marker from the report. Swipe or scroll wide tables to see all columns.

LIBERO — success rate (%)
Method Spatial Object Goal Long Average ↑
π₀ 98.0 96.8 94.4 88.4 94.4
PD-VLA 95.5 96.7 94.9 91.7 94.7
π₀.₅ 98.8 98.2 98.0 92.4 96.9
GR00T-N1.7 97.7 98.5 97.5 94.4 97.0
OpenVLA-OFT 97.6 98.4 97.9 94.5 97.1
Fast-WAM 98.2 100.0 97.0 95.2 97.6
Motus 96.8 99.8 96.6 97.6 97.7
VLA-Adapter 99.6 99.6 98.2 96.4 98.5
X-VLA 98.2 98.6 97.8 97.6 98.1
Cosmos-Policy 98.1 100.0 98.2 97.6 98.5
LingBot-VA 98.5 99.6 97.2 98.5 98.5
Spatial Forcing 99.4 99.6 98.8 96.0 98.5
Xiaomi-Robotics-0 98.8 100.0 98.8 97.2 98.7
UniWAM (Ours) 99.6 99.6 99.2 98.4 99.2
LIBERO-Plus — success rate (%)
Method Camera Robot Language Light Background Noise Layout Total
π₀ 79.6 21.1 72.5 84.7 86.2 68.3 69.4 67.4
GR00T-N1.6 92.6 33.5 80.1 93.6 95.4 93.6 75.0 79.4
OpenVLA-OFT 92.8 30.3 85.8 94.9 93.9 89.3 77.6 79.5
Spatial Forcing 95.2 47.9 73.5 91.2 95.6 92.2 74.8 80.5
MemoryVLA 91.4 48.6 79.4 95.2 95.3 94.0 75.7 81.9
π₀.₅ 87.6 78.4 80.0 92.6 91.6 91.4 81.8 86.2
ACoT-VLA 96.6 70.4 79.7 95.1 97.1 95.9 85.0 88.0
Kairos 95.5 72.6 86.8 97.7 95.8 96.8 81.5 89.0
UniWAM (Ours) 92.0 89.5 92.2 97.9 97.6 94.3 84.6 92.6
RoboTwin 2.0 — success rate (%)
Method C2C C2R Average
VLA
GR00T-N1.7† 43.6 20.7 32.2
StarVLA† 58.1 10.6 34.4
Xiaomi Robotics-0 62.90 18.20 40.55
Abot-M0 57.40 30.36 43.88
X-VLA 68.00 20.90 44.45
Spatial Forcing 77.20 26.74 51.97
π₀.₅ 70.70 46.00 58.35
GigaBrain-0.7 66.80 67.90 67.35
WAM
AHA-WAM† 64.3 3.2 33.8
FastWAM 70.20 1.20 37.70
X-WAM 70.00 25.80 47.90
4D-WAM† 81.5 41.8 61.7
OpenWAM-α† 89.4 48.7 69.0
UniWAM (ours) 75.14 68.32 71.73

Real-world instruction following and long-horizon execution

Across four instruction-following tasks, UniWAM achieves an average task success rate of 67.5% and an instruction-following rate of 82.5%. In long-horizon tabletop organization, average Task Progress is 5.0 out of 6, compared with 4.8 for π₀.₅ and 3.2 for Motus.

Real-world task success, instruction following, and long-horizon Task Progress.
Real-world task success, instruction following, and long-horizon Task Progress. View full size ↗

Platform

The real-world setup uses two AgileX Piper arms, each with six degrees of freedom and a parallel gripper. Each arm has a wrist-mounted camera, complemented by a third-person camera with a global view of the workspace.

Four instruction-following tasks and a multi-stage tabletop organization task on the dual-arm platform.
Four instruction-following tasks and a multi-stage tabletop organization task on the dual-arm platform. View full size ↗

Model Architecture

UniWAM connects a Physical Reasoner, a World Generator, and an Action Predictor through joint multimodal attention. The Physical Reasoner uses Qwen3-VL-2B-Instruct; the World Generator uses Wan2.2-TI2V-5B; the Action Predictor models a continuous action flow field.

Architecture and data overview from the tech report.
Architecture and data overview from the tech report. View full size ↗

Training Recipe

Data processing and annotation

The data pipeline cleans and annotates robot and human demonstrations. For EgoDex, EgoANT segments recordings into atomic manipulation subtasks and produces concise descriptions of each segment.

Stage 1 · Pre-training

Robot demonstrations, human egocentric data, and visual question answering jointly support action learning, future visual prediction, and semantic understanding. Physical language expresses local manipulation behavior in natural language, providing shared supervision for human and robot data.

EgoANT segmentation and labeling pipeline for human egocentric data.
EgoANT segmentation and labeling pipeline for human egocentric data. View full size ↗

Stage 2 · Post-training

Future visual noise augmentation reduces dependence on precise future predictions. History-conditioned flow matching initializes action generation from perturbed action history and refines it toward future actions; the visual source remains Gaussian noise.

History-conditioned action denoising, compared with generation initialized from Gaussian noise.
History-conditioned action denoising, compared with generation initialized from Gaussian noise. View full size ↗

Scaling human–robot co-training

The report compares robot-only and robot-plus-human training with VQA in both settings. Human data becomes beneficial at larger robot-data scales, alongside improvements in validation loss and downstream RoboTwin performance.

Data scaling: validation MSE and RoboTwin Clean2Rand success rate.
Data scaling: validation MSE and RoboTwin Clean2Rand success rate. View full size ↗

Ablation studies

Ablations examine pre-training data composition, the Physical Reasoner’s training strategy, and action generation. The full three-source configuration reaches 71.73% average success across RoboTwin C2C and C2R.

RoboTwin ablations of pre-training data, reasoner training, and action generation.
RoboTwin ablations of pre-training data, reasoner training, and action generation. View full size ↗

BibTeX

@misc{chen2026uniwam,
  title={UniWAM: Unified World-Action Model},
  author={Jiayi Chen and Wenxuan Song and Jingbo Wang and Shuai Zhou and Xicheng Gong and Zehua Fan and Ziyang Zhou and Junwu E and Haodong Yan and Fuhao Li and Qize Yu and Xu Huang and Pengwei Wang and Wen Chen and Shunbo Zhou and Haoang Li},
  year={2026},
  eprint={2610.02054},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2610.02054}
}