Vision-language-action models draw on large-scale visual and language knowledge, but typically predict actions directly from the current observation. World-action models add future-state reasoning. For manipulation, the useful future includes not only what the scene may become, but also where task-relevant changes occur and how those regions move.
Bridge-WA combines a Latent World Dynamics Module (LWDM) with WorldBridge. During training, frozen WAN VAE and DINOv3 encoders extract future-state and change targets from observed robot demonstrations; optical flow provides local motion targets. LWDM predicts these complementary world priors from the current multimodal context. WorldBridge conditions the action transformer through layer-specific multi-source attention, spatiotemporal biases, and reliability-gated feature modulation.
LWDM predicts future-state, temporal-change, and motion-flow representations from current observations and language. WorldBridge selectively supplies these priors to a flow-matching action expert using world-guided attention and reliability-gated modulation. Future-frame target encoders and cached supervision are used only during training.
Training targets from observed futures. Frozen WAN VAE and DINOv3 features represent future states and reveal regions that change between current and future observations. Optical flow supplies local displacement. Five learned-query heads predict these targets from the present multimodal context; no separate world-model pretraining is required for this version.
Selective world-to-action guidance. WorldBridge routes future, change, and motion memories to different action-transformer layers. Attention biases emphasize likely-changing and moving regions; learned reliability gates reduce the influence of world priors when they are less compatible with the current action features. Matched clean and augmented training further stabilizes actions under visual perturbations.
@misc{bai2026bridgewa,
title = {Bridge-WA: Learning Action-Relevant World Dynamics for Robotic Manipulation},
author = {Bai, Yongjie and Dai, Mingtong and Wang, Zhouxia and Wang, Hanting and Zhong, Qijun and Yan, Feng and Liu, Yang and Lin, Liang},
year = {2026},
url = {https://arxiv.org/abs/2607.02195},
note = {Project page: https://hcplab-sysu.github.io/BRIDGE-WA}
}