BadWAM: When World-Action Models Dream Right but Act Wrong
Qi Li, Xingyi Yang, Xinchao Wang
Read on arXiv →Key claim
WAMs are vulnerable to new adversarial attack types.
In plain English
World-action models (WAMs) are designed to couple action generation with future predictions, which is thought to enhance their robustness and safety. However, this paper reveals that the assumption of alignment between imagined futures and executed actions is fragile. It introduces BadWAM, a framework that characterizes new adversarial attacks specific to WAMs, demonstrating how small visual perturbations can lead to significant failures in task execution. Builders should consider these vulnerabilities when developing systems that rely on WAMs for embodied control.
Introduces a new framework for evaluating vulnerabilities in World-Action Models.
Presents empirical results across different WAM variants, demonstrating significant attack impacts.
Deep reliability assessment
The methodology supports the claim that small visual perturbations can desynchronize action generation from future prediction in WAMs, leading to task failures. However, the generalization of these results to all WAMs may be overclaimed without broader testing across diverse models and environments.
Reproducibility
No open source code or dataset is mentioned in the paper, making reproducibility challenging.
Key figure
Figure 1 likely illustrates the architecture of the BadWAM framework, highlighting the interaction between action generation and future prediction.
