Home /Research /Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color
MANIPULATION

Lights, Camera, Malfunction: When Illumination Robustness Leaves VLA Models Blind to Color

Marino Watanabe, Takami Sato, Kentaro Yoshioka

Year
2026
Access
Open access

Abstract

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for general-purpose robot manipulation; however, their transition to real-world environments reveals vulnerabilities to minor environmental perturbations. We propose FLARE, an optimized physical spotlight attack framework that exploits these vulnerabilities via targeted illuminations, dropping baseline task success rates to zero without any access to model internals. While adversarial training is the standard countermeasure, we identify a critical and previously underestimated defensive pitfall: naive data augmentations incorrectly condition VLA models to discard color as noise, collapsing their visual perception into a purely shape-biased processor. We expose this degradation through a diagnostic grayscale evaluation, in which the defended model maintains high success rates on grayscale inputs, while its success rate on benign, color-dependent real-world tasks drops to at most 47.5%, well below the undefended baseline. To address this, we propose ChromaGuard, a chroma-preserving adversarial training method. On a physical 6-DoF robotic platform, we demonstrate that ChromaGuard achieves 97.5% and 92.5% success rates in benign and attacked color-dependent tasks, respectively.

Keywords

VLA模型光照攻击对抗训练颜色感知机器人操控

Related papers

Browse all MANIPULATION papers