Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation

May 17, 2026 Β· Grace Period Β· πŸ› CVPR 2026 Workshop

⏳ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Nicanor Mayumu, Xiaoheng Deng, Patrick Mukala arXiv ID 2605.17268 Category cs.AI: Artificial Intelligence Cross-listed cs.CV, cs.RO Citations 0 Venue CVPR 2026 Workshop
Abstract
We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 100 diverse PhysicalAI-AV scenarios. Our main finding is that output natural-language rationales with trajectories may be significantly unfaithful: (i) overall reasoning fidelity is only 42.5%, with Chain-of-Causation matching scene reality less than half the time; (ii) 94 missed pedestrians in one-third of pedestrian-relevant scenes; (iii) 97.7% trajectory fragility under mild visual perturbations; and (iv) only 48.3% mean reasoning-action consistency, with 53.3% of inferences exhibiting low consistency, including 37.9% of stop-claimed cases where the model continues instead. We formalize faithfulness information-theoretically, define entity and action fidelity with verification criteria, and outline a four-component safety architecture aligned with these results.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

πŸ“œ Similar Papers

In the same crypt β€” Artificial Intelligence