Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

arXiv CS Wednesday 03 June 2026, 04:00 UTC By S Divakar Bhat, Toshihiko Yamasaki 1 min read

Key Points

arXiv:2606.02742v1 Announce Type: new Abstract: Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance queries. A common assumption is that consistent predictions across viewpoints reflect geometric grounding. We test this assumption and find the opposite: leading VLMs often produce view-invariant and consistent answers even when those answers are incorrect, indicating weak coupling between predictions and viewpoint-specific visual evidence. We introduce \textbf{ViewDiag}, a controlled multi-view evaluation protocol built from Hypersim, ScanNet, and KITTI360, comprising 176 object-pair tracks across 80 scenes with 2--10 views per track. The protocol evaluates models along three axes: metric accuracy, distributional concentration, and a latent feature probe for internal collapse that distinguishes decision collapse from representation collapse. Across diverse models, we observe a consistent pattern of high prediction stability paired with substantial error, clustering in a regime characterized by strong consistency but low accuracy. \noindent These results challenge the common use of cross-view consistency as a proxy for geometric understanding. Instead, we show that stable predictions may reflect prior-driven collapse rather than evidence-sensitive reasoning. ViewDiag provides a controlled benchmark and diagnostic framework for evaluating spatial VLMs beyond accuracy alone. The code and data can be found \href{https://github.com/SDivakarBhat/Consistent_Yet_Wrong.git}{here}

Hypersim (LOCATION) ViewDiag (PERSON)

Originally published by arXiv CS Read original →

Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

Related Stories

You can personalize your Instagram algorithm now — unless you want to see more posts from accounts you follow

Super Micro Seeks $7B in Equity Deal for AI Equipment

Ubisoft reportedly shuts down more studios and lays off staff in Barcelona and San Francisco

Anthropic CEO Says Government Should Be Able to Block New Models