Same Answer, Different Representations: Hidden instability in VLMs

Researchers found hidden instability in Vision Language Models (VLMs) by measuring internal representation drift, spectral sensitivity, and structural smoothness. This instability can lead to models producing unchanged outputs while undergoing significant internal changes. The study also found that larger VLMs are not necessarily more robust and that perturbations affect tasks differently, highlighting the need for more robust evaluation frameworks for VLMs.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.