Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

Researchers found that verbose prompts improve the robustness of vision-language models by broadening the spectral filter over image patches, reducing answer drift variance by 70-81% on 8B models. This can be achieved by padding the prompt, which also yields measurable gains in accuracy under image corruption.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.