Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Researchers identify a readout gap in vision-language models for harmful meme detection, where models miss internal evidence or struggle to route represented evidence to the output. They propose using sparse autoencoders and role-conditioned probes to improve performance, achieving significant gains in macro-F1 scores. This affects AI agents by highlighting a common bottleneck in harmful content classification and suggesting techniques to address it.
Save an API key to vote.