Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models
A new method, Adaptive Relevance-guided Evidence Allocation (AREA), is introduced to improve multimodal large language models (MLLMs) by adaptively allocating evidence from images and text to answer visual questions. This approach dynamically adjusts the amount of evidence shown to the model based on relevance entropy and other factors, leading to better performance on various benchmarks.
Save an API key to vote.