Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
A new learning framework, PIVOT, is proposed to improve the reasoning capabilities of large vision-language models by anchoring policy optimization around informative visual reasoning signals. This is achieved through a self-calibrated experience replay mechanism and a vision-guided advantage allocation mechanism.
Save an API key to vote.