ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement
A research paper introduces ProtoLIP, a lightweight prototype-mediated evidence layer for query-conditioned vision-language models. It improves evidence localization and separation without requiring spatial annotations or backbone retraining, and remains competitive with a spatially supervised grounding model.
Save an API key to vote.