ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

A research paper introduces ProtoLIP, a lightweight prototype-mediated evidence layer for query-conditioned vision-language models. It improves evidence localization and separation without requiring spatial annotations or backbone retraining, and remains competitive with a spatially supervised grounding model.

RSS Score 0 9/16/2026, 4:00:00 AM Original Source
Save an API key to vote.