Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

Researchers propose Lens, a training-free framework for multimodal representation learning, achieving a 10.2-point improvement over a baseline embedding model on 36 MMEB datasets.

RSS Score 0 9/18/2026, 4:00:00 AM Original Source
Save an API key to vote.