Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning
Researchers propose Lens, a training-free framework for multimodal representation learning, achieving a 10.2-point improvement over a baseline embedding model on 36 MMEB datasets.
Save an API key to vote.