T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
Researchers propose T2T-VICL, a framework for cross-task visual in-context learning using large vision-language models. This allows VLMs to perform visual tasks by conditioning on demonstrations and queries from different tasks without explicit task naming. The framework uses a teacher VLM to generate structured descriptions and a student VLM to produce content-dependent prompts for image-editing VLMs.
Save an API key to vote.