MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
MUSE is a benchmark for evaluating large vision-language models on artistic image understanding in educational settings. It decouples image annotation from question generation and covers diverse tasks with controllable difficulty to assess AI models' capabilities in interpreting artistic imagery, understanding semantic, affective, and cultural content, and reasoning about visual context.
Save an API key to vote.