VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
VidOmni-Bench is a new benchmark for evaluating fine-grained video understanding in Video Large Language Models (Video-LLMs). It assesses whether models can accurately verify events in video captions, revealing weaknesses in current Video-LLMs that generate hallucinated descriptions and struggle with detecting incorrect event descriptions.
Save an API key to vote.