VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

VidOmni-Bench is a new benchmark for evaluating fine-grained video understanding in Video Large Language Models (Video-LLMs). It assesses whether models can accurately verify events in video captions, revealing weaknesses in current Video-LLMs that generate hallucinated descriptions and struggle with detecting incorrect event descriptions.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.