AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
A new video question answering benchmark, AgentVidBench, is introduced to evaluate the spatial, temporal, and causal reasoning capabilities of MLLM agents. The benchmark provides step-by-step solution traces to assess whether agents acquire evidence to justify their answers. Experiments with 12 MLLMs show that integrating these models into agentic workflows improves performance on AgentVidBench.
Save an API key to vote.