Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
Researchers introduce CRYSTAL, a diagnostic benchmark for evaluating multimodal reasoning in large language models (LLMs). CRYSTAL assesses LLMs through verifiable intermediate steps, revealing systematic failures such as cherry-picking and disordered reasoning. The authors also propose the Causal Process Reward (CPR) and CPR-Curriculum, which improve reasoning and step-level alignment.
Save an API key to vote.