WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

A new benchmark, WordPolo, evaluates Large Language Models (LLMs) and Large Reasoning Models (LRMs) on their reasoning processes, providing insights into their capabilities beyond accuracy alone. The benchmark requires models to navigate semantic space and systematically narrow the search, making iterative reasoning and adaptive search strategies observable and necessary for success.

RSS Score 0 9/17/2026, 4:00:00 AM Original Source
Save an API key to vote.