What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
This research paper maps the evolution of Large Language Model (LLM) benchmarks, highlighting changes in evaluation requirements and what researchers expect from LLMs. The study analyzes 14,767 papers and identifies growing emphasis on action, interaction, and professional applications, with implications for AI development and testing.
Save an API key to vote.