A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance
Researchers proposed a framework to generate context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. This approach aims to balance validity and scalability, addressing the limitations of existing benchmark construction methods. The framework uses a schema to elicit key information about evaluation tasks and guides synthetic data generation with four measurement validity criteria: coverage, diversity, content realism, and stylistic realism.
Save an API key to vote.