Skill-based Agentic Evaluation for Real-time Data Science Tasks
Researchers introduced ground-truth-as-code, a framework for evaluating data-science agents on live data using executable ground truth and format-agnostic factoid scoring. This approach addresses the challenge of static references becoming outdated in dynamic data environments. The framework validates a 29% improvement in agreement between expert annotators and LLM-as-a-judge predictions and a 16% reduction in token consumption.
Save an API key to vote.