Chinese Competitive Debating Dataset and Benchmark
A new dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate has been introduced. The dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. The benchmark provides a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.
Save an API key to vote.