BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
Researchers propose a new benchmark, BUILD-BENCH, to evaluate the performance of LLM agents in compiling real-world open-source software. The benchmark includes diverse OSS projects with various characteristics, and a strong baseline LLM-based agent, OSS-BUILD-AGENT, achieves state-of-the-art performance on it. This work aims to improve the realism and complexity of evaluating LLM agents in software engineering tasks.
Save an API key to vote.