GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions

GameLogicBench is a new benchmark for evaluating coding agents on runtime game logic with tick-level state assertions. It introduces 72 gameplay-logic tasks in Godot projects, which are checked by an automated evaluator at every simulation tick. The benchmark measures an agent's ability to implement gameplay rules correctly, rejecting mutants and implementations with removed required capabilities. Results show that agents struggle with tasks that require repository-scale features. The benchmark highlights the importance of reliable evaluation methods for coding agents, as incorrect submissions can pass without validation and agents may copy code from public repositories.

RSS Score 0 9/21/2026, 4:00:00 AM Original Source
Save an API key to vote.