PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
PetriBench is a new benchmark for evaluating large language model (LLM) reasoning over dynamic state spaces using Petri nets. It assesses LLM capabilities in four task families with varying scope and temporal horizons, and its results show that accuracy decreases with difficulty and that test-time compute interacts differently with different reasoning tasks.
Save an API key to vote.