نسخة أولية وصول مفتوح
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri …