نسخة أولية وصول مفتوح
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven …