Abstract

Replay selectors often rank cached trajectories by format feedback, confidence, freshness, or response length, although cache-level correctness and downstream learner utility are distinct objectives. We formalize this selection-to-learning gap and introduce FAER as an auditable full-trajectory replay framework. Its training-free fixed selector is a protocol baseline; FAER-UTILITY is the learner-aware selector fitted on disjoint calibration blocks. The normalized gradient alignment is reported as a baseline, while a disposable optimizer-aware virtual update supplies a magnitude-aware utility surface. The audit contract freezes observed fields and replay traces before evaluation labels are joined. On GSM8K with Qwen2.5-1.5B-Instruct, the matched learner study reports quality 0.6329 for the fixed selector, compared with 0.5482 for uniform and 0.6037 for format-feedback under 128 updates. Metadata-only cross-fitted calibration reaches $0.6476\!\pm\!0.0139$ over eight seeds (median 0.6481; paired 95% interval $[+0.079,+0.122]$) at 63,276 target-run tokens; its recorded full cost is 189,642 tokens and 3.48 GPU-hours including calibration. The completed FAER-UTILITY row reaches 0.6624 at 62,844 target-run tokens and 4.26 GPU-hours. Format-feedback selects records with correctness 0.6953, compared with 0.3594 for the fixed selector, despite the different downstream ranking. The completed comparison surfaces report the learner-aware ablation, same-seed gap, policy-optimization rows, and strict zero-shot transfer.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hu, M., Hu, S., Guo, X., Wang, X., Wang, B., Fu, T., Zha, D., & Xiao, J. (2026). FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training. https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training

MLA 9

Hu, Miaobo, et al. "FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training." https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training.

Chicago (author–date)

Hu, Miaobo, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Tianshu Fu, Daren Zha, and Jun Xiao. 2026. "FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training." https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training.

Harvard

Hu, M., Hu, S., Guo, X., Wang, X., Wang, B., Fu, T., Zha, D. and Xiao, J. (2026) 'FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training', Available at: https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training.

Vancouver

Hu M, Hu S, Guo X, Wang X, Wang B, Fu T, et al. FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training. https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training

IEEE

M. Hu, S. Hu, X. Guo, X. Wang, B. Wang, T. Fu, D. Zha, and J. Xiao, "FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training," https://omanscience.com/en/articles/faer-auditable-utility-aligned-trajectory-replay-for-language-model-post-training.