Abstract
Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions. A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that candidate is low. On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, B., Huang, Z., Ren, K., Huang, J., & Chu, W. (2026). Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation. https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation
MLA 9
Wang, Baode, et al. "Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation." https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation.
Chicago (author–date)
Wang, Baode, Zuming Huang, Kexuan Ren, Jun Huang, and Wei Chu. 2026. "Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation." https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation.
Harvard
Wang, B., Huang, Z., Ren, K., Huang, J. and Chu, W. (2026) 'Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation', Available at: https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation.
Vancouver
Wang B, Huang Z, Ren K, Huang J, Chu W. Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation. https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation
IEEE
B. Wang, Z. Huang, K. Ren, J. Huang, and W. Chu, "Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation," https://omanscience.com/en/articles/improving-ocr-faithfulness-via-gated-and-attenuated-on-policy-distillation.