الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Yuan, M., Liang, X., Yang, J., Chen, Z., Zhang, Z., Wang, H., Wang, Y., & Li, J. (2026). الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات. https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression

MLA 9

Yuan, Mingqing, et al. "الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات." https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression.

شيكاغو (المؤلف–التاريخ)

Yuan, Mingqing, Xiaobo Liang, Junwei Yang, Ziwei Chen, Zeren Zhang, Hejin Wang, Yubin Wang, and Juntao Li. 2026. "الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات." https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression.

هارفارد

Yuan, M., Liang, X., Yang, J., Chen, Z., Zhang, Z., Wang, H., Wang, Y. and Li, J. (2026) 'الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات', Available at: https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression.

فانكوفر

Yuan M, Liang X, Yang J, Chen Z, Zhang Z, Wang H, et al. الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات. https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression

IEEE

M. Yuan, X. Liang, J. Yang, Z. Chen, Z. Zhang, H. Wang, Y. Wang, and J. Li, "الحكم في الفضاء الكامن: نمذجة مكافأة توليدية فعّالة عبر ضغط يحفظ الدلالات," https://omanscience.com/ar/articles/judging-in-latent-space-efficient-generative-reward-modeling-via-semantics-preserving-compression.