الملخص
When the human reference scores zero on a metric, the released GTRS-Dense label generator for NAVSIM marks every candidate trajectory in the scene as passing it. NAVSIM's authors introduced this human-reference forgiveness to avoid penalizing contextually justified maneuvers when scoring one trajectory, and warned that it could overlook important failures. In label generation it sets a whole column of 16,384 candidate targets to passing. To measure the consequences for supervision, we re-run the generator with the overwrite disabled and compare the pre-overwrite targets with the released labels on all 103,288 navtrain scenes. The rule erases a candidate distinction that the training loss reads on 11,237 of them (10.8793%). Firing usually changes most of a column: lane keeping carries 9,982 of the 13,042 forgiven loss columns, and its median forgiven column had 14,391 of 16,384 candidates failing before the overwrite. On held-out navtest scenes forgiven on lane keeping, the released lane-keeping head's median AUC against the pre-overwrite outcome is 0.7095; on unforgiven scenes matched on failing-candidate count it is 0.9807. For the Hydra-MDP checkpoint released with GTRS, whose configuration takes the same label file, the two values are 0.6627 and 0.9761. Continuing the released GTRS-Dense checkpoint for 300 optimizer steps with three paired seeds, we observe the forgiven-scene AUC 0.1086-0.1251 higher with pre-overwrite than with published targets, and a narrower gap between matched groups, still above zero. Scoring with forgiveness disabled, we observe lane keeping higher by 2.478-3.524 points on navtest scenes forgiven on any of five loss metrics, with lower adjacent-frame plan consistency. Both changes are larger there than on the rest. EPDMS, scored the same way, does not separate the two target sets.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Guo, J., Yang, J., Ye, J., Sun, Y., Xin, S., Zhang, K., & Yang, H. (2026). When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM. https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim
MLA 9
Guo, Jiaxuan, et al. "When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM." https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim.
شيكاغو (المؤلف–التاريخ)
Guo, Jiaxuan, Jingxin Yang, Jiaqi Ye, Youran Sun, Shuo Xin, Kejia Zhang, and Haizhao Yang. 2026. "When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM." https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim.
هارفارد
Guo, J., Yang, J., Ye, J., Sun, Y., Xin, S., Zhang, K. and Yang, H. (2026) 'When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM', Available at: https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim.
فانكوفر
Guo J, Yang J, Ye J, Sun Y, Xin S, Zhang K, et al. When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM. https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim
IEEE
J. Guo, J. Yang, J. Ye, Y. Sun, S. Xin, K. Zhang, and H. Yang, "When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM," https://omanscience.com/ar/articles/when-an-evaluation-rule-writes-training-labels-measuring-human-reference-forgiveness-in-navsim.