الملخص

Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating inference cost, response time, average correctness, and correctness across repeated request conditions. Controlled comparisons vary hypothesis visibility, requested outputs, and output order while keeping the contract and target judgment fixed. Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy. Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general stability advantage. Development diagnostics further reveal compensating corrections and regressions, as well as persistent errors. These findings motivate evaluating cost and response time alongside whether individual judgments remain correct as the request configuration changes. Code: https://github.com/ZF-Utokyo/Jev-Benchmark

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhang, F., Chen, Y., Xie, Z., Zhou, Y., Peng, S., Fan, L., Ji, X., Zheng, C., Shan, H., Yu, P. S., Liu, X., Chen, Y., Nakov, P., & He, S. (2026). Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding. https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding

MLA 9

Zhang, Fan, et al. "Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding." https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding.

شيكاغو (المؤلف–التاريخ)

Zhang, Fan, Yankai Chen, Zhuohan Xie, Yixi Zhou, Sijia Peng, Lei Fan, Xinhua Ji, Cunyuan Zheng, Huangyong Shan, Philip S. Yu, Xue Liu, Yu Chen, Preslav Nakov, and Songwei He. 2026. "Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding." https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding.

هارفارد

Zhang, F., Chen, Y., Xie, Z., Zhou, Y., Peng, S., Fan, L., Ji, X., Zheng, C., Shan, H., Yu, P. S., Liu, X., Chen, Y., Nakov, P. and He, S. (2026) 'Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding', Available at: https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding.

فانكوفر

Zhang F, Chen Y, Xie Z, Zhou Y, Peng S, Fan L, et al. Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding. https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding

IEEE

F. Zhang, Y. Chen, Z. Xie, Y. Zhou, S. Peng, L. Fan, X. Ji, C. Zheng, H. Shan, P. S. Yu, X. Liu, Y. Chen, P. Nakov, and S. He, "Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding," https://omanscience.com/ar/articles/same-scores-different-decisions-evaluating-jev-and-language-models-for-legal-document-understanding.