الملخص

Jev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev on 13 multiple-choice benchmarks covering knowledge, reasoning, and multilingual understanding, and compare it with 19 LLMs in three tiers: frontier, representative, and small. Jev is competitive with frontier LLMs on knowledge and commonsense benchmarks and obtains the best score on MMLU-Redux and ARC-Challenge. Outside mathematics, it also outperforms most representative LLMs and all small LLMs. However, it falls behind on mathematical word problems: on MathQA, it is 17.7 points below the frontier median and scores lower than all 19 LLMs. These results indicate that a specialized decision model can match general-purpose LLMs on decisions that rely mainly on knowledge, but not on decisions that require multi-step calculation.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Li, X., Chang, Q., Ning, J., Xu, C., Zhang, S., Zhang, Y., Luo, L., & Lin, H. (2026). Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks. https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks

MLA 9

Li, Xing, et al. "Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks." https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks.

شيكاغو (المؤلف–التاريخ)

Li, Xing, Qingcheng Chang, Jinzhong Ning, Changfeng Xu, Shenlong Zhang, Yijia Zhang, Ling Luo, and Hongfei Lin. 2026. "Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks." https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks.

هارفارد

Li, X., Chang, Q., Ning, J., Xu, C., Zhang, S., Zhang, Y., Luo, L. and Lin, H. (2026) 'Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks', Available at: https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks.

فانكوفر

Li X, Chang Q, Ning J, Xu C, Zhang S, Zhang Y, et al. Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks. https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks

IEEE

X. Li, Q. Chang, J. Ning, C. Xu, S. Zhang, Y. Zhang, L. Luo, and H. Lin, "Specialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual Tasks," https://omanscience.com/ar/articles/specialized-decision-models-vs-general-purpose-llms-benchmarking-jev-across-knowledge-reasoning-and-multilingual-tasks.