الملخص
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Bondielli, A., Passaro, L., Bacciu, D., & Lenci, A. (2026). Likelihood Ranking doesn't Scale Like Prompting in LLMs. https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms
MLA 9
Bondielli, Alessandro, et al. "Likelihood Ranking doesn't Scale Like Prompting in LLMs." https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.
شيكاغو (المؤلف–التاريخ)
Bondielli, Alessandro, Lucia Passaro, Davide Bacciu, and Alessandro Lenci. 2026. "Likelihood Ranking doesn't Scale Like Prompting in LLMs." https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.
هارفارد
Bondielli, A., Passaro, L., Bacciu, D. and Lenci, A. (2026) 'Likelihood Ranking doesn't Scale Like Prompting in LLMs', Available at: https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.
فانكوفر
Bondielli A, Passaro L, Bacciu D, Lenci A. Likelihood Ranking doesn't Scale Like Prompting in LLMs. https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms
IEEE
A. Bondielli, L. Passaro, D. Bacciu, and A. Lenci, "Likelihood Ranking doesn't Scale Like Prompting in LLMs," https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.