الملخص

LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Bondielli, A., Passaro, L., Bacciu, D., & Lenci, A. (2026). Likelihood Ranking doesn't Scale Like Prompting in LLMs. https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms

MLA 9

Bondielli, Alessandro, et al. "Likelihood Ranking doesn't Scale Like Prompting in LLMs." https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.

شيكاغو (المؤلف–التاريخ)

Bondielli, Alessandro, Lucia Passaro, Davide Bacciu, and Alessandro Lenci. 2026. "Likelihood Ranking doesn't Scale Like Prompting in LLMs." https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.

هارفارد

Bondielli, A., Passaro, L., Bacciu, D. and Lenci, A. (2026) 'Likelihood Ranking doesn't Scale Like Prompting in LLMs', Available at: https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.

فانكوفر

Bondielli A, Passaro L, Bacciu D, Lenci A. Likelihood Ranking doesn't Scale Like Prompting in LLMs. https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms

IEEE

A. Bondielli, L. Passaro, D. Bacciu, and A. Lenci, "Likelihood Ranking doesn't Scale Like Prompting in LLMs," https://omanscience.com/ar/articles/likelihood-ranking-doesn-t-scale-like-prompting-in-llms.