Abstract
In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chehbouni, K., Medjdoub, M., Carichon, F., Farnadi, G., & Cheung, J. C. K. (2026). LLJ Cards: Best practices for the Use of LLMs as Judges. https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges
MLA 9
Chehbouni, Khaoula, et al. "LLJ Cards: Best practices for the Use of LLMs as Judges." https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges.
Chicago (author–date)
Chehbouni, Khaoula, Melina Medjdoub, Florian Carichon, Golnoosh Farnadi, and Jackie Chi Kit Cheung. 2026. "LLJ Cards: Best practices for the Use of LLMs as Judges." https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges.
Harvard
Chehbouni, K., Medjdoub, M., Carichon, F., Farnadi, G. and Cheung, J. C. K. (2026) 'LLJ Cards: Best practices for the Use of LLMs as Judges', Available at: https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges.
Vancouver
Chehbouni K, Medjdoub M, Carichon F, Farnadi G, Cheung JCK. LLJ Cards: Best practices for the Use of LLMs as Judges. https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges
IEEE
K. Chehbouni, M. Medjdoub, F. Carichon, G. Farnadi, and J. C. K. Cheung, "LLJ Cards: Best practices for the Use of LLMs as Judges," https://omanscience.com/en/articles/llj-cards-best-practices-for-the-use-of-llms-as-judges.