Abstract
In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often benchmarked with open-domain question answering (QA) datasets containing questions and corresponding short reference answers. First, an LLM is used to generate answers to questions within the QA dataset. Then, some automated labeling strategy is used to label these answers as hallucinated or not by comparing them with the reference answers in the dataset. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness (whether the answer is fully supported by the reference) and factual correctness (whether the answer is free from contradictions and factually false specific claims). In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question-answer pairs spanning three commonly used QA datasets and three generator models, with labels targeting answer-level factual correctness. We evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants as automated labelers. Our experiments reveal substantial disagreement both among automated labeling strategies and between these labels and human annotations. Many strategies also exhibit strong directional error biases, and for most judge-generator pairs, replacing a faithfulness-oriented prompt with a factual-correctness prompt improves agreement with human annotations and reduces false-positive dominance, indicating that automated hallucination labels depend strongly on how the target criterion is specified. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Valjakka, J., Kivimäki, J., Mylläri, J., & Nurminen, J. K. (2026). The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation. https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation
MLA 9
Valjakka, Jorma, et al. "The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation." https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation.
Chicago (author–date)
Valjakka, Jorma, Juhani Kivimäki, Juha Mylläri, and Jukka K. Nurminen. 2026. "The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation." https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation.
Harvard
Valjakka, J., Kivimäki, J., Mylläri, J. and Nurminen, J. K. (2026) 'The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation', Available at: https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation.
Vancouver
Valjakka J, Kivimäki J, Mylläri J, Nurminen JK. The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation. https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation
IEEE
J. Valjakka, J. Kivimäki, J. Mylläri, and J. K. Nurminen, "The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation," https://omanscience.com/en/articles/the-labeling-problem-in-hallucination-detection-benchmarks-an-empirical-evaluation.