Abstract

In this paper, we provide an overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) task. Building on the success of the NTCIR-18 core task AEOLLM, we proposed AEOLLM-2 for NTCIR-19 to further investigate automatic evaluation methods for Large Language Models (LLMs), particularly in long-form text generation scenarios. In AEOLLM-2, we introduced a new subtask, Deep Research Evaluation, which focuses on the automatic evaluation of long-form deep research reports generated by LLMs. Participants developed evaluation methods to automatically assess the quality of these reports, and the performance of each method was measured by comparing its scores against human-annotated ground-truth labels. This year, we received 91 runs from 10 teams in total. This paper describes the background of the task, the dataset construction, the evaluation measures, the participants' methods, and the final evaluation results.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chen, J., Dong, Y., Li, H., Liu, Y., & Ai, Q. (2026). Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task. https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task

MLA 9

Chen, Junjie, et al. "Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task." https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task.

Chicago (author–date)

Chen, Junjie, Yuxi Dong, Haitao Li, Yiqun Liu, and Qingyao Ai. 2026. "Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task." https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task.

Harvard

Chen, J., Dong, Y., Li, H., Liu, Y. and Ai, Q. (2026) 'Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task', Available at: https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task.

Vancouver

Chen J, Dong Y, Li H, Liu Y, Ai Q. Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task. https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task

IEEE

J. Chen, Y. Dong, H. Li, Y. Liu, and Q. Ai, "Overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) Task," https://omanscience.com/en/articles/overview-of-the-ntcir-19-automatic-evaluation-of-llms-2-aeollm-2-task.