Abstract

Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wilcox, B., Gao, D., Kathiravelu, P., Causey, D., Sha, K., & Feng, Y. (2026). Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science. https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science

MLA 9

Wilcox, Benjamin, et al. "Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science." https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science.

Chicago (author–date)

Wilcox, Benjamin, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, and Yunhe Feng. 2026. "Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science." https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science.

Harvard

Wilcox, B., Gao, D., Kathiravelu, P., Causey, D., Sha, K. and Feng, Y. (2026) 'Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science', Available at: https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science.

Vancouver

Wilcox B, Gao D, Kathiravelu P, Causey D, Sha K, Feng Y. Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science. https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science

IEEE

B. Wilcox, D. Gao, P. Kathiravelu, D. Causey, K. Sha, and Y. Feng, "Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science," https://omanscience.com/en/articles/arctic-questions-missing-answers-a-dataset-and-benchmark-for-llm-abstention-in-arctic-science.