Abstract
Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Iqbal, F. S., Dutt, R., Das, S., Verma, A., & Choudhury, S. R. (2026). Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks. https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks
MLA 9
Iqbal, Fahmid Shahriar, et al. "Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks." https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks.
Chicago (author–date)
Iqbal, Fahmid Shahriar, Ritam Dutt, Soumitra Das, Arnav Verma, and Sagnik Ray Choudhury. 2026. "Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks." https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks.
Harvard
Iqbal, F. S., Dutt, R., Das, S., Verma, A. and Choudhury, S. R. (2026) 'Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks', Available at: https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks.
Vancouver
Iqbal FS, Dutt R, Das S, Verma A, Choudhury SR. Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks. https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks
IEEE
F. S. Iqbal, R. Dutt, S. Das, A. Verma, and S. R. Choudhury, "Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks," https://omanscience.com/en/articles/evaluating-multi-dimensional-generalization-of-large-language-models-in-temporal-extraction-tasks.