Abstract

This article proposes Selective Affective Layer Fine-Tuning (SALFT), an efficient adaptation framework for Video Vision Transformers in player arousal recognition from gameplay. To bypass computationally expensive full fine-tuning, SALFT introduces a selection criterion based on the L2-norm change in layer parameters after brief adaptation, directly measuring representational shifts and providing a more stable basis than gradient-based alternatives. Evaluated via five-fold cross-validation on the Arousal Video Game AnnotatIoN dataset, SALFT achieves performance comparable to full fine-tuning across all games without statistically significant degradation ($p>0.05$), while updating only $\approx$8% of parameters (over 92% reduction). Notably, in one game, SALFT consistently outperforms both full fine-tuning and the best baseline across all metrics and folds, reaching the theoretical minimum p-value (p=0.0625, exact two-sided Wilcoxon signed-rank test). In addition, we introduce an interpretability method to trace attention patterns, enhancing model transparency. These results establish SALFT as an effective and efficient approach for affective game computing.

Keywords

Subject

Publication details

DOI
10.1109/tg.2026.3691772
Journal
Not available
Open access
Green open access

Cite this article

APA 7

Xia, Y., Khan, I., Dewantoro, M. F., Ouyang, W., & Thawonmas, R. (2026). Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage. https://doi.org/10.1109/tg.2026.3691772

MLA 9

Xia, Yi, et al. "Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage." https://doi.org/10.1109/tg.2026.3691772.

Chicago (author–date)

Xia, Yi, Ibrahim Khan, Mury Fajar Dewantoro, Wenwen Ouyang, and Ruck Thawonmas. 2026. "Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage." https://doi.org/10.1109/tg.2026.3691772.

Harvard

Xia, Y., Khan, I., Dewantoro, M. F., Ouyang, W. and Thawonmas, R. (2026) 'Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage', doi:10.1109/tg.2026.3691772.

Vancouver

Xia Y, Khan I, Dewantoro MF, Ouyang W, Thawonmas R. Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage. doi:10.1109/tg.2026.3691772

IEEE

Y. Xia, I. Khan, M. F. Dewantoro, W. Ouyang, and R. Thawonmas, "Emoception: Selective Affective Layer Fine-Tuning of Video Vision Transformers for Player Arousal Change Recognition From Gameplay Footage," doi: 10.1109/tg.2026.3691772.