Abstract

Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensional evaluation on IEMOCAP, incorporating acoustic cues as natural language descriptions following the SpeechCueLLM approach. We evaluate six models spanning the LLaMA, GPT, and Qwen families under zero-shot prompting, few-shot prompting, and LoRA fine-tuning. LoRA fine-tuned LLaMA models substantially outperform prompt-engineered GPT models on both tasks despite GPT's larger scale, a gap we attribute to domain adaptation rather than model capacity. Our best model achieves a Valence CCC of 0.7822, a new state-of-the-art on IEMOCAP. Ablation studies confirm that textual audio descriptions meaningfully improve smaller models (+3.5 to 3.6 weighted F1) while contributing little for the largest model, suggesting audio cues are most valuable when linguistic capacity is limited. The performance asymmetry across VAD dimensions closely mirrors the annotator agreement hierarchy in IEMOCAP's own annotations.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hu, Y., & Choi, J. (2026). Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue. https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue

MLA 9

Hu, Yutong, and Jinho Choi. "Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue." https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue.

Chicago (author–date)

Hu, Yutong, and Jinho Choi. 2026. "Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue." https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue.

Harvard

Hu, Y. and Choi, J. (2026) 'Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue', Available at: https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue.

Vancouver

Hu Y, Choi J. Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue. https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue

IEEE

Y. Hu, and J. Choi, "Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue," https://omanscience.com/en/articles/beyond-text-llm-based-dimensional-emotion-evaluation-in-multimodal-dialogue.