Abstract

In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Nguyen, C. P., Nguyen, H. T., & Le, T. (2026). Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering. https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering

MLA 9

Nguyen, Cong Phu, et al. "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering." https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.

Chicago (author–date)

Nguyen, Cong Phu, Huy Tien Nguyen, and Tung Le. 2026. "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering." https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.

Harvard

Nguyen, C. P., Nguyen, H. T. and Le, T. (2026) 'Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering', Available at: https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.

Vancouver

Nguyen CP, Nguyen HT, Le T. Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering. https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering

IEEE

C. P. Nguyen, H. T. Nguyen, and T. Le, "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering," https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.