Abstract
In recent decades, artificial intelligence has made significant progress in understanding and interacting with images. One of the important applications of this technology is Visual Question Answering (VQA), a research field that requires computers to understand and answer questions about images in a natural manner. Despite extensive research and development in VQA for English, there have been very few similar efforts made for other languages, especially Vietnamese. This gap presents a significant challenge and opportunity for the advancement of VQA technology in the Vietnamese language context. By bridging this gap, the field of Vietnamese VQA not only enriches the diversity of research in artificial intelligence but also enables practical applications in various domains, such as education, healthcare, and entertainment, catering to Vietnamese-speaking populations worldwide. Thus, the exploration and development of Vietnamese VQA systems hold immense potential for advancing both research and practical applications in the intersection of computer vision and natural language processing. In this paper, we propose a Multi-layer Fusing Transformer model utilizing a cross attention module to combine multiple modality features of images and texts from different layers in an aggregated representation. Our architecture allows us extract information from low level to high level. Through detailed experiments and ablation studies, our model achieves promising results against the competitive baselines in ViVQA dataset for Vietnamese language.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Nguyen, C. P., Nguyen, H. T., & Le, T. (2026). Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering. https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering
MLA 9
Nguyen, Cong Phu, et al. "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering." https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.
Chicago (author–date)
Nguyen, Cong Phu, Huy Tien Nguyen, and Tung Le. 2026. "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering." https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.
Harvard
Nguyen, C. P., Nguyen, H. T. and Le, T. (2026) 'Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering', Available at: https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.
Vancouver
Nguyen CP, Nguyen HT, Le T. Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering. https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering
IEEE
C. P. Nguyen, H. T. Nguyen, and T. Le, "Fusing Visual and Textual Representations via Multi-layer Fusing Transformers for Vietnamese Visual Question Answering," https://omanscience.com/en/articles/fusing-visual-and-textual-representations-via-multi-layer-fusing-transformers-for-vietnamese-visual-question-answering.