الملخص

Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.

الكلمات المفتاحية

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Alla, S., Sichani, A. S., & Shyu, C. R. (2026). Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement. https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement

MLA 9

Alla, Srikar, et al. "Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement." https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement.

شيكاغو (المؤلف–التاريخ)

Alla, Srikar, Ali Shiri Sichani, and Chi-Ren Shyu. 2026. "Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement." https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement.

هارفارد

Alla, S., Sichani, A. S. and Shyu, C. R. (2026) 'Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement', Available at: https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement.

فانكوفر

Alla S, Sichani AS, Shyu CR. Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement. https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement

IEEE

S. Alla, A. S. Sichani, and C. R. Shyu, "Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement," https://omanscience.com/ar/articles/quantum-entangled-multimodal-fusion-networks-qemfn-resource-aware-hybrid-vision-language-fusion-via-trainable-entanglement.