الباحثون

Yang Luo

المنشورات 3

نسخة أولية وصول مفتوح

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Jianfei Zhao, Yifan Wang, Feng Zhang وآخرون · 2026

Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever ( …

نسخة أولية وصول مفتوح

VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction

Qiutong Chen, Yuchan Guo, Zhenlong Yuan وآخرون · 2026

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in video understanding, yet their reliance on retrospective summarization and text-centric priors often limits their ability to bridge unobserved causal transitions when applied to Video Event Prediction (VEP). To address this, we propose V …

نسخة أولية وصول مفتوح

S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection

Qiangqiang Zhou, Yang Luo, Yong Chen وآخرون · 2026

Alignment-free RGB-T salient object detection (RGB-T SOD) aims to identify salient objects from unregistered RGB and thermal image pairs without costly pre-alignment. However, spatial misalignment breaks pixel-wise correspondence and causes feature contamination during cross-modal fusion. To address this issue, we prop …

المؤلفون المشاركون