الباحثون

Xiao Yang

المنشورات 5

نسخة أولية وصول مفتوح

FlashGaze: Training-Free Multi-Scale Patch Pruning For Efficient Video Understanding

Ziye Zhu, Yanghao Zhou, Lixing Tan وآخرون · 2026

Multimodal Large Language Models (MLLMs) have demonstrated strong performance in video understanding, yet efficiently processing long, high-resolution videos remains challenging. Such videos often contain substantial spatiotemporal redundancy, and processing redundant visual tokens can incur avoidable computational ove …

نسخة أولية وصول مفتوح

IdeaScientist: Orchestrating Agents for Grounded Scientific Ideation

Jiarui Liu, Renjie Tao, Yiwei Liao وآخرون · 2026

Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an an …

نسخة أولية وصول مفتوح

MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

Guangzhi Xiong, Xinyuan Zhang, Xiao Yang وآخرون · 2026

Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text …

نسخة أولية وصول مفتوح

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Yu Xu, Yuxin Zhang, Xiao Yang وآخرون · 2026

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is s …

نسخة أولية وصول مفتوح

CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents

Xiao Yang, Yangchen Ou, Yuhan Gao وآخرون · 2026

Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically …

المؤلفون المشاركون