الباحثون

Xue Yang

المنشورات 8

نسخة أولية وصول مفتوح

IntactWorld: Joint World Modeling with Intact Features

Boming Tan, Xiangdong Zhang, Yan Xia وآخرون · 2026

While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning proc …

نسخة أولية وصول مفتوح

Reasoning-Informed Visual Editing

Xue Yang, Peiyuan Zhang, Yilun Zhu وآخرون · 2026

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the …

نسخة أولية وصول مفتوح

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Shuran Ma, JiaLe Li, Yuxin Dong وآخرون · 2026

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Ca …

نسخة أولية وصول مفتوح

MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Xilin Dang, Weilin Ruan, Xue Yang وآخرون · 2026

Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation …

نسخة أولية وصول مفتوح

Quantum Fidelity Landscape-Guided Prior Calibration for Single-Circuit QGAN Image Generation

Xue Yang, Rigui Zhou, Dax Enshan Koh وآخرون · 2026

Quantum Generative Adversarial Networks (QGANs) have emerged as representative generative models in the Noisy Intermediate-Scale Quantum (NISQ) era and have attracted increasing attention in quantum machine learning. However, most existing QGAN methods rely on patch-based decomposition strategies, which weaken the glob …

نسخة أولية وصول مفتوح

LoopVL: Recurrent Visual Intelligence

Zhe Qian, Ziyang Gong, Zhongxing Xu وآخرون · 2026

We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal train …

نسخة أولية وصول مفتوح

Bench2Dex: Benchmarking Visuo-Tactile Bimanual Dexterous Manipulation Across Dexterous Hands

Zhenjie Yang, Yideng Zhang, Dongjie Zhang وآخرون · 2026

Tactile sensing provides contact information that can be difficult to infer from vision alone, but tactile hardware for dexterous hands has not converged to a common design. Dexterous hands differ in finger structure, contact surfaces, and sensor layouts, while simulated tactile signals still differ from measurements p …

المؤلفون المشاركون