الباحثون

Yan Li

المنشورات 4

نسخة أولية وصول مفتوح

Reasoning-Informed Visual Editing

Xue Yang, Peiyuan Zhang, Yilun Zhu وآخرون · 2026

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the …

نسخة أولية وصول مفتوح

RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts

The genome holds the blueprint that governs the biological properties of the cell. Consequently, advancing our knowledge of genomic function is crucial both for a broader understanding of biology and for continued biomedical advances. The success of large language models on natural language and protein sequences has mo …

نسخة أولية وصول مفتوح

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Zelong Xu, Yan Li, Wenhe Hu وآخرون · 2026

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in …

نسخة أولية وصول مفتوح

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai وآخرون · 2026

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be …

المؤلفون المشاركون