الباحثون

Yifan Yang

المنشورات 7

نسخة أولية وصول مفتوح

Reasoning-Informed Visual Editing

Xue Yang, Peiyuan Zhang, Yilun Zhu وآخرون · 2026

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the …

نسخة أولية وصول مفتوح

Beyond Visual Enhancement: Adaptive Multi-Context Steering to Mitigate LVLM Hallucinations

Shuran Ma, JiaLe Li, Yuxin Dong وآخرون · 2026

Hallucination remains a significant challenge in Large Vision-Language Models (LVLMs). Existing training-free methods generally mitigate hallucinations through contrastive decoding or visual enhancement, often increasing the relative influence of visual evidence during generation. This raises a fundamental question: Ca …

نسخة أولية وصول مفتوح

Structure-aware Keypoint Localization for Videofluoroscopic Swallowing Study

Kai Zhou, Chuanshen Chen, Runhao Zeng وآخرون · 2026

Videofluoroscopic Swallowing Study (VFSS) is one of the gold standard for diagnosing swallowing disorders, providing dynamic X-ray imaging of the swallowing process. Automated kinematic analysis in VFSS relies fundamentally on precise anatomical keypoint localization. However, existing studies focus on limited keypoint …

نسخة أولية وصول مفتوح

Trust the View That Sees the Target: Mining Cross-View Conflicts for Reliability-Gated Disaster Damage Assessment

Yifan Yang · 2026 · 10.1145/3849732.3857333

After a disaster, building damage is assessed from overhead tiles and ground-level photographs, and most methods fuse the two views symmetrically, trusting both equally for every building. This paper focuses on the samples where that assumption fails: the conflict cases, on which two independently trained single-view m …

نسخة أولية وصول مفتوح

SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search

Zetong Zhou, Wentao Zhang, Jingyuan Wang وآخرون · 2026

Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, b …

نسخة أولية وصول مفتوح

Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery

Xu Wang, Yifan Yang, TingHao YU وآخرون · 2026

Sparse autoencoders (SAEs) expose features that help us understand and steer language models, but faithful reconstruction does not guarantee informative concepts. Token-level objectives reward lexical and formatting details alongside semantic content, all competing for a limited sparse budget. We introduce a family of …

نسخة أولية وصول مفتوح

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Qi Chen, Yunfei Chu, Haolin He وآخرون · 2026

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental qu …

المؤلفون المشاركون