الباحثون

Yifei Wang

المنشورات 9

نسخة أولية وصول مفتوح

OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation

Yueqi Wang, Zitian Guo, Yupeng Hou وآخرون · 2026

Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching …

نسخة أولية وصول مفتوح

Visualizing Distribution Coverage in Generative Diffusion Models

Yifei Wang, Xiaoyu Wu, Tsu-Jui Fu وآخرون · 2026

Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses …

نسخة أولية وصول مفتوح

Reliability Testing of Medical Model Performance under Distributed Deployment

Yifei Wang, Xiaohan Zhang, Youtao Ding وآخرون · 2026

Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to …

نسخة أولية وصول مفتوح

Scaling Video Generation for Reasoning: At What Cost?

Weihang Guo, Xiaoyu Wu, Yifei Wang وآخرون · 2026

We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inf …

نسخة أولية وصول مفتوح

What Paired Evaluations Reveal under Visual Perturbations

Yongda Wei, Chen Zhang, Yifei Wang وآخرون · 2026

Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggregate accuracy. We i …

نسخة أولية وصول مفتوح

Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers

Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semant …

نسخة أولية وصول مفتوح

Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

Xiaoyu Wu, Weihang Guo, Yifei Wang وآخرون · 2026

Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selecte …

نسخة أولية وصول مفتوح

Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents

Donghua Cai, Yongheng Deng, Yifei Wang وآخرون · 2026

Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While effective in control …

نسخة أولية وصول مفتوح

Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation

Jiaqi Zhai, Jingkai Zhao, Chen Yang وآخرون · 2026

Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geometry-grounded coordinate for steerable and force-responsive manipulation. Each arm owns thirteen signed translation, rotation, and hold atoms gro …

المؤلفون المشاركون