الباحثون

Shuo Xing

المنشورات 3

نسخة أولية وصول مفتوح

OpticalRec: Unified Optical Vision-Language Representation for Multimodal Recommendation

Yueqi Wang, Zitian Guo, Yupeng Hou وآخرون · 2026

Recent advances in vision-language modeling have substantially improved multimodal encoding, retrieval and reasoning. Yet for multimodal recommendation, encoding rich item vision-language semantic interactions remains a long-standing bottleneck, which hampers accurate item representation learning and user-item matching …

نسخة أولية وصول مفتوح

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Shuo Xing, Zilin Dai, Chengyuan Qian وآخرون · 2026

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLM …

نسخة أولية وصول مفتوح

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Shuo Xing, Pooja Verlani, Balu Adsumilli وآخرون · 2026

Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily foc …

المؤلفون المشاركون