الباحثون

Yifan Xing

المنشورات 5

نسخة أولية وصول مفتوح

SPLIT-RL: Staged Perception-Language Reasoning Training with Claim-Level Advantages

Raja Kumar, Rajat Koner, Ritwick Chaudhry وآخرون · 2026

Vision-Language (VL) reasoning requires a model to both extract relevant and accurate information from an image (visual reasoning, VR), and to infer the answer from it (language reasoning, LR). Reinforcement learning with verifiable rewards typically trains both through a single chain-of-thought with a final-answer rew …

نسخة أولية وصول مفتوح

DSV-Mem: Evaluating Multimodal Memory in Professional Workflows for MLLM Agents

Jike Zhong, Ritwick Chaudhry, Xuanbai Chen وآخرون · 2026

Conversational MLLM agents are increasingly expected to assist in professional workflows, from AI research and engineering design to product management and business operations. Yet this capability remains underexplored: existing benchmarks largely focus on informal, everyday interactions and personal-life scenarios fea …

نسخة أولية وصول مفتوح

ALoDLM: Adaptively Looped Diffusion Language Models

Liancheng Fang, Zhuowei Li, Youngeun Kim وآخرون · 2026

Diffusion language models (DLMs) enable fast generation by predicting multiple tokens in parallel, but their practical adoption remains limited by a persistent quality gap relative to comparably sized autoregressive (AR) models. We attribute this gap to a computation-difficulty mismatch: within a partially observed seq …

نسخة أولية وصول مفتوح

VoxelSage: Tool-Augmented 3D CT Analysis and Simulator-Shielded Sequential Resection Planning for Liver Tumors

Binghong Qian, Xuanhe Liu, Yifan Xing وآخرون · 2026

Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dimensional CT volume. Existing tools often handle these steps separately, while language models cannot reliably compute physical measurements from CT. To provide an integra …

نسخة أولية وصول مفتوح

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Xi Xiao, Tianchen Zhao, Youngeun Kim وآخرون · 2026

Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to supervise what latent toke …

المؤلفون المشاركون