الباحثون

Wei Chen

المنشورات 11

نسخة أولية وصول مفتوح

Momentum-Space Path Integral Approach to Non-Hermitian Symmetry Breaking

A quantum-classical correspondence for non-Hermitian symmetry breaking has recently been established using coordinate-space path integrals, providing a semiclassical understanding of spectral transitions at the level of individual eigenstates. Here we develop its dual formulation in momentum space by constructing the c …

نسخة أولية وصول مفتوح

Harnessing Multimodal Large Language Models for Training-Free Human-Object Interaction Detection

Zhaolin Cai, Huiyu Duan, Liu Yang وآخرون · 2026

Human-object interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Traditional supervised methods perform strongly but rely on task-specific training. Recent multimodal large language models (MLLMs) offer a promising route to training-free HOI detection through their broad vi …

نسخة أولية وصول مفتوح

Flash-OPD: Fast On-Policy Distillation

Wei Chen, Junle Chen, Yitong Yang وآخرون · 2026

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision comp …

نسخة أولية وصول مفتوح

You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue

Junle Chen, Wei Chen, Zhengjun Huang وآخرون · 2026

When a large language model handles a multi-turn task and a user proposes a change but ultimately rejects it, the model should continue as if nothing changed. We find a surprising failure: merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged. To systematical …

نسخة أولية وصول مفتوح

Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

Wei Chen, Xuanyu Zheng, Yancheng Long وآخرون · 2026

Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and t …

نسخة أولية وصول مفتوح

SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards

Yingchao Yu, Pengfei Sun, Wenxuan Pan وآخرون · 2026

Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we for …

نسخة أولية وصول مفتوح

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

Tianhang Guo, Yulin He, Wei Chen وآخرون · 2026

Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before l …

نسخة أولية وصول مفتوح

A Support-Enhanced Granular-Jamming Gripper for RL-based Grasping with Continuum Manipulators

Danyu Liu, Tianlin Zhang, Wei Chen وآخرون · 2026

Continuum manipulators provide dexterous motion in confined spaces, but structural compliance, hysteresis, and load-dependent deformation leave residual position and orientation errors that can undermine reliable contact with rigid grippers. To address this limitation, this paper presents a lightweight support-enhanced …

نسخة أولية وصول مفتوح

Imagine then Verify: Affordance-Targeted Active Perception for Task-Oriented Grasping in Cluttered Scenes

Jingzhi Cui, Xuefeng Liu, Feng Han وآخرون · 2026

Task-oriented grasping (TOG) requires robots to grasp functional parts of objects (e.g., the handle of a mug for pouring), yet these affordance regions are frequently occluded in cluttered scenes. Active perception via next-best-view (NBV) planning can resolve such occlusions by moving the camera for more informative o …

المؤلفون المشاركون