الباحثون

Wenhao Li

المنشورات 11

نسخة أولية وصول مفتوح

DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

Wenhao Li, Xianjing Meng, Qiangchang Wang وآخرون · 2026

Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making …

نسخة أولية وصول مفتوح

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

Wenhao Li, Xiu Su, Yu Han وآخرون · 2026

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack tempora …

نسخة أولية وصول مفتوح

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

Yiheng Lyu, Xueying Jiang, Wenhao Li وآخرون · 2026

How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, exte …

نسخة أولية وصول مفتوح

FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

Pu Wang, Yongcong Wang, Wenhao Li وآخرون · 2026

Existing video deraining methods typically exploit neighboring frames through either explicit alignment or implicit spatiotemporal aggregation. Explicit alignment relies on accurate motion estimation, which can become unreliable under dense rain, while implicit aggregation avoids alignment but lacks explicit guidance o …

نسخة أولية وصول مفتوح

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Wenhao Li, Zhibin Wu, Chong Xiao وآخرون · 2026

Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, …

نسخة أولية وصول مفتوح

Safety-Filtered Distributed Koopman-MPC

Shengjun Zhang, Wenhao Li, Zhenxin Lin وآخرون · 2026

Distributed model predictive control (DMPC) often constructs both predictions and collision constraints from neighbor trajectories, so packet loss can remove both. We separate these roles: received trajectories drive Koopman-MPC, while local sensing and shelf geometry define a hard-constrained quadratic program (QP) th …

نسخة أولية وصول مفتوح

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

Peidong Liu, Zhiyuan Xiang, Mingyang Li وآخرون · 2026

Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data …

المؤلفون المشاركون