الباحثون

Peng Wang

المنشورات 14

نسخة أولية وصول مفتوح

Transforming Image Editors into Video Editors

Feng Wang, Zijie Li, Ceyuan Yang وآخرون · 2026

Recent image editing systems have achieved impressive semantic understanding, visual fidelity, and instruction-following ability, while video editing remains substantially more difficult and costly. In this paper, we present a simple alternative to end-to-end video editing: instead of training a monolithic video editor …

نسخة أولية وصول مفتوح

LR-V2X: Loss-Resilient Collaborative Perception under Low-Bandwidth Communication

Kang Yang, Tianci Bu, Peng Wang وآخرون · 2026

Given the inherent unpredictability of packet loss in vehicular wireless communications, V2X collaborative perception can yield practical benefits only if agents can achieve reliable collaboration under lossy and low-bandwidth communication conditions. Existing dense BEV feature fusion methods depend on redundant BEV f …

نسخة أولية وصول مفتوح

AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation

Zhenxuan Zeng, Qingle Wu, Wei Suo وآخرون · 2026

Goal-oriented Vision-and-Language Navigation (VLN) requires agents to locate and reach targets described in natural language without prescribed routes. Air--ground collaboration is valuable for tasks requiring both wide-area search and fine-grained localization. However, systematic study of goal-oriented air--ground co …

نسخة أولية وصول مفتوح

UGOD: Uncertainty-Guided Opacity and Dropout for Sparse-View 3D Gaussian Splatting

Zhihao Guo, Peng Wang, Zidong Chen وآخرون · 2026

Sparse-view 3D Gaussian Splatting is prone to overfitting because limited observations leave many Gaussian primitives weakly constrained, yet their contributions are still accumulated through alpha blending. Without uncertainty estimation, the renderer cannot distinguish unreliable primitives from well-constrained ones …

نسخة أولية وصول مفتوح

ElectrolyteFM: Unifying Electrolyte Property Prediction through Cross-Property Knowledge Learning

Jiaxin Yu, Shuo Wang, Peng Wang وآخرون · 2026

Electrolyte formulation design requires balancing multiple physicochemical properties, yet existing models often focus on a limited subset. Learning each property in isolation can overlook transferable chemical information, whereas indiscriminate sharing can introduce cross-property interference. Our directed transfer …

نسخة أولية وصول مفتوح

P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration

Haizhao Jing, Zhenhao Shang, Haokui Zhang وآخرون · 2026

Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence leng …

نسخة أولية وصول مفتوح

SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization

Zhenhao Shang, Haizhao Jing, Haokui Zhang وآخرون · 2026

Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-textual token ratios …

نسخة أولية وصول مفتوح

Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning

Shengchao Hu, Peng Wang, Qiyang Zhou وآخرون · 2026

Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, which weakens action gro …

نسخة أولية وصول مفتوح

CFCH: Coarse-Fine Collaborative Hierarchical Learning for Anterior Segment Disease Analysis

Peng Wang, Haohan Zou, Yanlin Wu وآخرون · 2026

Accurate classification of anterior segment diseases is crucial for ophthalmic screening and diagnosis. However, slit-lamp image analysis remains challenging due to substantial variability in imaging conditions and the intrinsic anatomical-disease hierarchy of ocular pathologies. Existing methods typically formulate th …

نسخة أولية وصول مفتوح

CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels

A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with …

المؤلفون المشاركون