الباحثون

Yan Huang

المنشورات 7

نسخة أولية وصول مفتوح

EgoPhys: Estimating Peak Contact Force and Mechanical Work from Egocentric Manipulation Video

Zhuo Dong, Jianhua Yang, Haohao Li وآخرون · 2026

Physically grounded manipulation of articulated objects requires understanding both the maximum forces encountered during contact and the work performed as their parts move. Peak contact force and mechanical work quantify these complementary aspects, but estimating them from egocentric video is challenging because phys …

نسخة أولية وصول مفتوح

ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models

Yuan Xu, Yixiang Chen, Qisen Ma وآخرون · 2026

Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the vis …

نسخة أولية وصول مفتوح

AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

Hanzhi Zhang, Qiao Zhang, Qinglei Cao وآخرون · 2026

Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-traini …

نسخة أولية وصول مفتوح

SymNetPro: LOS-Aware Directional Multi-Transmitter Localization from Sparse Radio Observations

Directional multi-transmitter localization from sparse received-power observations is difficult because the receiver observes only the source-unresolved aggregate field: multiple directional sources superpose, building blockage fragments their visible regions, and stronger sources can mask weaker ones. We present SymNe …

نسخة أولية وصول مفتوح

Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment

Chenxiao Li, Yunhe Feng, Dongfang Liu وآخرون · 2026

The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with other modalities to align different entities. Although simple, such strategies overlook the potential noise in the images and their semantic misalignment with corresponding …

المؤلفون المشاركون