الباحثون

Heng Fan

المنشورات 8

نسخة أولية وصول مفتوح

AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

Hanzhi Zhang, Qiao Zhang, Qinglei Cao وآخرون · 2026

Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-traini …

نسخة أولية وصول مفتوح

ArticuTable: Generating Instance-Level Interactive Rigid-Articulated 3D Tabletop Scenes from a Single Image

Kai Lv, Yibo Yin, Lijun Guo وآخرون · 2026

Embodied agents benefit from 3D environments that combine visual fidelity to real-world observations with physical interactivity. Existing single-image tabletop reconstruction methods recover plausible scene geometry but typically represent objects as monolithic rigid bodies, limiting interaction to whole-object rigid …

نسخة أولية وصول مفتوح

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

Liang Peng, Shizhuo Mu, Bohan Tan وآخرون · 2026

3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search …

نسخة أولية وصول مفتوح

SymNetPro: LOS-Aware Directional Multi-Transmitter Localization from Sparse Radio Observations

Directional multi-transmitter localization from sparse received-power observations is difficult because the receiver observes only the source-unresolved aggregate field: multiple directional sources superpose, building blockage fragments their visible regions, and stronger sources can mask weaker ones. We present SymNe …

نسخة أولية وصول مفتوح

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Shaohua Dong, Zexuan Meng, Haiyan Sun وآخرون · 2026

In this paper, we propose RGBD20K, a novel dataset for facilitating the development of more robust and general RGB-D semantic segmentation by encompassing abundant categories and high-quality annotations. RGBD20K possesses several attractive properties: (1) Expanded Semantic Space. In particular, it covers 160 fine-gra …

نسخة أولية وصول مفتوح

Knowing When to Trust Images: Reliability-Aware Multi-modal Entity Alignment

Chenxiao Li, Yunhe Feng, Dongfang Liu وآخرون · 2026

The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with other modalities to align different entities. Although simple, such strategies overlook the potential noise in the images and their semantic misalignment with corresponding …

المؤلفون المشاركون