الباحثون

Zhen Lei

المنشورات 5

نسخة أولية وصول مفتوح

From Suppression to Repair: Mitigating Object Hallucination in Large Vision-Language Models via Localized Distribution Alignment

Chen Zhao, Xingping Dong, Jiachun Shi وآخرون · 2026

Object hallucination remains a major obstacle for large vision-language models (LVLMs) to generate reliable content. An intuitive mitigation strategy is to suppress hallucination-related components in hidden representations. However, these components may also contain useful information, and suppressing them can weaken …

نسخة أولية وصول مفتوح

Geometric Similarity in VLM Low-Level Vision Representations

Shao-Jun Xia, Huixin Zhang, Zhen Lei وآخرون · 2026

Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field …

نسخة أولية وصول مفتوح

VisionHOPE: Visual Backbones as Self-Modifying Learning Systems

Siran Peng, Tianshuo Zhang, Tianyu Fu وآخرون · 2026

Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an image. Across this …

نسخة أولية وصول مفتوح

Object-Centric Conditioning for Visuomotor Flow Matching

Jijie Li, Xu Yang, Junhong Zou وآخرون · 2026

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical mo …

المؤلفون المشاركون