الباحثون

Lei Zhang

المنشورات 16

نسخة أولية وصول مفتوح

Outage-Aware Robust Sensing for UAV ISAC Systems

Qiming Li, Lei Zhang, Haoran Xu وآخرون · 2026

Integrated sensing and communications (ISAC) infrastructure can provide an external perception layer for aerial robots, but dynamic blockage makes the set of informative sensing links at the next update uncertain. A single predicted availability pattern can miss low-information outcomes, whereas protecting all patterns …

نسخة أولية وصول مفتوح

UniData: Universal Multimodal Instruction Generation Pipeline

Jiaqi Tang, Yi-Feng Wu, Yuting Zhang وآخرون · 2026

Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they …

نسخة أولية وصول مفتوح

HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

Xiangtao Kong, Shuaizheng Liu, Rongyuan Wu وآخرون · 2026

Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models. This paradi …

نسخة أولية وصول مفتوح

Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation

Zhen Guo, Rongyuan Wu, Qiaosi Yi وآخرون · 2026

Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challengin …

نسخة أولية وصول مفتوح

MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

Hang Wang, Chao Shen, Lei Zhang وآخرون · 2026

The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, …

نسخة أولية وصول مفتوح

MMPostTrainBench: Benchmarking Autonomous Research for Multimodal Post-Training

Yuxin Liu, Yuxuan Wang, Zhenxin Lei وآخرون · 2026

Autonomous research seeks sustained model improvements through iterative experimentation and feedback. LLM agents show promise in automating machine learning and language-model post-training, but their ability to sustain multimodal improvement remains unclear. We introduce MMPostTrainBench, a benchmark spanning eight t …

نسخة أولية وصول مفتوح

On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection

Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited training data. We compare a …

نسخة أولية وصول مفتوح

Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization

Wei Zhao, Yangshuo Zou, Chengxiang Ding وآخرون · 2026

We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent superv …

نسخة أولية وصول مفتوح

LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation

Zhengqiang Zhang, Lingchen Sun, Rongyuan Wu وآخرون · 2026

Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rat …

نسخة أولية وصول مفتوح

OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization

Zhiyuan Ma, Wenbo Hu, Wang Zhao وآخرون · 2026

Despite recent advancements in 3D generation, models often struggle to produce assets with high visual fidelity. To bridge this gap, we propose OREO, an alignment framework that enhances the realism of 3D generators by leveraging rich 2D diffusion priors. Instead of relying on static datasets, OREO establishes a dynami …

نسخة أولية وصول مفتوح

Less Is More in the Long Tail: Stage-Adaptive Sample Selection for Annotation-Efficient Dense Prediction

Xiaofei Du, Lei Zhang, Shuyu Yan وآخرون · 2026

Deep learning performance generally improves with increasing training data, yet this scaling is fundamentally constrained by annotation cost in large-scale dense prediction tasks with long-tailed category distributions, where pixel- or voxel-level annotation is prohibitively expensive. We propose SASS (Stage-Adaptive S …

نسخة أولية وصول مفتوح

General Collaborative Intelligence: Architecting Cognition for Resilient Multi-Agent Ecosystems

Lei Zhang, Chun Ye, Le Yang وآخرون · 2026

Multi-agent unmanned systems are moving from isolated, ego-centric sensing toward collaborative intelligence, in which distributed agents exchange compact features to overcome a local observation trap that no single agent can escape: occlusions, finite sensor range, and environmental degradation. The field has matured …

نسخة أولية وصول مفتوح

ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

Ju Dong, Yu Fu, Jian Chen وآخرون · 2026

Force-aware Vision-Language-Action (VLA) policies improve contact-rich manipulation, but typically combine task-level motion and contact-dependent adjustment in a single action prediction. Demonstrations provide no explicit labels for decomposing that prediction into a reusable reference action and a correction. We pre …

المؤلفون المشاركون