الباحثون

Jun Zhang

المنشورات 17

نسخة أولية وصول مفتوح

Learning Which Correspondences to Trust: Confidence-Weighted Event-Camera Localization in LiDAR Maps

Panagiotis Kiousis, Kuangyi Chen, Jun Zhang وآخرون · 2026

Localizing an event camera against a pre-built LiDAR map can be cast as dense optical-flow estimation between a rendered depth view and an event image, followed by a Perspective-n-Point (PnP) solver over the induced 3D-2D correspondences. Existing pipelines rely on geometric consensus during pose estimation, but do not …

نسخة أولية وصول مفتوح

Equal Path Cost, Unequal Output Effects: Understanding Perturbation Propagation in Diffusion Models

Wei Guo, Yaowen Zhang, Xingtong Ge وآخرون · 2026

Diffusion models have achieved remarkable success in generative modeling, with their sampling procedures routinely modified to control generation and improve efficiency. These modifications introduce perturbations along the sampling trajectory, raising a central question: how do such perturbations affect generated outp …

نسخة أولية وصول مفتوح

AstraSR: Real-World Thermal Super-Resolution with GPT-6 Astra

Mengyuan Li, Changhong Fu, Jun Zhang وآخرون · 2026

Real-world thermal super-resolution (SR) is constrained by limited sensor resolution and the difficulty of obtaining corresponding high-resolution (HR) observations for direct model supervision. Conventional SR methods typically construct training pairs by treating captured thermal images with real-world degradations a …

نسخة أولية وصول مفتوح

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang وآخرون · 2026

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording …

نسخة أولية وصول مفتوح

MegaAvatar: Controllable Talking Avatar Generation

Junyao Gao, Sibo Liu, Weidong Zhang وآخرون · 2026

This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with previous talking-avatar methods that mainly rely on audio or reference-image conditioning, we introduce additional SMPL-X-derived 3D guidance, enabling global control over …

نسخة أولية وصول مفتوح

Enhancing Autoregressive Video Generation via Representation Adversarial Distillation

Fangyu Lin, Xingtong Ge, Lunjie Zhu وآخرون · 2026

Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) prim …

نسخة أولية وصول مفتوح

AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

Yijie Bian, Kai Zhang, Wei Guo وآخرون · 2026

Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scal …

نسخة أولية وصول مفتوح

Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

Xingtong Ge, Yutong Wang, Lunjie Zhu وآخرون · 2026

Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Me …

نسخة أولية وصول مفتوح

Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation

Chenjian Gao, Zhihao Hu, Jianqi Ma وآخرون · 2026

Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rol …

نسخة أولية وصول مفتوح

WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon

Haiyu Zhang, Wenqiang Sun, Tengfei Wang وآخرون · 2026

Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this pap …

نسخة أولية وصول مفتوح

GPARA: Graph-Posterior-Aligned Refinement and Active Acquisition for Grounding Diffusion Priors

Wangqian Chen, Hao Wang, Yumeng Zhang وآخرون · 2026

Active grounding of a frozen diffusion prior requires jointly determining where new measurements should be taken and how they should be used to refine the current reconstruction. Posterior-ensemble-based methods can estimate acquisition utility from generated samples, but require repeated ensemble generation as observa …

نسخة أولية وصول مفتوح

ReplayLens: Auditing Agents' Use of Outcomes

Dong Xu, Zhangfan Yang, Jiantao Wu وآخرون · 2026

When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship in the stored histo …

نسخة أولية وصول مفتوح

Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

Dong Xu, Zhangfan Yang, Jiantao Wu وآخرون · 2026

Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure th …

نسخة أولية وصول مفتوح

Learning to Optimize UAV Path Planning for Data Sensing in Wireless Sensor Networks

Sijie Ma, Zeyuan Ma, Weijia Cao وآخرون · 2026

UAVs have emerged as highly flexible platforms for data sensing in Wireless Sensor Networks (WSNs). Path planning for UAVs in such tasks plays a key role to assure remote sensing effectiveness and friendly energy consumption. However, existing approaches show two key limitations: i) they are primarily hand-crafted with …

نسخة أولية وصول مفتوح

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

ZhuoXin Liu, Zhiming Ma, Ying Zhang وآخرون · 2026

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages s …

المؤلفون المشاركون