الباحثون

Yi Zhang

المنشورات 19

نسخة أولية وصول مفتوح

Attributing HOW, Not Just WHICH: Counterfactual Response Trajectories for Diffusion Models

Haoqian Zhang, Ziyuan Yang, Zerui Shao وآخرون · 2026

Diffusion models have achieved remarkable success in image generation, yet tracing their outputs to individual training examples remains challenging. Existing attribution methods often compress factor-specific effects into scalar responses, making distinct internal changes indistinguishable. This is particularly limiti …

نسخة أولية وصول مفتوح

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge وآخرون · 2026

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac …

نسخة أولية وصول مفتوح

Virtual model control for compliant reaching under uncertainties

Yi Zhang, Daniel Larby, Fumiya Iida وآخرون · 2026

Virtual Model Control (VMC) is an approach to design a controller for force-controlled robots in complex uncertain environments. While this method was primarily investigated for legged robot locomotion in the past, it can be more generally applicable to other types of robotic systems. This paper investigates the VMC fr …

نسخة أولية وصول مفتوح

PWM: Personalized World Models with Online Reinforcement Learning

Zhexin Lou, Guancheng Lu, Zeyu Zhang وآخرون · 2026

Pretrained world models can generate diverse environments, yet users often want to explore a particular scene specified by their own video. This requires learning the scene's visual identity while retaining the quality of action-conditioned generation. We introduce Personalized World Models (PWM), a framework for custo …

نسخة أولية وصول مفتوح

Rethinking What to Cache in Few-Step Diffusion Transformers: Solver-Aware Target Selection

Shuo Yang, Lihao Fang, Yi Zhang وآخرون · 2026

Diffusion Transformers (DiTs) can generate high-quality images and videos, but generating each sample requires multiple costly DiT forward passes. Two common ways to accelerate DiT sampling are step distillation, which reduces the number of sampling steps, and caching, which skips some DiT evaluations by reusing a tens …

نسخة أولية وصول مفتوح

Enhancing Autoregressive Video Generation via Representation Adversarial Distillation

Fangyu Lin, Xingtong Ge, Lunjie Zhu وآخرون · 2026

Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context and can propagate through subsequent rollouts, leading to detail degradation, structural drift, and unstable motion. Existing distribution matching distillation (DMD) prim …

نسخة أولية وصول مفتوح

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Yilun Liu, Yi Zhang, Ganyu Wu وآخرون · 2026

Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed d …

نسخة أولية وصول مفتوح

Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

Weiqiang Wang, Zhuokun Chen, Yusheng Dai وآخرون · 2026

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visua …

نسخة أولية وصول مفتوح

Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

Xingtong Ge, Yutong Wang, Lunjie Zhu وآخرون · 2026

Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Me …

نسخة أولية وصول مفتوح

Learning to Retrieve Missing Evidence for Long-Term Memory QA

Yi-Xuan Deng, Yi Zhang, Wei Liu وآخرون · 2026

Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on …

نسخة أولية وصول مفتوح

From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning

Clinical CT and UHRCT cannot resolve individual trabeculae, whereas synchrotron radiation microCT (SRuCT) provides high-resolution references but is not applicable for in vivo imaging. The two domains differ by a 32x resolution gap, are only coarsely paired, and exhibit severe physical differences including partial vol …

نسخة أولية وصول مفتوح

De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift

Mengyuan Liu, Yuhang Wen, Yi Zhang وآخرون · 2026 · 10.1007/s11263-026-03037-1

Skeleton sequences can represent both individual actions and multi-entity interactions, encompassing human bodies, hands, objects, and robots. Existing approaches to recognize skeleton-based actions and interactions usually adopt a late fusion strategy, which expects individuals are independent and identically distribu …

نسخة أولية وصول مفتوح

SetOPD: From Few Visual Exemplars to Multimodal Candidate Sets for Remote-Sensing Open-Prompt Detection

Jinlong Hu, Yi Zhang, Zhiqi Xia وآخرون · 2026

Open-prompt detectors allow users to specify targets with text, visual exemplars, or both. We argue that existing designs underuse the visual modality in two ways. First, multiple exemplars are commonly compressed into a single class-level embedding. This textualizes visual prompting: the resulting vector plays the rol …

نسخة أولية وصول مفتوح

It's the Geometry, Not the Model: Effective Rank and Subspace Alignment in Functional Connectivity Classification

Xiao Fan, Jingyuan Li, Yubo Han وآخرون · 2026

Resting-state functional connectivity (FC) is widely used to classify brain phenotypes and disorders. Most pipelines use the full connectome and seek gains through model design. We instead examine how FC geometry constrains classification and cross-site transfer. Across-subject FC variation concentrates in a small effe …

نسخة أولية وصول مفتوح

Hi-OPD: Hierarchy-Aware Open-Prompt Detection for Remote Sensing Images

Jinlong Hu, Yi Zhang, Zhiqi Xia وآخرون · 2026

Hi-OPD addresses a failure mode left uncontrolled by flat open-prompt training: descendant retrieval need not persist under ancestor queries when multi-source remote sensing annotations exhibit inconsistent granularity and missing labels. A detector may localize \textit{car} and \textit{van} under atomic prompts yet mi …

نسخة أولية وصول مفتوح

Transferring the Intelligence of VLMs to Robotic Control

Meng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang وآخرون · 2026

Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly g …

المؤلفون المشاركون