الباحثون

Xu Yang

المنشورات 10

نسخة أولية وصول مفتوح

Reasoning-Informed Visual Editing

Xue Yang, Peiyuan Zhang, Yilun Zhu وآخرون · 2026

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the …

نسخة أولية وصول مفتوح

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Hanmo Chen, Chengcheng Liu, Tianxiao Chen وآخرون · 2026

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sou …

نسخة أولية وصول مفتوح

Proprioceptive Force Estimation for Quadruped Locomotion and Human-Robot Interaction

Run Wang, Xu Yang, Alapati Tuerxun وآخرون · 2026

Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy are jointly trained …

نسخة أولية وصول مفتوح

Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning

Yihan Zhou, Rui Yan, Mingcong Li وآخرون · 2026

Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleoperation to provide fram …

نسخة أولية وصول مفتوح

TSGate: Timestep-Aware Gated Attention for Diffusion Transformers

Boyu Zhang, Yifan Liu, Shuxia Lin وآخرون · 2026

Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-of-domain (OOD) prom …

نسخة أولية وصول مفتوح

PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders

Chenduo Hao, Chuanbao Gao, Pinjun Zeng وآخرون · 2026

In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce PULSE (Paired Utili …

نسخة أولية وصول مفتوح

Object-Centric Conditioning for Visuomotor Flow Matching

Jijie Li, Xu Yang, Junhong Zou وآخرون · 2026

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical mo …

نسخة أولية وصول مفتوح

CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents

Tao Huang, Guosen Wu, Guolong Zheng وآخرون · 2026

Privacy leakage in LLM agents is commonly evaluated within individual components such as memory, retrieval, or tool-use pipelines, which makes it difficult to distinguish internal exposure from information that an external observer can actually recover. We present CIPL (Channel Inversion for Privacy Leakage), a channel …

نسخة أولية وصول مفتوح

Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

Yutong Yao, Yanjie Cao, Guanhua Chen وآخرون · 2026

Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely …

المؤلفون المشاركون