الباحثون

Wei Xu

المنشورات 8

نسخة أولية وصول مفتوح

TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs

Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP …

نسخة أولية وصول مفتوح

Efficient Neural Field Learning via Adaptive Coverage and Focused Sampling

Guang Zhao, Xihaier Luo, Huan-Hsin Tseng وآخرون · 2026

Implicit neural representations (INRs) provide a flexible framework for modeling high-dimensional continuous fields, but their training is often inefficient due to uniform subsampling that ignores spatial heterogeneity. Existing adaptive sampling methods partially address this issue by prioritizing high-error samples, …

نسخة أولية وصول مفتوح

Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

Xianghan Wei, Xiaoda Yang, Zhi Wang وآخرون · 2026

While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free appro …

نسخة أولية وصول مفتوح

Beyond Oracle Communication: Benchmarking Interactive Intent Alignment Under Miscommunication and Evolving User Intent

Zheyuan Zhang, Mengyuan Chao, Ke Xiao وآخرون · 2026

Modern LLM agents increasingly tackle complex tasks through interactive, long-horizon exchanges with users, while existing benchmarks generally assume that users always accurately and sufficiently communicate a fixed intent. However, this oracle communication assumption rarely holds in practice: users may miscommunicat …

نسخة أولية وصول مفتوح

Humanoid Loco-Manipulation With Discrete VLA Model

Wenxin Shao, Siqi Chai, Kun Li وآخرون · 2026

Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inf …

نسخة أولية وصول مفتوح

DroneWAM: Efficient World Action Model for Drone Visual Navigation

Liang Yao, Fan Liu, Hongbo Lu وآخرون · 2026

World-action models give visual navigation agents a way to anticipate how candidate actions will change future observations and to act from the predicted consequences. For drones, this capability must operate under tight accuracy and efficiency constraints. We present DroneWAM, an efficient world-action model for drone …

نسخة أولية وصول مفتوح

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Dehai Min, Daoan Zhang, Yiming Zeng وآخرون · 2026

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable …

المؤلفون المشاركون