الباحثون

Zhi Wang

المنشورات 16

نسخة أولية وصول مفتوح

RealtimeWAM: How Fast Can I Run My World Action Model?

Huanan Liu, Ye Li, Kangye Ji وآخرون · 2026

World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural …

نسخة أولية وصول مفتوح

Organize Primitives into Semantic Parts: Reinforcement Reasoning for 3D Segmentation

Xiaoming Gong, Ruoyu Wu, Zhenhong Sun وآخرون · 2026

Primitive-based 3D segmentation offers a compact and explicit alternative to dense surface prediction, naturally supporting structural abstraction and boundary localization. However, geometric decomposition alone does not determine how primitives should be organized into semantic parts: a single part may span multiple …

نسخة أولية وصول مفتوح

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao وآخرون · 2026

Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce ChunkTrust, which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware H …

نسخة أولية وصول مفتوح

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

Xinhao Yang, Wenhao Wu, Ning Lv وآخرون · 2026

3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behavio …

نسخة أولية وصول مفتوح

CurvSpec: Adaptive Multi-Curvature Learning for Partial Relevant Video Retrieval

Zhen Liu, Letian Li, Jinpeng Wang وآخرون · 2026 · 10.1145/3767308.3835167

Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging t …

نسخة أولية وصول مفتوح

Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

Xianghan Wei, Xiaoda Yang, Zhi Wang وآخرون · 2026

While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free appro …

نسخة أولية وصول مفتوح

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

Zhimin Wang, Meiyuan Zhu, Duo Wu وآخرون · 2026

Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For examp …

نسخة أولية وصول مفتوح

Clinical Trajectory Alignment for Medical Vision-Language Pre-training

Huimin Yan, Xian Yang, Zhi Wang وآخرون · 2026

Medical vision-language pre-training largely follows a visit-level image-report matching paradigm, aligning paired images and reports at individual visits. While effective for static cross-modal correspondence, this paradigm provides limited supervision for longitudinal clinical change, such as whether abnormalities im …

نسخة أولية وصول مفتوح

CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

Shuzhao Xie, Lelin Wang, Guying Lin وآخرون · 2026

Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose …

نسخة أولية وصول مفتوح

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

Duo Wu, Haifeng Wang, Rongwei Lu وآخرون · 2026

Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introdu …

نسخة أولية وصول مفتوح

Weaponizing Ground Truth: Data Poisoning Attacks by Exploiting Boundary Misalignment Between Antivirus Software and Learning-Based Detectors

Jieshuai Yang, Zhi Wang, Yan Jia وآخرون · 2026

Machine-learning (ML)-based malware detectors are commonly trained using labels obtained from antivirus (AV) engines and aggregation services (e.g., VirusTotal). This practice assumes AV-generated labels provide reliable supervision. However, small byte-level modifications can substantially alter AV verdicts while leav …

نسخة أولية وصول مفتوح

Breaking the Black Box: Byte-Level Boundary Inference of Real-World Antivirus Systems

Jieshuai Yang, Zhi Wang, Yan Jia وآخرون · 2026

Existing approaches for understanding the detection logic of real-world antivirus (AV) software infer only binary malware/benign decisions from black-box queries, providing limited insight into the fine-grained decision-critical regions that govern AV detection. In this paper, we present \textbf{AVHunter}, the first fr …

نسخة أولية وصول مفتوح

HarnessPAI: An Evolving Harness for Physical AI

Xin Wang, Wenhao Wu, Menghao Zhang وآخرون · 2026

Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabi …

المؤلفون المشاركون