الباحثون

Kun Wang

المنشورات 7

نسخة أولية وصول مفتوح

GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA

Kun Wang, Yupeng Hu, Ruping Cao وآخرون · 2026

GoldenViewVQA requires models to jointly answer driving-scene questions and identify the camera view containing the supporting visual evidence, making precise evidence localization as important as answer correctness. We present \textbf{CoVeR-VQA}, a training-free multi-stage verification and correction framework for gr …

نسخة أولية وصول مفتوح

VisualErase: Dual-Branch Visual Trajectory Redirection for Robust Concept Erasure in Text-to-Image Diffusion Models

Qianlong Xiang, Miao Zhang, Kun Wang وآخرون · 2026

Concept erasure is essential for the safe deployment of text-to-image diffusion models, as they may reproduce harmful, copyrighted, or privacy-sensitive content learned from unconstrained large-scale data. Existing methods typically erase unwanted concepts while preserving general generation capability by redirecting t …

نسخة أولية وصول مفتوح

XTurnix: Large-Scale Self-Supervised Turn Control through Two-State Binary Decisions

Zhanxun Liu, Yifan Duan, Hengtao Wu وآخرون · 2026

General turn-taking behavior in real-time dialogue systems requires deciding whether to keep listening or start responding while listening, and whether to continue or stop while speaking. Existing turn detectors use heterogeneous, task-specific label spaces and are often trained on limited annotations or evaluated on i …

نسخة أولية وصول مفتوح

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

Xuhua Chen, Zhenhan Yin, Yuan Zhang وآخرون · 2026

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0 …

نسخة أولية وصول مفتوح

Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models

Qiankun Li, Yuechen Zhang, Bowen Chen وآخرون · 2026

Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial in …

المؤلفون المشاركون