الباحثون

Alois Knoll

المنشورات 8

نسخة أولية وصول مفتوح

DTFormer: Text-Guided Semantic Alignment for RGB-D Segmentation

Ziang Wei, Yinlong Liu, Yan Xia وآخرون · 2026

RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel-level supervision, lacking direct high-level semantic constraints. This raises a central question-can external knowledge such as language priors inject stronger semanti …

نسخة أولية وصول مفتوح

Learning Chaos Without Seeing Chaos: Extrapolation of Global Dynamics in Autoregressive Transformers

Yilun Liu, Yi Zhang, Ganyu Wu وآخرون · 2026

Autoregressive models are trained to predict a system's behavior one step at a time, and recursive generation allows the learned dynamics to unfold over long horizons. To what extent can such dynamics learned from local observations recover broader organization of an underlying system that was only partially observed d …

نسخة أولية وصول مفتوح

VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness

Hang Li, Mingxin Zhang, Zihan Wu وآخرون · 2026

Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without dist …

نسخة أولية وصول مفتوح

Timed Rule-Based Supervision of an End-to-End Autonomous Parking Policy

Kejia Gao, Liguo Zhou, Lei Yu وآخرون · 2026

We study whether a manually specified runtime supervisor can correct recurring failures of an existing end-to-end parking policy in a fixed CARLA parking lot. The vision-based Transformer architecture is inherited from Yang et al.; our contribution is a timed, rule-based Parametric Safety Shield (PSS) applied to its co …

نسخة أولية وصول مفتوح

React When You Need To: Event-Triggered Asynchronous Inference for VLA Policies

Yansong Wu, Huaqing Li, Tianding Hou وآخرون · 2026

Vision-Language-Action (VLA) models commonly predict action chunks, limiting their ability to react to environmental changes during execution. Existing asynchronous inference methods improve reactivity but typically rely on a fixed inference gap. In this paper, we propose an event-guided dynamic inference strategy that …

نسخة أولية وصول مفتوح

VT-Bridge: Bridging Pretrained Foundation VLAs to VTLAs via Lightweight Residual Adaptation

Yansong Wu, Tuo Yang, Rongping Zhao وآخرون · 2026

Vision-Tactile-Language-Action (VTLA) models have demonstrated clear advantages over Vision-Language-Action (VLA) models in contact-rich manipulation. However, developing VTLA models is severely constrained by the massive amounts of vision-tactile data and computational resources required. To address this bottleneck, w …

المؤلفون المشاركون