الباحثون

Kailun Yang

المنشورات 17

نسخة أولية وصول مفتوح

SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction

Yuhang Wang, Kai Luo, Yuanfan Zheng وآخرون · 2026

Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incompatible image features, misaligned historical observations, a …

نسخة أولية وصول مفتوح

FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

Wenya Su, Kai Luo, Di Wen وآخرون · 2026

Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes mu …

نسخة أولية وصول مفتوح

EmbPASS: Towards Cross-Embodiment Open Panoramic Segmentation

Pujun Guo, Yuanfan Zheng, Fei Teng وآخرون · 2026

Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneous embodied platforms exhibit substantial differences in observation viewpoints and spatial layouts, giving rise to cross-embodiment observation shifts that pose addition …

نسخة أولية وصول مفتوح

OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

Runtong Wu, Fei Teng, Di Wen وآخرون · 2026

Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection ( …

نسخة أولية وصول مفتوح

UniTrackPLA: Unified Panorama-Language-Action Model for Instruction-Guided Navigation and Dynamic Person Tracking

Pengfei Qi, Haoran Lin, Sizhuang Chen وآخرون · 2026

General-purpose embodied robots should support both navigation toward language-specified destinations and dynamic person tracking under arbitrary initial target azimuths. However, existing methods typically rely on forward-facing observations and address these tasks with separate policies, limiting omnidirectional perc …

نسخة أولية وصول مفتوح

LensBridge: Frequency-Guided Compound Degradation Adaptation for Lens Aberration Correction and Veiling Glare Removal

Xiaolong Qian, Zhonghua Yi, Qi Jiang وآخرون · 2026

Simplified optical systems often exhibit residual lens aberrations and Veiling Glare (VG), resulting in spatially varying blur and contrast reduction. Large-scale Lens Libraries (LensLib) enable reusable aberration correction models by covering diverse Point Spread Functions (PSFs), but their aberration-only training d …

نسخة أولية وصول مفتوح

EgoRefine: Ego-Referenced Predictive Alignment and Trajectory-Conditioned Reliability-Aware Fusion for Asynchronous Collaborative Perception

Lingzhao Kong, Yongsheng Zang, Yu Kang وآخرون · 2026

Collaborative perception enables connected agents to share complementary observations for 3D object detection, extending sensing range and mitigating occlusion. Under asynchronous communication, however, cooperative features arrive with temporal delay. Existing prediction-based methods compensate for these features mai …

نسخة أولية وصول مفتوح

HEIR: Learning Human-Entity Interactions with Functional Roles

Di Wen, Wenhao Guo, Yuedong Tan وآخرون · 2026

Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score indiv …

نسخة أولية وصول مفتوح

Beyond Isolated Entities: Relation-Aware Multi-Entity Modeling for Unsupervised Video Anomaly Detection

Video anomaly detection for patrol robots and surveillance systems must recognize abnormal interactions among familiar entities. Existing pixel-reconstruction and isolated-entity methods may fail when individual entities appear normal but their spatial or motion relations are abnormal. This work presents Interaction-Ce …

نسخة أولية وصول مفتوح

PanOVOcc: Panoramic Embodied Open-Vocabulary Occupancy Mapping with Long-term Spatial Voxel Memory

Di Kuang, Mengfei Duan, Yuhang Wang وآخرون · 2026

Persistent semantic occupancy mapping is essential for embodied scene understanding. However, perspective-based systems provide limited spatial coverage, while existing panoramic methods primarily predict local volumes from single observations. We introduce PanOVOcc, a training-free framework for persistent open-vocabu …

نسخة أولية وصول مفتوح

PanoFuse: Panorama-Enhanced Vision-Language-Action Learning with Decoupled Semantic-Geometric Routing

Peng Xu, Haoran Lin, Wanjun Jia وآخرون · 2026

Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions, …

نسخة أولية وصول مفتوح

Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models

Di Wen, Ruodi Zhang, Kailun Yang وآخرون · 2026

Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captured temporal structur …

نسخة أولية وصول مفتوح

RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision

Ruiping Liu, Shaofang Quan, Qian Yin وآخرون · 2026

Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether …

نسخة أولية وصول مفتوح

OmniMimic: Dynamics-completed Motion Augmentation for Multi-style Omnidirectional Quadruped Locomotion

Sheng Wu, Guoqiang Zhao, Zhe Yang وآخرون · 2026

Animal demonstrations provide quadruped robots with natural and distinctive gait styles that are difficult to specify through hand-crafted rewards. However, their narrow directional coverage leaves little style-consistent supervision for backward, lateral, and turning commands. We present OmniMimic, a training framewor …

نسخة أولية وصول مفتوح

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

Zhikun Zhou, Kunyu Peng, Runyi Yang وآخرون · 2026

Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independentl …

نسخة أولية وصول مفتوح

INSPECT: Learning Robot View Selection from Assistant Use

Di Wen, Kailun Yang, Wenhao Guo وآخرون · 2026

Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, whic …

المؤلفون المشاركون