الباحثون

Kunyu Peng

المنشورات 11

نسخة أولية وصول مفتوح

FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

Wenya Su, Kai Luo, Di Wen وآخرون · 2026

Fisheye cameras give mobile robots a single-sensor, low-cost view of their surroundings, yet the COCO-pretrained detectors that practitioners routinely reuse fail on them: strong radial distortion warps local image structure, while boundary compression shrinks objects to near-invisible sizes. Full fine-tuning closes mu …

نسخة أولية وصول مفتوح

OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

Runtong Wu, Fei Teng, Di Wen وآخرون · 2026

Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection ( …

نسخة أولية وصول مفتوح

HEIR: Learning Human-Entity Interactions with Functional Roles

Di Wen, Wenhao Guo, Yuedong Tan وآخرون · 2026

Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score indiv …

نسخة أولية وصول مفتوح

Algebraic Consistency Alone Does Not Certify Temporal Structure in Latent Action Models

Di Wen, Ruodi Zhang, Kailun Yang وآخرون · 2026

Latent action models infer a code for the transition between two frames of action-free video. Recent methods regularise this code to compose additively and reverse antisymmetrically, and report order-of-magnitude reductions in the resulting errors as a label-free certificate that the code has captured temporal structur …

نسخة أولية وصول مفتوح

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

Zhiyuan Gao, Di Wen, Yanxiang Zhan وآخرون · 2026

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in …

نسخة أولية وصول مفتوح

RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision

Ruiping Liu, Shaofang Quan, Qian Yin وآخرون · 2026

Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether …

نسخة أولية وصول مفتوح

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

Zhikun Zhou, Kunyu Peng, Runyi Yang وآخرون · 2026

Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independentl …

نسخة أولية وصول مفتوح

INSPECT: Learning Robot View Selection from Assistant Use

Di Wen, Kailun Yang, Wenhao Guo وآخرون · 2026

Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, whic …

المؤلفون المشاركون