الباحثون

Yuan Liu

المنشورات 12

نسخة أولية وصول مفتوح

TKCAM: Text and Keyframe to Camera Trajectory Generation

Haozhe Yang, Zhiyang Dou, Zekai Gu وآخرون · 2026

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional …

نسخة أولية وصول مفتوح

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

Wenbo Chen, Tianfu Li, Haoxuan Xu وآخرون · 2026

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it …

نسخة أولية وصول مفتوح

TrackFish3D: Self-Supervised 3D Tracking of Schooling Fish from Multi-view Videos

Patt Phurtivilai, Zhiyang Dou, Yifan Wu وآخرون · 2026

Quantifying collective fish behavior requires accurate trajectories, yet multi-view 3D tracking remains challenging due to frequent occlusions, visually similar individuals, and the long-standing scarcity of identity annotations. We present TrackFish3D, a geometry-driven self-supervised framework for dense multi-camera …

نسخة أولية وصول مفتوح

JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation

Yafeng Chen, Boya Dong, Yankun Huang وآخرون · 2026

We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation …

نسخة أولية وصول مفتوح

SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models

Tianfu Li, Haoxuan Xu, Wenbo Chen وآخرون · 2026

Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive i …

نسخة أولية وصول مفتوح

FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection

Ben Liang, Chao Sui, Junqi Bai وآخرون · 2026

In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the …

نسخة أولية وصول مفتوح

PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

Zhe Zhu, Yiheng Zhang, Peng Li وآخرون · 2026

Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these s …

نسخة أولية وصول مفتوح

S4R: Scaling for Rigid-Body Interpenetration Resolution

Zhiyang Dou, Ang Zhao, Chen Peng وآخرون · 2026 · 10.1145/3842510

Rigid-body interpenetration frequently occurs in procedurally assembled and generated scenes and must be removed before downstream applications such as physical simulation. We present S4R (Scaling for Rigid-Body Interpenetration Resolution), a scale-continuation method for static interpenetration repair. S4R first unif …

نسخة أولية وصول مفتوح

GrowMTP: Can RL Grow Its Own Draft Head?

Minghua He, Lingzhe Zhang, Yuan Liu وآخرون · 2026

Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing subs …

المؤلفون المشاركون