الباحثون

Chunchao Guo

المنشورات 9

نسخة أولية وصول مفتوح

Position Forcing: Self-Conditioning 3D Generation

Ziheng Ouyang, Zeqiang Lai, Jiarui Chen وآخرون · 2026

Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation qu …

نسخة أولية وصول مفتوح

OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

Zhenyang Liu, Chenjie Cao, Yisu Zhang وآخرون · 2026

Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual traj …

نسخة أولية وصول مفتوح

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

Jiarui Chen, Zeqiang Lai, Jiangshan Wang وآخرون · 2026

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trac …

نسخة أولية وصول مفتوح

FlowHMR: Physically Plausible Motion Capture from Video

Zhanke Wang, Chengfeng Zhao, Qing Shuai وآخرون · 2026

We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and train the network with geometric supervision. However, recovering human motion from monocular video is inherently ambigu …

نسخة أولية وصول مفتوح

Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization

Zhen Zhou, Zhiwei Ning, Puhua Jiang وآخرون · 2026

Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without …

نسخة أولية وصول مفتوح

Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?

Jiangshan Wang, Zeqiang Lai, Jiayi Guo وآخرون · 2026

Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work …

نسخة أولية وصول مفتوح

CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving

Yu Meng, Baining Zhao, Junta Wu وآخرون · 2026

Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-view driving video g …

نسخة أولية وصول مفتوح

WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon

Haiyu Zhang, Wenqiang Sun, Tengfei Wang وآخرون · 2026

Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time responsiveness. In this pap …

المؤلفون المشاركون