الباحثون

Xiangyu Yue

المنشورات 6

نسخة أولية وصول مفتوح

Position Forcing: Self-Conditioning 3D Generation

Ziheng Ouyang, Zeqiang Lai, Jiarui Chen وآخرون · 2026

Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation qu …

نسخة أولية وصول مفتوح

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

Jiarui Chen, Zeqiang Lai, Jiangshan Wang وآخرون · 2026

Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trac …

نسخة أولية وصول مفتوح

Vela: Scaling Vision-Language-Action Models with Adaptive Action Curve Parametrization

Yifan Li, Jiaxu Wang, Dongming Wu وآخرون · 2026

Most vision-language-action models represent future motion as fixed-rate action chunks, tying temporal resolution and prediction horizon to a fixed output budget. This pointwise representation wastes capacity on highly correlated neighboring actions, leaves temporal continuity and smoothness to be learned implicitly, a …

نسخة أولية وصول مفتوح

EviRover: Reinforcing Agentic Perception Beyond a Glance

Kaixuan Fan, Kaituo Feng, Tianshuo Peng وآخرون · 2026

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi …

نسخة أولية وصول مفتوح

Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?

Jiangshan Wang, Zeqiang Lai, Jiayi Guo وآخرون · 2026

Native 3D texture generation synthesizes colors directly in 3D space for a given geometry, conditioned on multi-view reference images. It is generally believed that training such models requires large-scale, high-quality real 3D asset data, whose acquisition remains a long-standing and challenging problem. In this work …

نسخة أولية وصول مفتوح

Watch, Recall, Act: Always-On Robots in Concurrent Embodied Streams

Ding Yi, Peiwen Sun, Chenchu Rong وآخرون · 2026

An always-on robot faces an endless stream that never resets: instructions arrive and lapse, the scene changes, and its own past actions reshape what it must reason about. Today's action models are built for the opposite: a fixed instruction, no mid-task intervention, single-step reasoning. In an open-ended world a rob …

المؤلفون المشاركون