الباحثون

Jianfei Cai

المنشورات 8

نسخة أولية وصول مفتوح

Seek-and-View Reasoning for Multi-View Spatial Understanding

Qixiang Chen, Cheng Zhang, Fucai Ke وآخرون · 2026

Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, w …

نسخة أولية وصول مفتوح

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Zhuokun Chen, Xi Lin, Xiyu Wu وآخرون · 2026

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fix …

نسخة أولية وصول مفتوح

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

Zhuokun Chen, Feng Chen, Xi Lin وآخرون · 2026

Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we …

نسخة أولية وصول مفتوح

Feedforward Novel View Synthesis for Heterogeneous Cameras

Meng Wei, Cheng Zhang, Boying Li وآخرون · 2026

Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may …

نسخة أولية وصول مفتوح

Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

Xindi Yang, Baolu Li, Liam Lee وآخرون · 2026

Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agen …

نسخة أولية وصول مفتوح

Reconstructing the Dynamic World: A Representation-Centric View of 4D Scene Reconstruction

Ziren Gong, Guo Chen, Yongjia Li وآخرون · 2026

4D scene reconstruction aims to recover the evolving geometry, appearance, and motion of dynamic environments from visual observations. Despite substantial progress in neural scene representations, reconstructing dynamic scenes remains challenging due to non-rigid motion, occlusions, temporal inconsistencies, and the t …

نسخة أولية وصول مفتوح

Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History

Weiqiang Wang, Zhuokun Chen, Yusheng Dai وآخرون · 2026

Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visua …

نسخة أولية وصول مفتوح

FeCoSplat: Feedback-Guided Compression for Feed-Forward 3D Gaussian Splatting

Yuxuan Li, Yihang Chen, Yufeng Zhang وآخرون · 2026

Feed-forward 3D Gaussian Splatting (3DGS) enables efficient novel-view synthesis from sparse multi-view images, yet its representations remain costly to store and transmit. Existing approaches compress either the input images, incurring heavy receiver-side reconstruction, or the reconstructed Gaussian primitives, which …

المؤلفون المشاركون