الباحثون

Bin Yang

المنشورات 11

نسخة أولية وصول مفتوح

V-CoLA: Vision Token Compression with Linear Attention

Hao Jiang, Yiru Mao, Tianpeng Bu وآخرون · 2026

Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating …

نسخة أولية وصول مفتوح

SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces

Rongxue Li, Meng Yang, Yiru Mao وآخرون · 2026

Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM …

نسخة أولية وصول مفتوح

QiYao-I: A Manifold Based Foundation Model for Irregular Multivariate Time Series Forecasting

Linfeng Wang, Ruitong Zhang, Kai Zhao وآخرون · 2026

Irregular multivariate time series forecasting is a challenging yet important problem in real-world applications, where observations are often irregularly sampled and asynchronously recorded across variables. Existing time series foundation models are mostly built on regularly sampled sequences, making them difficult t …

نسخة أولية وصول مفتوح

Do Neural PDE Solvers Learn the Right Dynamics?

Haonan Li, Yue Song, Bin Yang وآخرون · 2026

Neural PDE solvers can achieve low prediction errors, but do they reproduce the dynamics of the systems they model? Prediction scores alone offer an incomplete answer: they measure agreement with reference solutions but provide limited insight into how errors accumulate, nearby states diverge, or extreme events arise. …

نسخة أولية وصول مفتوح

Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

Duanning Chen, Ke He, Bin Yang وآخرون · 2026

Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterp …

نسخة أولية وصول مفتوح

HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling

Chunyi Hou, Xiangfei Qiu, Hanyin Cheng وآخرون · 2026

Most existing time series generators rely on a two-stage modeling paradigm: the first stage learns discrete latent representations of time series; the second stage performs autoregressive modeling on these discrete latents through next token prediction. However, this paradigm suffers from two stage-specific limitations …

نسخة أولية وصول مفتوح

QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities

Hanyin Cheng, Linfeng Wang, Zhengbo Qu وآخرون · 2026

Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities …

نسخة أولية وصول مفتوح

XMatch: Enhancing Covariate-Aware Time Series Forecasting through Tree-Structured Exogenous Matching

Ziyang Zhang, Hanyin Cheng, Xiangfei Qiu وآخرون · 2026

Future exogenous variables provide valuable information for forecasting endogenous time series. Existing covariate-aware methods primarily learn the direct influence of exogenous variables on endogenous variables. However, these effects can be complex and change with the pattern of the exogenous variables, making them …

نسخة أولية وصول مفتوح

OneWorld: Learning Consistent Physics Across Actions in World Models

Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying inco …

المؤلفون المشاركون