الباحثون

Bo Li

المنشورات 6

نسخة أولية وصول مفتوح

RealtimeWAM: How Fast Can I Run My World Action Model?

Huanan Liu, Ye Li, Kangye Ji وآخرون · 2026

World Action Models (WAMs) combine visual dynamics modeling with action generation, but their high inference latency limits responsive robot control. Recent efforts accelerate inference by removing explicit future-video generation at test time, as in FastWAM, an approach that requires a specially tailored architectural …

نسخة أولية وصول مفتوح

Complementary Supervised and Self-Supervised Representations for Out-of-Distribution Graph Learning

Qingying Hao, Zikang Chen, Chuxuan Hu وآخرون · 2026

Out-of-distribution (OOD) generalization remains challenging for graph neural networks (GNNs), as graph distributions can vary substantially across time and domains. Supervised and self-supervised graph representation learning are guided by distinct objectives and offer different perspectives on graph representations. …

نسخة أولية وصول مفتوح

Do ResNets Route? Sparse Interaction Experts in Residual Networks

Liang Yan, Siying Chen, Kaijie Chen وآخرون · 2026

Residual networks execute every block for every input, yet their functional contributions need not be input independent. We formulate a trained ResNet as a set function over binary residual-branch masks and apply Möbius inversion to decompose its output exactly into individual residual corrections and higher-order inte …

نسخة أولية وصول مفتوح

Hear the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation

Hanmo Chen, Chengcheng Liu, Tianxiao Chen وآخرون · 2026

Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requires the perceived sou …

نسخة أولية وصول مفتوح

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

Chengyang Zhang, Wenchuan Zhang, Bo Li وآخرون · 2026

Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even whe …

نسخة أولية وصول مفتوح

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

Liyang Fan, Chi Wei, Yitai Li وآخرون · 2026

Multimodal coding agents are expected to turn visual inputs into usable artifacts, and they act through a harness, the layer of tools, context management, and execution environment around the model. Existing evaluations often isolate short tool calls, API traces, or screenshot resemblance, and a low score under these p …

المؤلفون المشاركون