الباحثون

Jianwei Zhang

المنشورات 6

نسخة أولية وصول مفتوح

TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

Xin Wang, Hao Yu, Zhengyang Zhuge وآخرون · 2026

Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization ac …

نسخة أولية وصول مفتوح

Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training

Shuyuan Tu, Qi Tian, Yinming Huang وآخرون · 2026

Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informativ …

نسخة أولية وصول مفتوح

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents …

نسخة أولية وصول مفتوح

MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models

Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan وآخرون · 2026

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain co …

نسخة أولية وصول مفتوح

M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease

Lujia Zhong, Shuo Huang, Jianwei Zhang وآخرون · 2026

While various multimodal methods combining imaging and tabular data for Alzheimer's disease (AD) diagnosis were proposed, they are often limited in generalization across cohorts. In-context learning (ICL) has demonstrated excellent generalization performances and high flexibility in foundational tabular models such as …

نسخة أولية وصول مفتوح

ForceDelta-VLA: Distilling Force-Conditioned ActionCorrections for Contact-Rich Manipulation

Ju Dong, Yu Fu, Jian Chen وآخرون · 2026

Force-aware Vision-Language-Action (VLA) policies improve contact-rich manipulation, but typically combine task-level motion and contact-dependent adjustment in a single action prediction. Demonstrations provide no explicit labels for decomposing that prediction into a reusable reference action and a correction. We pre …

المؤلفون المشاركون