الباحثون

Huan Wang

المنشورات 5

نسخة أولية وصول مفتوح

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

Qitong Wang, Xinwei Niu, Mingluo Su وآخرون · 2026

The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding …

نسخة أولية وصول مفتوح

Universal Textual Teaching for LLMs

Zhanyi Lu, Huan Wang · 2026

Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across model …

نسخة أولية وصول مفتوح

MetaKernelBench: Measuring GPU Kernel Knowledge Transfer Beyond Code

Xueyi Chen, Shiyu Liu, Xin Jin وآخرون · 2026

Recent GPU kernel optimization agents retain what they learn in knowledge bases or as distilled skills. Kernel benchmarks score each attempt's implementation for correctness and speed but leave the reuse value of retained experience unmeasured. We introduce MetaKernelBench, which measures whether experience distilled f …

نسخة أولية وصول مفتوح

Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning

Muhammad Asif Ali, Wenqing Wang, Huan Wang وآخرون · 2026

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redunda …

المؤلفون المشاركون