الباحثون

Wei Lin

المنشورات 6

نسخة أولية وصول مفتوح

Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

Akshit Singh, Shyam Marjit, Wei Lin وآخرون · 2026

Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computati …

نسخة أولية وصول مفتوح

On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training

Shufan Shen, Zhongni Hou, Junshu Sun وآخرون · 2026

The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generaliza …

نسخة أولية وصول مفتوح

SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

Miteto Wei, Xiaohan Wang, Zehao Chen وآخرون · 2026

On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-cons …

نسخة أولية وصول مفتوح

TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs

Zihan Lin, Xiaohan Wang, Jie Cao وآخرون · 2026

Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or leave the benefit of ex …

نسخة أولية وصول مفتوح

IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents

Angqing Jiang, Gaoming Zhang, Chaoqun Zhang وآخرون · 2026

On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state. …

المؤلفون المشاركون