الباحثون

Zhijun Chen

المنشورات 2

نسخة أولية وصول مفتوح

Stochastic Teacher Intervention for Agentic On-Policy Distillation

Junnan Liu, Linhao Luo, Zhijun Chen وآخرون · 2026

On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape sub …

نسخة أولية وصول مفتوح

Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

Qili Zhang, Qianren Mao, Hanze Cai وآخرون · 2026

Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final …

المؤلفون المشاركون