الباحثون

Qianren Mao

المنشورات 3

نسخة أولية وصول مفتوح

Stochastic Teacher Intervention for Agentic On-Policy Distillation

Junnan Liu, Linhao Luo, Zhijun Chen وآخرون · 2026

On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape sub …

نسخة أولية وصول مفتوح

Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

Qili Zhang, Qianren Mao, Hanze Cai وآخرون · 2026

Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final …

نسخة أولية وصول مفتوح

Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

Zhenting Huang, Bo Jiang, Junnan Liu وآخرون · 2026

Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reli …

المؤلفون المشاركون