الباحثون

Muyang Li

المنشورات 5

نسخة أولية وصول مفتوح

GraphOPD: Graph-Augmented On-Policy Distillation for LLM Agents

Bohan Lin, Liyi Chen, Zhuoning Guo وآخرون · 2026

On-policy distillation post-trains large language model agents by supplying dense, step-level guidance from a teacher policy when the reinforcement-learning reward is sparse and arrives only once per trajectory. Existing instantiations allocate this guidance by the size of the teacher-student divergence at each step, o …

نسخة أولية وصول مفتوح

Constant-Curvature Sliced Gromov-Wasserstein for Heterogeneous Cross-Curvature Alignment

Shanglin Li, Wenjing Lu, Muyang Li وآخرون · 2026

Recent advances in representation learning have highlighted the utility of constant-curvature models, such as hyperbolic and spherical spaces, for modeling complex data. Mixed-curvature models further enhance this by integrating multiple constant-curvature components. However, these models typically learn each componen …

نسخة أولية وصول مفتوح

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Zhengyu Fang, Seoyeon Hong, Jie Yang وآخرون · 2026

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialis …

نسخة أولية وصول مفتوح

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Muyang Li, Jie Yang, Zhengyu Fang وآخرون · 2026

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta …

نسخة أولية وصول مفتوح

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

Suqin Yuan, Runqi Lin, Muyang Li وآخرون · 2026

Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show t …

المؤلفون المشاركون