نسخة أولية وصول مفتوح
DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. …