الباحثون

Zixiang Ni

المنشورات 2

نسخة أولية وصول مفتوح

ReTaCo: Residual-Target Control for On-Policy Distillation

Zixiang Ni, Zhuo Hu, Renjie Cao وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it …

نسخة أولية وصول مفتوح

CARAT: Do Materials LLMs Reason or Recite?

Jiajun Wu, Jian Yang, Zixiang Ni وآخرون · 2026

When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names …

المؤلفون المشاركون