الباحثون

Ziqun Bao

المنشورات 1

نسخة أولية وصول مفتوح

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Ziqun Bao, Xinyu Zhang, Yuchen Shao وآخرون · 2026

Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in check …

المؤلفون المشاركون