الباحثون

Yubin Jing

المنشورات 1

نسخة أولية وصول مفتوح

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or prefere …

المؤلفون المشاركون