نسخة أولية وصول مفتوح
ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases that degrade downstrea …