نسخة أولية وصول مفتوح
Dynamic Minimax Regret Optimization for Robust LLM Post-Training
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the train …