Preprint Open access
AutoLoCo: Communication Efficient Distributed LLM Training via Adaptive Synchronization
The pre-training of Large Language Models (LLMs) is increasingly conducted across multiple data centers. As training scales to a larger number of accelerators, the fraction of time spent on computation decreases, while the fraction spent on communication increases. Therefore, frequent synchronization becomes a growing …