Authors

Jungseul Ok

Publications 3

Preprint Open access

Dynamic Expert Pruning for Multi-Agent Systems

Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing met …

Co-authors