1.43× Faster MoE Training: LoongForge Redefines EP Expert Load Balancing with Topology-Aware Optimal Transport
PyTorch Forums
1.43× Faster MoE Training: LoongForge Redefines EP Expert Load Balancing with Topology-Aware Optimal Transport
MoE is now the default architecture for frontier models, and the direction of the next generation is clear: more experts, sparser activation. DeepSeek-V4-Pro raised the number of routed experts per layer from 256 in V3 to 384, while the number of experts activated per token actually dropped from 8 to 6. Those architectural gains come at a price: the complexity moves into the training system. How many tokens each GPU has to process depends entirely on routing results, and the more numerous and f...
0 comments
No comments yet.