Hand-built MoE expert parallelism: routing-skew capture, fused grouped-GEMM Triton kernels, 2-all2all dispatch/combine, overlap ablations — proving by counter-example why DeepEP exists.
cuda triton moe nccl mixture-of-experts all-to-all deepseek gpu-profiling distributed-inference kernel-optimization deepep grouped-gemm expert-parallelism
-
Updated
Jul 18, 2026 - Python