Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube.webp 1920w" sizes="(max-width: 768px) 100vw, 768px" title="green-cube" />Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-625x352.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-1536x864.png 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/08/green-cube.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="green-cube" /><p>Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE models that match or exceed the performance of dense model counterparts at a fraction of the training compute. MoE models provide efficient training through conditional computation. Instead of one dense feed-forward network (FFN) shared…</p>
<p><a href="https://developer.nvidia.com/blog/accelerating-dropless-moe-training-in-jax-with-nvidia-transformer-engine/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>
Read original article ↗
Related Articles
From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley P
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency
How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories
For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale
Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each comp