Open App →
Back to News Feed
Nvidia Developer Blog September 21, 2026 neutral

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

NVIDIAAI / LLMNVIDIA / GPUMemorySupply Chain
<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-625x351.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1.webp 1035w" sizes="(max-width: 768px) 100vw, 768px" title="grid-robot-arm-cleaning-plate" />The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-768x432.png" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-768x432.png 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-179x101.png 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-300x169.png 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-625x351.png 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-645x363.png 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-660x370.png 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-500x281.png 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-160x90.png 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-362x204.png 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-196x110.png 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-1024x576.png 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1-960x540.png 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/grid-robot-arm-cleaning-plate-1.webp 1035w" sizes="auto, (max-width: 768px) 100vw, 768px" title="grid-robot-arm-cleaning-plate" /><p>The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability that enables a single TensorRT network to execute across multiple GPUs using NCCL-backed distributed collectives while retaining TensorRT inference optimizations. It is fully supported starting with TensorRT 11.0.</p> <p><a href="https://developer.nvidia.com/blog/simplifying-model-serving-across-multiple-gpus-with-nvidia-tensorrt-multi-device-integration-in-nvidia-dynamo-triton/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>
Read original article ↗

Related Articles

NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories

Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooli

Nvidia Blog · September 21, 2026

LG Energy Solution's Battery Energy Storage System Qualifies as NVIDIA DSX Ready BESS

LG Energy Solution's BESS offering has qualified as DSX Ready BESS Advanced AC-coupled ESS solution helps accelerate tim

PR Newswire · September 21, 2026

From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale

Today, Egypt’s AI builders gathered in the Grand Egyptian Museum for a reception that highlighted the nation’s rapidly g

Nvidia Blog · September 21, 2026

Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2

Weather-sensitive industries increasingly have access to observations that offer an earlier, more local view of changing

Nvidia Developer Blog · September 21, 2026