Open App →
Back to News Feed
Nvidia Developer Blog September 9, 2026 positive

When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving

NVIDIAAI / LLMNVIDIA / GPU
<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo.webp 1999w" sizes="(max-width: 768px) 100vw, 768px" title="multimodal-dynamo" />Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-195x110.jpg 195w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/multimodal-dynamo.webp 1999w" sizes="auto, (max-width: 768px) 100vw, 768px" title="multimodal-dynamo" /><p>Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill and decode stages. It is most effective for image-heavy prompts, short-to-medium outputs, and quantized mixture-of-experts (MoE) models. This post shows when and how to use EPD disaggregation with NVIDIA Dynamo to achieve up to 5x…</p> <p><a href="https://developer.nvidia.com/blog/when-to-use-encode-prefill-decode-disaggregation-to-accelerate-multimodal-model-serving/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>
Read original article ↗

Related Articles

CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs

Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVI

Nvidia Developer Blog · September 9, 2026

ARPA-H Selects UpDoc to Lead Development of Autonomous Clinical AI System, in an Initiative Joined by Microsoft, OpenAI and NVIDIA

UpDoc awarded up to $9.2 million under ARPA-H's ADVOCATE program to develop and validate an autonomous clinical AI syste

PR Newswire · September 9, 2026

NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC

At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming to

Nvidia Blog · September 9, 2026

NVIDIA Vera: Rebuilding the CPU for Agentic AI

Source: Jonathon Evans and Polychronis Xekalakis, “NVIDIA Vera CPU,” Hot Chips 2026. Performance figures are NVIDIA clai

SemiWiki · September 8, 2026