From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry
<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4.webp 1920w" sizes="(max-width: 768px) 100vw, 768px" title="image1" />NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two...<img width="768" height="432" src="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg" class="webfeedsFeaturedVisual wp-post-image" alt="" style="display: block; margin-bottom: 5px; clear:both;max-width: 100%;" link_thumbnail="" decoding="async" loading="lazy" srcset="https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-768x432.jpg 768w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-179x101.jpg 179w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-300x169.jpg 300w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-625x352.jpg 625w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1536x864.jpg 1536w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-645x363.jpg 645w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-660x370.jpg 660w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-500x281.jpg 500w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-160x90.jpg 160w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-362x204.jpg 362w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-196x110.jpg 196w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-1024x576.jpg 1024w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4-960x540.jpg 960w, https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/image1-4.webp 1920w" sizes="auto, (max-width: 768px) 100vw, 768px" title="image1" /><p>NVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two parts. Time-to-rack runs from silicon leaving the fab to an assembled system arriving on a data center floor. Time-to-token covers everything thereafter: power, cooling, networking, and the software stack that makes the infrastructure…</p>
<p><a href="https://developer.nvidia.com/blog/from-wafer-out-to-first-token-codifying-supply-chain-expertise-with-nemotron-and-palantir-foundry/" rel="nofollow" data-wpel-link="internal" target="_self">Source</a></p>
Read original article ↗
Related Articles
When to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model Serving
Encode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates t
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVI
UpDoc awarded up to $9.2 million under ARPA-H's ADVOCATE program to develop and validate an autonomous clinical AI syste
NVIDIA Brings Real-Time AI to Broadcast, Sports and Global Streaming at IBC
At the IBC conference, running Sept. 11-14 in Amsterdam, the creative, technology and business communities are coming to