Open App →
Back to News Feed
Nvidia Blog August 24, 2026 By Shruti Koparkar neutral

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIAAI / LLMNVIDIA / GPU
According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why?  Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]
Read original article ↗

Related Articles

Agentrys Raises $24.5 Million to Build Agentic Design Automation for Chipmakers

Platform from Mark Ren, who pioneered AI for chip design at NVIDIA, enables engineering teams to build and own a self-im

Semiconductor Digest · August 26, 2026

Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding

Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers

Nvidia Developer Blog · August 26, 2026

Top Ten Semiconductor Companies in Q2

NVIDIA USA $125.7 Billion AI GPUs & Data Center Accelerators Samsung Electronics South Korea $72.7 Billion (Memory/DS Di

Electronics Weekly · August 26, 2026

Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo

When an LLM engine process fails, the standard recovery path involves a cold restart. This requires loading weights into

Nvidia Developer Blog · August 25, 2026