GTC March 2025 Keynote with NVIDIA CEO Jensen Huang
Jensen Huang (Founder and CEO · NVIDIA)
NVIDIA GTC 2025 · Keynote
Overview
NVIDIA CEO Jensen Huang's GTC March 2025 keynote delivered a comprehensive vision for the future of artificial intelligence, emphasizing a profound platform shift toward accelerated computing and the emergence of "AI factories." The talk centered on the rapid evolution of AI from its perceptual and generative stages to the sophisticated capabilities of agentic AI and physical AI, which are now driving an unprecedented demand for computational power. Huang articulated how NVIDIA is addressing this demand with a full-stack approach, encompassing groundbreaking hardware architectures like Blackwell and its successors, advanced interconnect technologies, and a vast ecosystem of software and frameworks designed to accelerate every layer of the AI workflow.

Key moments
- 0:00 Opening vision: AI's impact across industries
- 2:40 Jensen Huang welcomes to GTC, no scripts
- 4:48 Unveiling GeForce 5090: Blackwell generation performance
- 6:10 AI's journey: from perception to generative computing
- 7:55 Understanding Agentic AI: perception, reasoning, action
- 8:45 Next wave: Physical AI enabling robotics
- 11:00 Three fundamental AI challenges: data, training, scaling
GTC March 2025 Keynote with NVIDIA CEO Jensen Huang
Speakers: NVIDIA (NVIDIA)
Conference: NVIDIA GTC
YouTube: https://www.youtube.com/watch?v=_waPvOwL9Z8
Overview
NVIDIA CEO Jensen Huang's GTC March 2025 keynote delivered a comprehensive vision for the future of artificial intelligence, emphasizing a profound platform shift toward accelerated computing and the emergence of "AI factories." The talk centered on the rapid evolution of AI from its perceptual and generative stages to the sophisticated capabilities of agentic AI and physical AI, which are now driving an unprecedented demand for computational power. Huang articulated how NVIDIA is addressing this demand with a full-stack approach, encompassing groundbreaking hardware architectures like Blackwell and its successors, advanced interconnect technologies, and a vast ecosystem of software and frameworks designed to accelerate every layer of the AI workflow.
This keynote is critical because it outlines NVIDIA's strategy for enabling the next era of AI, which is characterized by systems that can reason, plan, and interact with the physical world. The projected computational requirements are staggering—easily 100 times more than previously anticipated—necessitating a complete re-architecture of data centers into specialized AI factories. Huang's presentation not only unveiled significant technological advancements but also provided a clear roadmap for how these innovations will democratize AI, extending its reach from cloud hyperscalers to enterprise, edge, and ultimately, into the realm of general-purpose robotics, transforming industries and unlocking new frontiers of scientific discovery.
Background
▶ Watch: Opening vision: AI's impact across industries (0:00)
The journey of AI, as contextualized by Jensen Huang, has seen several distinct waves. It began with perception AI (computer vision, speech recognition) about a decade ago, followed by the generative AI revolution in the last five years. Generative AI fundamentally shifted computing from a retrieval-based model, where pre-stored content was fetched, to a generative model, where AI understands context and creates new content on demand—be it text, images, video, or even scientific data like proteins and chemicals. This paradigm shift has permeated every layer of computing, leading to a major breakthrough in the last two to three years: agentic AI.
Agentic AI represents a significant leap, endowing AI with agency. Such systems can perceive and understand context, reason about problems, plan solutions, and take action using various tools and modalities, including interpreting websites or videos. This reasoning capability, built on techniques like Chain of Thought, best-of-N consistency checking, and path planning, involves generating sequences of "thinking tokens" that drastically increase computational requirements. Following agentic AI, the next wave is physical AI, which enables robotics by embedding an understanding of the three-dimensional physical world—including concepts like friction, inertia, cause and effect, and object permanence—into AI models.
These successive waves of AI have exposed three fundamental challenges in the ML/systems space:
- The Data Problem: AI is data-driven, requiring vast amounts of digital experience to learn.
- The Training Problem (without human-in-the-loop): To learn at superhuman rates and scales, AI training cannot be limited by human demonstration.
- The Scaling Law Problem: Finding algorithms and architectures where providing more resources consistently makes the AI smarter.
Critically, Huang noted that the industry largely underestimated the computational requirement for AI. The scaling law of AI is proving to be "hyper-accelerated," with agentic AI demanding easily 100 times more computation than anticipated just a year prior. This exponential growth stems from the need to generate many more tokens for reasoning and to process them much faster to maintain interactive responsiveness. This immense demand signals that general-purpose computing has run its course, necessitating a platform shift to accelerated computing powered by GPUs, transforming data centers into specialized AI factories dedicated to generating tokens.
Key Findings
▶ Watch: Unveiling GeForce 5090: Blackwell generation performance (4:48)
The keynote unveiled a series of pivotal advancements and observations driving the next era of AI:
- Explosive Computational Demand: Agentic AI, with its reasoning capabilities, requires an astounding 100 times more computation for inference than previous models, necessitating trillions of tokens for training through advanced methods like Reinforcement Learning with Verifiable Results (RLVR).
- Blackwell Architecture Dominance: The Blackwell GPU architecture (specifically the Grace Blackwell NVLink 72 rack) delivers a 25x performance improvement over the previous Hopper generation at iso-power for general inference workloads. For complex reasoning models, this performance gain escalates to 40x. Blackwell is now in full production, experiencing incredible demand from hyperscalers.
- Extreme Scale-Up with NVLink 72: NVIDIA introduced a revolutionary disaggregated NVLink 72 architecture, enabling 1 ExaFLOP of AI computation in a single liquid-cooled rack. This system effectively treats 72 GPUs as one massive GPU, overcoming traditional scaling limitations and memory bandwidth bottlenecks (570 terabytes per second).
- Dynamo: The AI Factory Operating System: A new open-source software stack, NVIDIA Dynamo, was announced. It acts as the "operating system" for AI factories, managing complex inference workloads, including dynamic prefill/decode phases, various parallelism techniques (tensor, pipeline, expert), in-flight batching, and KV cache management, to maximize throughput and quality of service.
- Predictable Annual Roadmap: NVIDIA committed to an annual cadence for its AI infrastructure roadmap, showcasing Blackwell Ultra (second half 2025), Vera Rubin (new CPU, GPU, and networking in second half 2026), and Vera Rubin Ultra (second half 2027), providing clarity for long-term AI factory planning and investment.
- Pioneering Silicon Photonics: NVIDIA announced its first co-packaged optical (CPO) silicon photonic system, a 1.6 terabit per second CPO utilizing a novel micro-ring resonator modulator (MRM) technology. This innovation drastically reduces power consumption (from 30W per transceiver to co-packaged solution, saving tens of megawatts per data center) and cost for ultra-scale AI networks, enabling scaling to millions of GPUs.
- New Computing Categories for AI: Two new personal computing systems were introduced: DGX Spark, a personal server, and DGX Station, a liquid-cooled workstation (20 petaflops), designed for AI researchers and software engineers, making powerful AI development accessible.
- Isaac Groot N1 for Humanoid Robotics: NVIDIA unveiled Isaac Groot N1, a generalist foundation model for humanoid robots, which is now open-sourced. It features a dual-system architecture for "fast" and "slow" thinking, trained using synthetic data generation (Omniverse, Cosmos) and physics-based simulation.
- Newton: A New Physics Engine for Robotics: A partnership with DeepMind and Disney Research resulted in Newton, a GPU-accelerated physics engine. Designed for fine-grain rigid and soft body simulation at super-realtime speeds, Newton is crucial for training complex robotic motor skills and tactile feedback.
- Omniverse Blueprint for AI Gigafactories: NVIDIA's Omniverse platform now includes a "blueprint" for designing and optimizing AI gigafactories through digital twins, allowing engineers to simulate and optimize power, cooling, and network topologies (with partners like Cadence and Schneider Electric) before physical construction.
- Enterprise AI Acceleration: Key partnerships with companies like Cisco (for Spectrum-X integration into enterprise networking), Dell, and the entire storage industry (for GPU-accelerated, semantics-based storage) were highlighted, alongside the open-sourcing of NVIDIA Inference Microservices (NIMS) models for enterprise-ready agentic AI.
Technical Deep Dive
▶ Watch: AI's journey: from perception to generative computing (6:10)
The technical core of the keynote centered on NVIDIA's full-stack approach to powering the new waves of AI, from silicon to systems to software.
Model Architectures and Training Paradigms
The shift to agentic AI and physical AI necessitates models that can reason, plan, and understand the physical world. Agentic AI relies on advanced reasoning techniques like Chain of Thought, where models break down problems step-by-step, best-of-N consistency checking to ensure reliable answers, and path planning to explore multiple solutions. This iterative reasoning dramatically increases the number of tokens generated per query, driving the 100x computational demand.
For training these complex models, especially without human-in-the-loop limitations, Reinforcement Learning with Verifiable Results (RLVR) is a key breakthrough. By leveraging existing knowledge bases (e.g., mathematical rules, logical puzzles) or generating synthetic data, AI can be rewarded for progressively better solutions over millions of examples and trials, generating trillions of tokens in the process.
Physical AI, the foundation for robotics, requires models to understand real-world physics concepts like friction, inertia, cause and effect, and object permanence. NVIDIA addresses this through simulation.
Hardware Architectures and Systems Design
The Blackwell GPU architecture is a cornerstone. Each Blackwell GPU (B200) effectively contains two GPU dies within a single package, offering increased transistor count, memory, and bandwidth over Hopper. Key to Blackwell's efficiency is its support for new precisions like FP8 and FP4 (4-bit floating point), which significantly reduce energy consumption per operation, crucial for power-limited AI factories.
The true innovation in scaling comes with the Grace Blackwell NVLink 72 rack. This system represents an "ultimate scale-up" design, moving beyond traditional PCIe Express and InfiniBand for intra-node communication. It leverages disaggregated NVLink switches (18 switches across 9 trays) to connect 72 Grace Blackwell Superchips (each comprising a Grace CPU and a Blackwell GPU) into a single, massive GPU effectively. This allows every GPU to communicate with every other GPU at full bandwidth simultaneously, achieving 570 terabytes per second of memory bandwidth and 1 ExaFLOP of AI computation in a single liquid-cooled rack. This liquid cooling is critical for managing the 120 kilowatt power draw and compressing 600,000 components into a dense, efficient form factor.
NVIDIA's roadmap extends this scale-up vision with:
- Blackwell Ultra: An upgrade in late 2025, featuring 1.5x more flops, 1.5x more memory, and 2x more networking bandwidth.
- Vera Rubin: Arriving in late 2026, this will be an entirely new architecture with a new CPU (twice the performance of Grace, 50W), new GPU, new CX9 networking, and new HBM4 memory, supporting MVLink 144.
- Vera Rubin Ultra: Slated for late 2027, this extreme scale-up system will feature MVLink 576, a 600kW rack, 2.5 million parts, and deliver 15 ExaFLOPS of scaled-up computation and 4,600 terabytes per second of scale-up bandwidth.
Networking Innovations
Scaling out AI factories to hundreds of thousands or even millions of GPUs requires equally advanced networking. NVIDIA offers InfiniBand for ultra-low latency, high-bandwidth connections, and Spectrum-X for bringing InfiniBand-like qualities (congestion control, low latency) to Ethernet networks.
A major breakthrough is NVIDIA's co-packaged optical (CPO) silicon photonics system. Traditional pluggable optical transceivers (e.g., 30W, $1000 each) become a significant power and cost burden at massive scales (e.g., 180W and $6000 per GPU for 6 transceivers). NVIDIA's CPO system integrates optical components directly into the package, using micro-ring resonator modulators (MRM). This technology offers superior density and power efficiency compared to traditional Mach-Zehnder modulators. The photonic IC is stacked with an electronic IC, micro-lenses, and fiber arrays using TSMC's CoWoS (Chip-on-Wafer-on-Substrate) 3D packaging. This enables a 1.6 terabit per second CPO, reducing power consumption by tens of megawatts in a data center and enabling switches with an unprecedented 512 Radix (ports).
Software Stack and Ecosystem
NVIDIA's software stack is comprehensive, starting with CUDA-X libraries, which accelerate diverse scientific and industrial applications. Examples include:
- Cai numeric: Zero-change drop-in acceleration for NumPy.
- cuLitho: Computational lithography for chip manufacturing.
- Ariel: Turns GPUs into 5G radios for RAN.
- cuOpt: Mathematical optimization, now open-sourced, for supply chain, logistics, and resource planning.
- Parabricks, MONAI, Earth-2, cuQuantum, CUDA-Q, QDSS, CDF, Warp: Libraries for genomics, medical imaging, weather prediction, quantum computing, CAE, dataframes (Spark, Pandas), and physics simulation, respectively.
The NVIDIA Dynamo operating system is critical for managing the complexity of AI factory inference. It dynamically orchestrates workloads by distinguishing between prefill (flops-intensive context processing and thinking, e.g., reading a PDF) and decode (bandwidth-intensive token generation, e.g., chatbot responses). Dynamo intelligently allocates GPU resources, manages KV cache routing, and handles various parallelism and batching techniques to optimize for either low latency or high throughput depending on the workload.
For robotics and autonomous vehicles, NVIDIA Omniverse serves as the operating system for physical AI. It integrates with Cosmos for synthetic data generation, allowing developers to create infinite, grounded, and diverse virtual environments for training. Isaac Lab is used for post-training robot policies. The new Newton physics engine, a collaboration with DeepMind and Disney Research, provides GPU-accelerated, fine-grain rigid and soft body simulation, crucial for tactile feedback and fine motor skills, integrated with the Moco framework. Isaac Groot N1, the generalist foundation model for humanoid robots, uses a dual-system architecture (slow thinking for reasoning, fast thinking for precise actions) and is trained within this Omniverse simulation pipeline.
Finally, for enterprise AI, NVIDIA offers NVIDIA Inference Microservices (NIMS), open-source, enterprise-ready models that can be deployed anywhere, from new DGX Spark and DGX Station systems to cloud and on-premise servers. NVIDIA is also working with the storage industry to enable GPU-accelerated, semantics-based storage systems that embed raw data into knowledge for conversational querying.
Experimental Setup & Results
▶ Watch: Next wave: Physical AI enabling robotics (8:45)
The keynote presented compelling empirical evidence and performance comparisons to underscore the significance of NVIDIA's new architectures and software.
One of the most striking comparisons involved the performance of Blackwell against the prior-generation Hopper architecture for AI inference. Jensen Huang presented a "factory diagram" with two axes: tokens per second throughput of the entire factory (Y-axis) and tokens per second response time for an individual user (X-axis). The goal is to maximize the area under the curve, representing the balance between high-quality (smarter, faster-responding) AI and high-volume token generation.
- Hopper Baseline: A 1-megawatt data center based on Hopper GPUs could achieve roughly 100 tokens per second per user experience and a maximum factory throughput of about 2.5 million tokens per second (under highly batched, high-latency conditions).
- Blackwell Advancements:
- Blackwell with MVLink 8 (FP8): Simply upgrading to Blackwell GPUs with FP8 precision provided an immediate performance boost.
- Blackwell with FP4 Quantization: Introducing FP4 (4-bit floating point) quantization significantly improved energy efficiency, allowing more computation within the same power envelope.
- Grace Blackwell NVLink 72 Rack: The full NVLink 72 scale-up architecture delivered a substantial leap.
- Dynamo Integration: The NVIDIA Dynamo operating system further extended the performance frontier by intelligently managing workloads.
The cumulative result of these innovations is a 25x performance improvement for Blackwell over Hopper at iso-power. This means that within the same power budget (the ultimate limiter for data center revenue), Blackwell can deliver 25 times more AI inference performance.
A specific demonstration highlighted the difference between traditional LLMs and reasoning models:
- Problem: Optimally seating guests at a wedding table with complex constraints (traditions, photogenic angles, feuding families).
- Traditional LLM (Llama 3, non-reasoning): Generated an answer quickly with under 500 tokens but made mistakes, rendering the output incorrect and wasted.
- Reasoning Model (R1): Took longer, generating over 8,000 tokens (almost 9,000 tokens in the example) to reason through various scenarios, test its own answers, and produce the correct, optimal seating arrangement. This illustrated the dramatic increase in token generation and computation required for smarter, agentic AI.
For such complex reasoning models, Blackwell's performance advantage was even more pronounced: it demonstrated 40 times the performance of Hopper for these next-generation workloads.
The impact of these advancements on AI factory scale was also quantified:
- A 100-megawatt AI factory built with Hopper required 45,000 GPU dies spread across 1,400 racks to produce 300 million tokens per second.
- The same 100-megawatt factory, when built with Blackwell, required only 8,000 GPU dies in 200 racks to achieve the same 300 million tokens per second. This translates to massive reductions in physical footprint, power consumption per operation, and overall cost of ownership.
Finally, the keynote showcased the Omniverse blueprint for AI Factory digital twins. This involved simulating a 1-gigawatt AI factory using NVIDIA Omniverse, integrating 3D layout data of NVIDIA DGX SuperPODs with advanced power and cooling systems (from partners like Vertiv and Schneider Electric using EAP for power block simulation and Cadence Reality for cooling) and network topology from NVIDIA AIR. Real-time simulation allowed engineers to iterate and run large-scale "what-if" scenarios in seconds, optimizing Total Cost of Ownership (TCO) and Power Usage Effectiveness (PUE) before physical construction.
Practical Implications
▶ Watch: Three fundamental AI challenges: data, training, scaling (11:00)
The advancements outlined in Jensen Huang's GTC keynote have profound practical implications across the AI ecosystem:
For Practitioners (Developers, Data Scientists)
- Access to Superhuman AI: The rise of agentic AI means practitioners will have access to AIs that can reason, plan, and use tools, enabling them to tackle far more complex problems than ever before.
- Democratized Development: New systems like DGX Spark and DGX Station bring massive computational power to individual desks and labs, accelerating development cycles.
- Rich Software Ecosystem: The expansive CUDA-X libraries, Omniverse for simulation, Isaac Lab for robotics, and open-source NIMS models provide a comprehensive toolkit for building and deploying advanced AI solutions across diverse domains.
- Synthetic Data as a Solution: The ability to generate vast amounts of high-quality synthetic data through Omniverse and Cosmos overcomes the limitations and cost of real-world data collection, particularly for robotics and complex scenario testing.
For Infrastructure Teams (Data Center Operators, Cloud Providers)
- Shift to AI Factories: Data centers must fundamentally transform into specialized "AI factories," designed for generating tokens at scale. This requires rethinking power, cooling, and network architectures.
- Liquid Cooling is Essential: The high-density and power requirements of Grace Blackwell NVLink 72 racks (120kW) make liquid cooling a necessity, moving away from traditional air-cooled setups.
- Strategic Planning with Roadmaps: NVIDIA's annual product cadence (Blackwell Ultra, Vera Rubin, Vera Rubin Ultra) provides a predictable roadmap, allowing infrastructure teams to plan multi-year investments in land, power, and capital expenditure.
- Networking Overhaul: The demand for massive scale-out requires advanced networking like InfiniBand and Spectrum-X, with silicon photonics becoming crucial for cost and power efficiency at the largest scales (millions of GPUs).
- Digital Twins for Design: The Omniverse blueprint for AI Factories offers an indispensable tool for designing, optimizing, and operating these complex infrastructures, reducing errors and accelerating deployment time.
For Model Builders (Researchers, AI Engineers)
- Focus on Reasoning and Multimodality: The emphasis on agentic AI encourages the development of models capable of complex reasoning, planning, and integrating information from multiple modalities.
- Leverage Reinforcement Learning: RLVR with synthetic data generation becomes a primary method for training robust, intelligent models at scales unachievable with human supervision.
- Inference Optimization Criticality: Given the 100x increase in inference computation, optimizing every aspect of the inference pipeline (e.g., prefill/decode management with Dynamo, parallelism, batching, quantization like FP4) is paramount for both performance and cost-effectiveness.
- Generalist Foundation Models: The introduction of Isaac Groot N1 signals a future where generalist models can be post-trained for a variety of robotic embodiments and tasks, accelerating robotics development.
For Deployers (Enterprise IT, Edge Computing)
- Enterprise AI Transformation: Partnerships with companies like Cisco and the storage industry mean that enterprise IT will soon have access to GPU-accelerated infrastructure and open-source, enterprise-ready NIMS models for deploying AI agents and advanced analytics on-prem or in hybrid clouds.
- AI at the Edge: The integration of AI into 5G radio networks (Ariel) and autonomous vehicles (GM partnership, Halos safety framework) demonstrates the expansion of AI to the edge, revolutionizing communications and transportation.
- Tradeoffs and Limitations: Deployers must carefully balance the tradeoff between AI quality (smarter, faster responses per user) and throughput (total tokens generated by the factory). Power consumption is highlighted as the ultimate limiter for data center revenue, making energy-efficient architectures and software crucial. The sheer scale and complexity of building and operating these AI factories represent significant challenges in terms of capital investment, engineering expertise, and operational management.
Key Takeaways
- AI's Inflection Point: AI is undergoing a profound platform shift, evolving into agentic and physical forms that demand exponentially more computation (100x increase for reasoning AI) and are driving the transformation of data centers into "AI factories."
- Blackwell's Performance Leap: NVIDIA's Blackwell architecture, particularly the Grace Blackwell NVLink 72 rack, delivers unprecedented scale-up, offering 25x performance over Hopper at iso-power for inference and up to 40x for complex reasoning models, enabling 1 ExaFLOP in a single liquid-cooled rack.
- The AI Factory Operating System: NVIDIA Dynamo is introduced as the open-source operating system for AI factories, intelligently managing complex inference workloads (prefill/decode, parallelism, batching) to maximize efficiency and quality of service.
- Full-Stack Strategy Across All Domains: NVIDIA is extending its full-stack approach (chips, systems, interconnects, software) across cloud, enterprise (with NIMS and partnerships like Cisco), and robotics (with Omniverse, Cosmos, Newton physics engine, and Isaac Groot N1 generalist robot model).
- Innovation for Scale-Out: Breakthroughs like NVIDIA's co-packaged optical silicon photonics (1.6 Tb/s CPO with micro-ring resonator modulators) are critical for cost-effectively scaling AI networks to hundreds of thousands or millions of GPUs while drastically reducing power consumption.
- Predictable Annual Roadmap & Digital Twins: NVIDIA's commitment to an annual architecture roadmap (Blackwell Ultra, Vera Rubin, Vera Rubin Ultra) allows for long-term planning of massive AI infrastructure investments, further aided by Omniverse blueprints for digital twin-based AI factory design and optimization.
About the Speaker(s)
The keynote was delivered by Jensen Huang, the visionary founder and CEO of NVIDIA. With over 33 years at the helm, Huang has consistently guided NVIDIA at the forefront of computing innovation, from revolutionizing graphics with GeForce to enabling the modern AI revolution with CUDA. He is deeply passionate about the transformative power of accelerated computing and its impact on scientific discovery and human progress, often sharing anecdotes about how NVIDIA's work enables scientists to achieve their life's work within their lifetime. Huang's presentations are known for their technical depth, his enthusiastic delivery, and his ability to articulate complex technological shifts with clarity and vision, making GTC a highly anticipated event for the AI and computing community.
Reviews
Simon Wisk (Open Source Developer & AI Tooling Expert) — WEAK
This is a product keynote, not a conference talk, and reviewing it as engineering content is like reviewing a car commercial for its mechanical accuracy. Jensen Huang is a great presenter and NVIDIA is shipping genuinely impressive hardware, but nothing in this 'talk' constitutes engineering insight an engineer can act on. It's a roadmap announcement wrapped in vision language, and the article summarizing it is essentially a cleaned-up press release.
Jensen Hitch (AI Compute Platform CEO) — MUST SEE
Jensen Huang's GTC March 2025 keynote is not a product announcement — it is a systems-level thesis about what AI infrastructure looks like when reasoning becomes the primary workload. The 100x compute demand claim is not marketing; it follows directly from token economics and Chain-of-Thought scaling, and every architectural decision presented — NVLink 72, FP4, Dynamo's prefill/decode disaggregation, co-packaged silicon photonics — traces back to a specific physical constraint that number creates. The full-stack coherence here, from transistor precision to data center power topology to robotics simulation pipelines, is what separates a platform shift from a benchmark press release.