NVIDIA CEO Jensen Huang Live GTC Paris Keynote at VivaTech 2025

Jensen Huang (Founder and CEO · NVIDIA)

NVIDIA GTC 2025 · Keynote

Overview

NVIDIA CEO Jensen Huang's keynote at GTC Paris, delivered live at VivaTech 2025, unveiled a sweeping vision for the future of artificial intelligence, positioning it not merely as a technology but as the bedrock of a new industrial revolution. The talk centered on the concept of AI factories – specialized data centers designed to generate "intelligent tokens" – and introduced the Grace Blackwell computing platform as the foundational "thinking machine" for this era. Huang detailed how NVIDIA's full-stack approach, from silicon to software, is enabling breakthroughs in agentic AI, embodied AI (robotics), and industrial AI through digital twins. The keynote emphasized Europe's critical role in this transformation, highlighting significant investments in AI infrastructure and partnerships across the continent.

Watch on YouTube

Visual summary for NVIDIA CEO Jensen Huang Live GTC Paris Keynote at VivaTech 2025 by Jensen Huang
Visual summary for NVIDIA CEO Jensen Huang Live GTC Paris Keynote at VivaTech 2025 by Jensen Huang

Key moments

  1. 0:00 Opening cinematic and Jensen Huang's welcome
  2. 4:26 NVIDIA's mission: Accelerated computing platform
  3. 5:40 Diverse applications accelerated by NVIDIA libraries
  4. 8:20 Unveiling CUDA Q for quantum computing
  5. 9:40 Quantum computing is reaching an inflection point
  6. 10:40 Future supercomputers: QPU and GPU integration
  7. 11:40 CUDA Q accelerated on Grace Blackwell 200

NVIDIA CEO Jensen Huang Live GTC Paris Keynote at VivaTech 2025

Speakers: NVIDIA (NVIDIA)

Conference: NVIDIA GTC

YouTube: https://www.youtube.com/watch?v=X9cHONwKkn4

Overview

NVIDIA CEO Jensen Huang's keynote at GTC Paris, delivered live at VivaTech 2025, unveiled a sweeping vision for the future of artificial intelligence, positioning it not merely as a technology but as the bedrock of a new industrial revolution. The talk centered on the concept of AI factories – specialized data centers designed to generate "intelligent tokens" – and introduced the Grace Blackwell computing platform as the foundational "thinking machine" for this era. Huang detailed how NVIDIA's full-stack approach, from silicon to software, is enabling breakthroughs in agentic AI, embodied AI (robotics), and industrial AI through digital twins. The keynote emphasized Europe's critical role in this transformation, highlighting significant investments in AI infrastructure and partnerships across the continent.

Huang articulated that AI is fundamentally reshaping industries and national infrastructures, mirroring the impact of electricity and the internet. The sheer computational demands of next-generation AI, particularly reasoning models and multi-agent systems, necessitate a dramatic leap in computing power, which the Blackwell architecture is designed to deliver. Beyond raw performance, the talk showcased NVIDIA's commitment to democratizing AI development and deployment through platforms like NVIDIA NIM and DGX Cloud Lepton, fostering a vibrant ecosystem for both proprietary and open-source models.

The significance of this talk extends beyond technological announcements; it's a strategic declaration of NVIDIA's intent to power the global AI infrastructure. By framing AI data centers as "factories" and national assets, Huang underscored the profound economic and societal implications. The emphasis on simulation in NVIDIA Omniverse for training robots and optimizing industrial processes, coupled with the introduction of new enterprise-grade AI systems, illustrates a comprehensive strategy to integrate advanced AI into every facet of the physical and digital world.

Background

▶ Watch: Opening cinematic and Jensen Huang's welcome (0:00)

NVIDIA's journey began with a mission to create a new computing platform capable of accelerating tasks beyond the reach of traditional CPUs. This led to the development of accelerated computing and the CUDA programming model, initially applied to molecular dynamics and, famously, computer graphics via GeForce GPUs. Over the years, NVIDIA built a vast ecosystem of over 400 CUDA-X libraries, each accelerating specific domains from computational lithography (cuLitho) and genomics (Parabricks) to medical imaging (MONAI) and classical machine learning (cuDF, cuML). This library-driven approach was crucial because accelerating applications requires not just faster processors, but a fundamental reformulation of algorithms for highly parallel execution.

The talk highlighted the evolution of AI itself, tracing its trajectory from the "AlexNet big bang" of 2012, which ushered in the era of deep learning and perception AI. This was followed by the generative AI revolution of the past five years, characterized by multimodal models capable of generating text, images, and other content. Huang then introduced the "new wave of AI": agentic AI and embodied AI. Agentic AI, he explained, focuses on the fundamental cycles of intelligence – perception, reasoning, planning, and execution – allowing AI to break down complex problems, learn, use tools, and solve novel tasks. Embodied AI, or robotics, represents the physical manifestation of this intelligence, enabling AI to generate local motion, walk, grab, and use tools in the real world.

A significant challenge in scaling AI has been the limitations of Moore's Law, which provides only a ~2x performance increase every 3-5 years. The burgeoning complexity of AI models, particularly reasoning models that engage in internal monologues and iterative problem-solving (e.g., chain-of-thought, tree-of-thought, self-reflection), demands far greater computational power. This necessitated a paradigm shift beyond incremental improvements, leading to the development of the Grace Blackwell architecture. Furthermore, the keynote touched upon the emerging field of quantum computing, noting its inflection point with the demonstration of logical qubits and envisioning a future where QPUs (Quantum Processing Units) collaborate with GPUs for pre-processing, control, and error correction in quantum-classical computing.

Key Findings

▶ Watch: Diverse applications accelerated by NVIDIA libraries (5:40)

The keynote unveiled several pivotal advancements and strategic directions:

  • Grace Blackwell (GB200/GB300) as the "Thinking Machine": NVIDIA's latest architecture delivers a 30-40x performance leap over the previous Hopper generation for reasoning models, far surpassing Moore's Law. This is achieved by designing it as "one giant virtual GPU" capable of supporting complex, multi-step AI reasoning.
  • Revolutionary MVLink Spine: This custom interconnect provides an unprecedented 130 terabytes per second of all-to-all bandwidth across 144 Blackwell dies (72 packages) within a single rack, exceeding the world's entire internet peak traffic. It transforms scaling up computing into a coherent, high-bandwidth fabric.
  • Mass Production of AI Supercomputers: NVIDIA is now producing 1,000 GB200 systems per week, each weighing nearly two tons and comprising 1.2 million components. A single GB200 rack is more performant than the 2018 Sierra supercomputer (which consumed 10 MW, compared to GB200's 100 kW).
  • Agentic AI Platform: NVIDIA introduced a comprehensive framework for building and deploying AI agents, leveraging NVIDIA Nemo Neotron LLMs, Nemo Retriever (multimodal search), and a general agent blueprint called AIQ. This platform provides tools for onboarding, curating data, evaluating, guard-railing, and deploying agents.
  • Embodied AI and Robotics Acceleration: The NVIDIA Thor computer and a specialized operating system, combined with transformer models, enable robots to learn complex local motion and articulation. NVIDIA Omniverse provides the photorealistic, physics-abiding simulation environment for training these robots with vast amounts of synthetic data.
  • Industrial AI Revolution via Digital Twins: Leveraging Omniverse, industries are building digital twins of factories, warehouses, and even fusion reactors for design, planning, optimization, and operation. This allows for simulation and refinement before physical implementation.
  • NVIDIA NIM and DGX Cloud Lepton for Democratized AI: NIM (NVIDIA Inference Microservice) packages enhanced open-source models (like Llama, Mistral) with reasoning capabilities and extended context, available for download. DGX Cloud Lepton provides a unified "cloud of clouds" platform for deploying NIMs and AI agents across public clouds, regional clouds, private clouds, and even local workstations with one-click integration with platforms like Hugging Face.
  • Quantum Computing Integration (CUDA Q): Announced that the entire quantum algorithm stack is accelerated on Grace Blackwell 200, enabling both classical simulation of qubits and quantum-classical collaboration with QPUs.

Technical Deep Dive

▶ Watch: Unveiling CUDA Q for quantum computing (8:20)

The technical core of the keynote revolved around the Grace Blackwell architecture and its role in powering the next generation of AI. Huang emphasized that the GB200 and GB300 systems are designed as "thinking machines," architected to handle the burgeoning demands of reasoning models. Unlike traditional one-shot chatbots, these models engage in complex internal processes like chain of thought, tree of thought, and self-reflection, requiring significantly more compute to generate thousands or even tens of thousands of tokens per query.

At the heart of Blackwell is the concept of a "one giant virtual GPU." To achieve this, NVIDIA disaggregated the traditional GPU node, integrating two Grace CPUs directly connected to four Blackwell GPUs within a single liquid-cooled compute tray. These trays are then interconnected using the groundbreaking MVLink system. MVLink is described as a memory semantics interconnect and a compute fabric, not merely a network. The centerpiece is the MVLink spine, a custom blind-mated backplane utilizing 5,000 copper cables to directly connect 144 Blackwell GPU dies (packaged into 72 units) across an entire rack. This spine delivers an astonishing 130 terabytes per second of all-to-all bandwidth, ensuring that all GPUs can communicate simultaneously without blocking, effectively shrinking a supercomputer into a 60-pound chassis. The manufacturing process itself is an engineering marvel, involving chip-on-wafer-on-substrate technology, attaching 32 Blackwell dies and 128 HBM (High Bandwidth Memory) stacks on a custom silicon interposer. The system also integrates ConnectX7 Super NICs for scale-out communications and Bluefield 3 DPUs to offload and accelerate networking, storage, and security tasks.

For agentic AI, NVIDIA offers a comprehensive software stack built around NVIDIA Nemo. This includes Nemo Neotron, which consists of world-class reasoning Large Language Models (LLMs). These models are often enhanced versions of open-source models (e.g., Llama, Mistral) that undergo extensive post-training, neural architecture search, reinforcement learning, and context extension to improve performance and reasoning capabilities. Nemo Retriever provides multimodal semantic search capabilities. The entire toolkit, integrated into the AI ops ecosystem, supports the full lifecycle of an AI agent, from initial onboarding and data curation to evaluation, guard-railing, and continuous deployment. This modular approach allows for the creation of systems of specialized AI models, rather than a single monolithic AI, enabling greater flexibility and continuous integration/continuous delivery (CI/CD).

Embodied AI and robotics are powered by the NVIDIA Thor computer, a self-contained dev kit featuring the Thor chip. This platform runs a specialized operating system for robotics and leverages transformer models that take sensor inputs and instructions to generate path plans and motor controls for complex articulation (arms, fingers, legs). A critical component for training these robots is NVIDIA Omniverse, a simulation environment that generates vast amounts of diverse, realistic synthetic data. Omniverse is designed to be photorealistic and strictly obey the laws of physics, which is crucial for training robots whose perception systems rely on photons and whose physical interactions demand accurate simulation. The example of "Greck," a humanoid robot learning to walk and manipulate in a virtual Omniverse environment before deployment in the physical world, powerfully illustrated this capability. The NVIDIA Drive platform for autonomous vehicles, featuring the Halo safety system, also exemplifies this full-stack approach, encompassing chip architecture, system design, a functional-safe operating system, and AI models trained on diverse real-world and synthetic data.

Finally, CUDA Q represents NVIDIA's foray into quantum computing. It's presented as an extension of CUDA for quantum-classical computing based on GPUs. This involves using classical GPUs for simulating qubits and quantum algorithms, as well as for the computationally intensive pre-processing, control, and error correction required when QPUs are integrated into supercomputers. The announcement of CUDA Q's availability for Grace Blackwell signifies NVIDIA's readiness to support the emerging quantum computing ecosystem.

Experimental Setup & Results

▶ Watch: Future supercomputers: QPU and GPU integration (10:40)

The keynote highlighted several impressive benchmarks and production capabilities demonstrating the efficacy of NVIDIA's new platforms:

  • Blackwell vs. Hopper Performance: The Grace Blackwell architecture delivers a 30-40 times performance increase for reasoning models compared to the previous Hopper generation. This significant leap is attributed to Blackwell's design as a "thinking machine" optimized for multi-step reasoning, which generates substantially more tokens per query than traditional generative AI.
  • MVLink Spine Bandwidth: The custom MVLink spine within a Blackwell rack achieves an unprecedented 130 terabytes per second (TB/s) of all-to-all bandwidth. Huang emphasized that this figure is "more than the data rate of the peak traffic of the world's entire internet traffic on this backplane," underscoring its capability to unify a massive number of GPUs into a single virtual entity.
  • GB200 Production Scale: NVIDIA is currently mass-producing 1,000 GB200 systems per week. Each system is a "two-ton," "1.2 million part" supercomputer, consuming 100 kilowatts (kW) of power. A single GB200 rack is stated to be more performant than the entire Sierra supercomputer from 2018, which consumed 10 megawatts (MW). This scale of production for high-performance supercomputers is unprecedented.
  • Agentic AI Token Generation: Huang noted that a single prompt into an AI agent designed to solve a complex problem (e.g., starting a food truck) could generate 10,000 times more tokens than a similar prompt in an original chatbot. This explosion in token generation directly necessitates the massive performance improvements offered by Blackwell.
  • Nemo Neotron Performance: NVIDIA's post-trained, enhanced open models, branded as Nemo Neotron, consistently achieve "top of the leaderboard" performance in benchmarks. Specific examples mentioned were significant improvements over Llama 8B, 70B, and 405B models, demonstrating the effectiveness of NVIDIA's optimization techniques including neural architecture search, reinforcement learning, and context extension.
  • Autonomous Vehicle (AV) Team Success: NVIDIA's AV team has won the end-to-end self-driving car challenge at CVPR two years in a row, showcasing the robustness and effectiveness of their NVIDIA Drive platform and Halo safety system. This indicates a high level of practical performance and reliability in a safety-critical domain.
  • DGX Cloud Lepton Deployment: The platform enables on-demand access to a global network of GPUs across various clouds (Lambda, AWS, GCP, Yoda, Nebus) and local systems (DGX Station, DGX Spark). It offers fast provisioning, scalable node deployment, real-time monitoring of GPU performance and model convergence, and one-click deployment of NIM endpoints or models across multiple clouds/regions for distributed inference.

These results collectively underscore the dramatic advancements in AI compute, the scalability of NVIDIA's hardware and software stack, and the practical capabilities of their AI models and agents.

Practical Implications

▶ Watch: CUDA Q accelerated on Grace Blackwell 200 (11:40)

The implications of Jensen Huang's keynote are far-reaching, impacting practitioners, infrastructure teams, model builders, and deployers across various industries.

For infrastructure teams and data center operators, the concept of AI factories is transformative. These are no longer passive storage facilities but active, revenue-generating manufacturing plants for "intelligent tokens." This requires a fundamental rethinking of data center design, scaling, orchestration, and provisioning. The sheer scale and power density of GB200 systems (two tons each, 100 kW per rack) necessitate advanced liquid cooling solutions and robust supply chains for mass production. Integrating these "AI-native" systems into traditional enterprise IT environments (x86, Linux, VMware, Red Hat, Dell EMC, etc.) is a critical challenge, addressed by new offerings like the RTX Pro server and partnerships with IT providers.

Model builders and AI developers gain unprecedented capabilities. The Blackwell architecture's 30-40x performance leap for reasoning models enables the creation of far more sophisticated agentic AI systems that can plan, reason, and self-reflect, moving beyond simple generative tasks. The NVIDIA Nemo framework, with its Neotron LLMs and comprehensive toolkit, democratizes access to enhanced, state-of-the-art open models, allowing companies to fine-tune and adapt them with their proprietary data for regional languages and specific enterprise use cases. This supports data sovereignty and cultural relevance.

Practitioners and deployers benefit significantly from DGX Cloud Lepton, a "cloud of clouds" platform that simplifies the deployment of AI agents and NIMs across diverse environments – public, regional, private, and edge. This unified deployment strategy reduces complexity and enables seamless scaling, allowing developers to focus on building AI rather than managing infrastructure. The integration with platforms like Hugging Face further streamlines the model training and deployment pipeline.

The rise of industrial AI and digital twins (powered by Omniverse) means that industries can design, plan, optimize, and operate complex systems (factories, warehouses, even fusion reactors) entirely in a virtual environment before physical implementation. This reduces costs, accelerates innovation, and improves efficiency. For robotics, the ability to train embodied AI in photorealistic, physics-abiding virtual worlds addresses the massive data requirements and programming complexities that have historically hindered widespread robot adoption. The NVIDIA Thor platform promises to make robots more adaptable and teachable, potentially addressing global labor shortages.

Tradeoffs and Limitations include the high initial investment in these advanced AI infrastructures, the complexity of integrating liquid-cooled, high-density systems, and the ongoing need for specialized talent to manage and develop AI solutions. While NVIDIA is working to ease deployment, the sheer scale of the technology still presents a barrier to entry for smaller entities. However, the availability of desktop-scale DGX Spark and enterprise RTX Pro servers aims to bring powerful AI capabilities closer to developers and traditional IT environments. Safety and reliability, particularly for autonomous vehicles (NVIDIA Drive's Halo system) and industrial applications, remain paramount and are deeply integrated into NVIDIA's full-stack approach. The vision of AI as a national infrastructure emphasizes the strategic importance for countries to invest in indigenous AI capabilities, potentially leading to geopolitical considerations and competition.

Key Takeaways

  • AI Factories are the New Infrastructure: Data centers are evolving into "AI factories" designed to generate intelligent tokens, becoming critical revenue-generating infrastructure for nations and companies, akin to electricity or the internet.
  • Blackwell: The Era of Thinking Machines: NVIDIA's Grace Blackwell architecture (GB200/GB300) delivers a 30-40x performance leap for complex reasoning AI models, designed as a "one giant virtual GPU" with revolutionary MVLink interconnects (130 TB/s bandwidth).
  • Agentic AI for Problem Solving: The next wave of AI focuses on agentic capabilities – perception, reasoning, planning, and execution – enabling AI to break down problems, use tools, and learn, powered by NVIDIA's Nemo platform and enhanced open models (NIM).
  • Embodied AI and Digital Twins Drive Robotics: NVIDIA Omniverse provides photorealistic, physics-accurate simulation environments for training embodied AI and robots, addressing data challenges and accelerating the development of humanoids and autonomous systems (like the Thor platform and Halo safety system for AVs).
  • Democratization of AI Development & Deployment: NVIDIA NIM offers enhanced, downloadable AI models, while DGX Cloud Lepton provides a unified, multi-cloud platform for deploying AI agents and models, making advanced AI accessible across diverse compute environments, from public clouds to desktop workstations.
  • Europe at the Forefront of Industrial AI: Europe is making significant investments in AI infrastructure, with NVIDIA partnering to build AI factories and an industrial AI cloud, leveraging the continent's strong industrial heritage to drive a new industrial revolution through digital twins and AI integration.

About the Speaker(s)

Jensen Huang is the founder and CEO of NVIDIA, a company he co-founded in 1993. Throughout the keynote, Huang demonstrated his deep technical understanding and visionary leadership, articulating NVIDIA's long-standing commitment to accelerated computing and its evolution into the AI era. He highlighted the company's journey from creating GPUs for computer graphics to powering the global AI revolution, emphasizing the importance of a full-stack approach from silicon to software. Huang's passion for innovation was evident as he showcased groundbreaking technologies like the Grace Blackwell architecture, agentic AI platforms, and embodied AI in robotics, consistently framing these advancements within the context of a new industrial revolution. His engagement with European partners and governments underscored his belief in the global impact of AI and NVIDIA's role in building indigenous AI infrastructure worldwide.

Reviews

Simon Wisk (Open Source Developer & AI Tooling Expert) — PASS

A Jensen Huang keynote is a product launch event dressed up as a conference talk. This one follows the formula: enormous benchmark numbers, sweeping historical analogies, and a parade of partnership logos. There is no engineering here — just marketing copy with silicon branding. Nothing in this article would help an engineer build, design, or evaluate anything.

Jensen Hitch (AI Compute Platform CEO) — STRONG ACCEPT

Huang delivers a coherent, systems-level argument for why the AI compute transition is a platform shift on the order of electricity — not an incremental benchmark improvement. The MVLink spine as a memory-semantics fabric, the framing of AI factories as industrial output units, and the full-stack reasoning from silicon through deployment to national infrastructure strategy all hold up under scrutiny. The 30-40x reasoning model performance claim is grounded in a real architectural insight — reasoning models generate orders of magnitude more tokens, so the correct unit of measure is tokens-per-watt at production scale, not FLOP counts — and Blackwell is explicitly designed around that…

→ Top-rated talks at NVIDIA GTC 2025

All talks from NVIDIA GTC 2025