NVIDIA GTC Washington, D.C. Keynote with CEO Jensen Huang
Jensen Huang (Founder and CEO · NVIDIA)
NVIDIA GTC 2025 · Keynote
Overview
Jensen Huang's keynote at NVIDIA GTC Washington, D.C., delivered a sweeping vision of the future of computing, positioning artificial intelligence and accelerated computing as the twin pillars of a new industrial revolution. The talk emphasized NVIDIA's foundational role in this transformation, from inventing the GPU and the CUDA programming model to pioneering AI factories and digital twins. Huang articulated how the convergence of these technologies is not merely enhancing existing industries but creating entirely new ones, driving an unprecedented demand for computational power that traditional computing models can no longer satisfy.

Key moments
- 0:00 Opening montage: America's innovation and tech history
- 2:20 AI: The new industrial revolution, powered by Nvidia
- 4:18 Jensen Huang takes the stage, welcomes to GTC
- 6:00 Nvidia's new computing model: accelerated computing introduced
- 6:35 The end of Moore's Law and the need for acceleration
- 7:00 GPU and CUDA: Foundations of accelerated computing
- 8:05 Nvidia's 'treasure': The CUDA programming model and ecosystem
- 9:00 Key NVIDIA libraries: CU litho, cuDNN, Megatron Core, Monai
NVIDIA GTC Washington, D.C. Keynote with CEO Jensen Huang
Speakers: Jensen Huang, CEO, NVIDIA
Conference: NVIDIA GTC
YouTube: https://www.youtube.com/watch?v=lQHK61IDFH4
Overview
Jensen Huang's keynote at NVIDIA GTC Washington, D.C., delivered a sweeping vision of the future of computing, positioning artificial intelligence and accelerated computing as the twin pillars of a new industrial revolution. The talk emphasized NVIDIA's foundational role in this transformation, from inventing the GPU and the CUDA programming model to pioneering AI factories and digital twins. Huang articulated how the convergence of these technologies is not merely enhancing existing industries but creating entirely new ones, driving an unprecedented demand for computational power that traditional computing models can no longer satisfy.
The keynote, set against a backdrop of American innovation and re-industrialization, underscored the strategic importance of leading in AI and advanced computing for national security and economic prosperity. Huang unveiled a series of groundbreaking advancements, including new platforms for 6G telecommunications, quantum computing, enterprise AI, robotics, and autonomous vehicles. His address was a powerful declaration of NVIDIA's commitment to extreme co-design across the entire computing stack—from chips to systems, software, and applications—to unlock exponential performance gains and make AI universally accessible and cost-effective.
This talk matters immensely as it outlines NVIDIA's comprehensive strategy to meet the escalating demands of the AI era, characterized by what Huang termed "two exponentials": the increasing computational requirements of AI models and the exponential growth in their usage as they become smarter and more valuable. It highlights the shift from general-purpose computing to specialized, accelerated architectures, and from software as "tools" to AI as "workers" capable of performing complex tasks. The keynote serves as a roadmap for how NVIDIA, in collaboration with a vast ecosystem of partners, intends to build the foundational infrastructure for the next generation of global innovation.
Background
▶ Watch: Opening montage: America's innovation and tech history (0:00)
The genesis of NVIDIA's vision for accelerated computing stems from a fundamental challenge to the traditional trajectory of computing: the slowdown of Moore's Law and the cessation of Dennard scaling nearly a decade ago. For decades, chip performance and power efficiency improved predictably, allowing general-purpose CPUs to handle increasingly complex tasks. However, as physical limits are approached, simply adding more transistors no longer yields exponential performance gains or power efficiency improvements. This reality, observed by NVIDIA 30 years ago, necessitated a radical rethinking of computing architecture.
NVIDIA's response was the invention of the GPU and the CUDA parallel programming model. This approach recognized that many computational problems, particularly those involving massive datasets and parallelizable operations, could be vastly accelerated by offloading them from sequential CPUs to highly parallel GPUs. However, this required a fundamentally different programming paradigm, demanding the reinvention of algorithms, creation of new libraries, and rewriting of applications. This 30-year dedication led to the development of over 350 CUDA X libraries, spanning diverse domains from computational lithography (CU Litho) and sparse solvers to data analytics (QDF) and deep learning (cuDNN, Megatron Core). These libraries effectively made accelerated computing accessible and productive, laying the groundwork for the AI revolution.
The emergence of AI, particularly large language models and generative AI, has further intensified the demand for accelerated computing. AI is not just a new application; it's a "new industrial revolution" that reinvents the entire computing stack. Unlike traditional software, which involves hand-coding on CPUs, AI is characterized by machine learning—intensive training on data running on GPUs. This shift requires a new type of infrastructure, a new approach to software development, and a re-evaluation of how computational resources are deployed and managed. The problem this talk addresses is how to sustain the exponential growth of AI capabilities and usage in an era where conventional scaling has ended, and how to harness this power for broad societal and economic benefit.
Key Findings
▶ Watch: Jensen Huang takes the stage, welcomes to GTC (4:18)
The keynote delivered several pivotal findings and announcements that collectively paint a picture of NVIDIA's strategy for the AI era:
- AI as a New Computing Model and Industry: AI is fundamentally reinventing the computing stack, moving beyond traditional hand-coded software on CPUs to machine learning on GPUs. It's not just a tool but "work" itself, with AI agents capable of using tools and performing tasks, thus addressing a much larger segment of the global economy (the "hundred trillion dollar economy" beneath the "trillion dollar IT tools industry"). This necessitates the creation of AI factories—specialized data centers designed to produce valuable "tokens" (the computational unit of AI) at high rates and low costs.
- Two Exponentials Driving Demand: The AI industry is experiencing two compounding exponential demands:
- Scaling Laws: The continuous need for more computation to enable pre-training, post-training (skills and reasoning), and thinking (inference) in AI models. Huang noted that "thinking is hard" and requires extraordinary computational load.
- Smarter Models, More Usage: As AI models become more intelligent, grounded, and capable of complex reasoning and problem-solving, their utility increases, leading to exponential growth in user adoption and, consequently, even greater demand for compute.
- Extreme Co-Design as the Solution to Moore's Law Limits: With Moore's Law largely ended, sustaining exponential performance gains requires extreme co-design—a holistic approach that re-architects everything from the ground up: new fundamental computer architectures, chips, systems, software, model architectures, and applications. This allows for compounding performance benefits far beyond the incremental gains of traditional chip design.
- Grace Blackwell (GB200) and Vera Rubin Architectures: NVIDIA unveiled the Grace Blackwell MVLink72 as the most extreme co-designed computer in modern times, delivering 10x generational performance over the H200 (NVIDIA's previous generation) for AI inference. This system achieves the lowest cost per token generation globally. The subsequent Vera Rubin generation, already in the lab, represents further advancements in rack-scale, cableless, liquid-cooled AI supercomputing, promising even greater performance and efficiency.
- New Platforms for Strategic Industries:
- NVIDIA Arc for 6G: A new product line for telecommunications, combining Grace CPU, Blackwell GPU, and ConnectX networking to enable software-defined, programmable wireless communication with integrated AI processing.
- NVQLink for Quantum Computing: An interconnect architecture directly connecting quantum processors (QPUs) with NVIDIA GPUs, leveraging the CUDAQ platform for quantum error correction, calibration, control, and hybrid simulations.
- Omniverse DSX for AI Factories: A digital twin blueprint for designing, optimizing, and operating gigascale AI factories, integrating building, power, and cooling with NVIDIA's AI infrastructure stack.
- Drive Hyperion for Robo-Taxis: A standardized, sensor-rich computing platform on wheels for autonomous vehicles, enabling the deployment of AI chauffeurs.
- Commitment to Open-Source AI and Ecosystem Expansion: NVIDIA is now a leading contributor to open-source AI models across language, physical AI, and biology domains, recognizing their critical role for startups, researchers, and national innovation. The company continues to deeply integrate its CUDA X libraries and AI models into major cloud platforms (AWS, Google Cloud, Microsoft Azure, Oracle) and enterprise SaaS solutions (ServiceNow, SAP, Synopsys, Cadence, CrowdStrike, Palantir).
- Re-industrialization and Manufacturing in America: NVIDIA is bringing manufacturing back to the U.S., with full production of Blackwell systems in Arizona and future AI factories (like the one with Foxconn in Texas) being built domestically, emphasizing national security and job creation.
Technical Deep Dive
▶ Watch: The end of Moore's Law and the need for acceleration (6:35)
NVIDIA's keynote delved into the intricate technical underpinnings of its accelerated computing and AI ecosystem, showcasing innovations across hardware, software, and systems design.
Accelerated Computing Foundation: CUDA X Libraries
At the core of NVIDIA's long-term strategy is CUDA, a parallel computing platform and programming model. Huang highlighted its 30-year evolution, ensuring compatibility across generations (now CUDA 13, with CUDA 14 upcoming) to foster developer adoption. Built atop CUDA are 350+ CUDA X libraries, each meticulously redesigned for accelerated computing. Examples include:
- CU Litho: Computational lithography, critical for chip manufacturing (used by TSMC, Samsung, ASML).
- Sparse solvers: For CAE (Computer-Aided Engineering) applications.
- Co-op: Numerical optimization, achieving record-breaking performance in problems like the traveling salesperson problem.
- Warp Python solver: For CUDA-accelerated simulation.
- QDF: A data frame approach accelerating SQL data frame databases.
- cuDNN and Megatron Core: Foundational libraries for deep neural networks and training extremely large language models.
- Monai: The leading medical imaging AI framework.
- Ariel: Genomics processing.
- cuQuantum: For quantum computing.
These libraries are the "treasure of our company," enabling ecosystem partners and opening new markets.
NVIDIA Arc: Powering 6G with AI
The NVIDIA Arc (Aerial Radio Network Computer) is a revolutionary platform for 6G telecommunications. It integrates three key NVIDIA technologies:
- Grace CPU: NVIDIA's Arm-based data center CPU.
- Blackwell GPU: The latest generation of NVIDIA's AI-accelerated GPUs.
- ConnectX (Mellanox) Networking: High-performance, low-latency interconnects.
These components collectively run the Aerial CUDA X library, a software-defined wireless communication system. Arc enables AI for RAN (Radio Access Network), using AI (e.g., reinforcement learning) to improve spectral efficiency by dynamically adjusting beamforming based on real-time context (traffic, mobility, weather). This can significantly reduce the 1.5-2% of global power consumed by spectral inefficiency. Furthermore, AI on RAN transforms the wireless network into an edge industrial robotics cloud, extending cloud computing capabilities directly to base stations for new AI applications at the edge.
NVQLink and CUDAQ: Hybrid Quantum-Classical Computing
NVIDIA introduced NVQLink, a new interconnect architecture that directly links quantum processors (QPUs) with NVIDIA GPUs. This is crucial for overcoming the fragility of qubits, which remain stable for only a few hundred operations while meaningful problems require trillions. NVQLink facilitates quantum error correction by moving terabytes of data between quantum hardware and GPUs thousands of times per second.
The underlying software platform is CUDAQ, an extension of CUDA designed for hybrid quantum-GPU computing. CUDAQ allows QPUs and GPUs to work collaboratively, with computation moving back and forth within microseconds—the essential latency for quantum-classical cooperation. This architecture is scalable, designed to support future quantum computers with tens to hundreds of thousands of qubits for control, co-simulation, and error correction.
The AI Factory and Extreme Co-Design
Huang introduced the concept of an AI factory—a specialized data center whose sole purpose is to produce "tokens" (the fundamental computational unit of AI, representing everything from words and images to 3D structures, chemicals, and robot actions) that are valuable, generated at incredible rates, and cost-effectively. Unlike general-purpose data centers, AI factories are purpose-built.
To achieve this, NVIDIA employs extreme co-design, a holistic approach spanning the entire computing stack:
- New computer architecture: Designing systems from the ground up.
- New chips: Developing specialized processors like Blackwell, Grace, and Bluefield.
- New systems: Integrating these chips into rack-scale supercomputers.
- New software: Optimizing CUDA, CUDA X libraries, and AI frameworks.
- New model architectures: Understanding future AI model needs.
- New applications: Developing agentic AI systems.
This integrated approach enables compounding exponential performance gains, far exceeding traditional generational improvements.
Grace Blackwell (GB200) and Vera Rubin Architectures
The Grace Blackwell (GB200) MVLink72 is a prime example of extreme co-design. It's a rack-scale computer where 72 Blackwell GPUs and Grace CPUs are interconnected into a single, giant virtual GPU fabric using NVLink 72. This allows all chips to work together as one, enabling a single GB200 superchip to handle four "experts" within a multi-trillion parameter AI model, compared to 32 experts per GPU in previous generations due to interconnect limitations. This architectural shift significantly boosts inference speed and token generation efficiency.
For scale-out beyond a single rack, NVIDIA uses Spectrum Ethernet (Spectrum X), an AI-optimized Ethernet technology, and Spectrum X Gigascale (XGS) to connect multiple data centers.
The upcoming Vera Rubin generation represents the third iteration of this rack-scale design, aiming for even greater integration and performance. It features a completely cableless, 100% liquid-cooled design. The Vera Rubin compute tray includes a Vera Rubin super chip with 8 ConnectX-9 Super NICs, 8 CPXs, 2 Bluefield 4 DPUs (for offloading networking, storage, security, and crucially, a new context processor for KV caching to manage ever-growing AI memory), 2 Vera CPUs, and 4 Rubin packages (containing 8 Rubin GPUs). This node, with its MVLink switch and Spectrum X Ethernet (featuring silicon photonics and co-ax options), is engineered for unprecedented data throughput and low latency, capable of handling the entire global internet's peak traffic in one second across its spine.
Omniverse DSX: Digital Twins for AI Factories
NVIDIA Omniverse DSX (Data Center Simulation and Experience) is a blueprint and digital twin platform for designing, optimizing, and operating gigascale AI factories. It allows for the co-design of the physical infrastructure (building, power, cooling) with NVIDIA's AI compute stack. Partners like Jacobs Engineering, Siemens, Schneider Electric, and Vertiv use Omniverse to create SIM-ready open USD assets, simulate thermals and electricals with CUDA-accelerated tools (EAB, Cadence), and optimize compute density and layout. Once designed, partners deliver pre-fabricated modules for faster deployment.
During operation, the digital twin acts as an operating system, with AI agents (from FIDRA, Emerald AI) optimizing power consumption and detecting anomalies. This digital twin approach is critical for orchestrating the complex interplay of 1.2 million components, 2 tons of hardware, and 130 trillion transistors in a single GB200 MVLink72 rack.
Physical AI and Robotics
Physical AI, which understands the physical world, laws of physics, causality, and permanence, requires a three-computer architecture:
- GB200: For training AI models.
- Omniverse computer: For simulating the AI within a digital twin (e.g., of a factory or robot), leveraging generative AI, computer graphics, sensor simulation, and ray tracing.
- Jetson Thor robotics computer: For operating the robot in the real world.
This stack enables the development of robotic factories (e.g., Foxconn's facility in Houston for manufacturing NVIDIA AI infrastructure), where Fanuc manipulators, FII manual manipulators, and AMRs (Autonomous Mobile Robots) are orchestrated and monitored by Vision AI agents (built on NVIDIA Metropolis and Cosmos) within the Omniverse digital twin.
NVIDIA is also collaborating with companies like Figure (humanoid robots), Agility (warehouse automation), Johnson & Johnson (surgical robots), and Disney Research (adorable Disney Blue robot, using the Newton simulator for physically-aware learning).
NVIDIA Drive Hyperion: Autonomous Vehicles
NVIDIA Drive Hyperion is a standardized, end-to-end computing platform for robo-taxis and autonomous vehicles. It features a comprehensive sensor suite (surround cameras, radars, LiDAR) providing high-level cocoon perception and redundancy for safety. This platform allows AV system developers (Wayve, Waymo, Aurora, Momenta, Nuro, WeRide) to deploy their AI on a common, robust chassis. NVIDIA's partnership with Uber aims to connect these Hyperion-powered cars into a global robo-taxi network.
Experimental Setup & Results
▶ Watch: GPU and CUDA: Foundations of accelerated computing (7:00)
The keynote highlighted several significant performance benchmarks and market indicators demonstrating the impact of NVIDIA's innovations:
- Grace Blackwell (GB200) Performance: Jensen Huang cited benchmarks by SemiAnalysis, indicating that the Grace Blackwell architecture delivers 10 times the performance per GPU compared to the H200 (NVIDIA's previous generation, described as the "second best GPU in the world") across AI workloads. This substantial leap is attributed to the extreme co-design approach rather than just an increase in transistor count.
- Lowest Cost Token Generation: A key metric for AI factories is the cost-effectiveness of generating tokens. The Grace Blackwell MVLink72 system was shown to produce the lowest cost tokens in the world. Despite being NVIDIA's "most expensive computer," its unparalleled token generation capability (tokens per second) relative to its total cost of ownership (TCO) makes it the most economical solution for large-scale AI deployment.
- CSP Capex Investment: The top six Cloud Service Providers (CSPs) – Amazon, Corewave, Google, Meta, Microsoft, and Oracle – are making substantial capital expenditures (capex) in AI infrastructure. Huang revealed that these CSPs are investing heavily in NVIDIA's new architecture, specifically the Grace Blackwell MVLink72, indicating a broad industry shift towards accelerated computing and AI factories for optimal TCO.
- Blackwell Production and Growth: NVIDIA projects extraordinary growth for Blackwell. They have visibility into half a trillion dollars of cumulative Blackwell and early Reuben ramps through 2026 (excluding China). In the first several quarters of production (approximately 3.5 quarters), 6 million Blackwell units (each containing two GPUs in one large package) have been shipped. The projected total of 20 million Blackwell GPUs through 2026 represents a five-times growth rate compared to the entire lifetime shipments of the previous-generation Hopper GPUs.
- Vera Rubin Performance: The next-generation Vera Rubin rack-scale computer is designed to deliver 100 petaflops of performance. Huang noted that this single rack could replace 25 racks of the DGX-1 (the system delivered to OpenAI nine years prior), signifying a 100x performance increase in a significantly smaller footprint.
- Omniverse DSX Impact: Optimizations achieved through Omniverse DSX for a 1-gigawatt AI factory can deliver billions of dollars in additional revenue per year, by maximizing compute density, optimizing layouts, and reducing build times. This demonstrates the significant economic value of digital twin technology in AI factory design and operation.
These results collectively underscore the profound impact of NVIDIA's full-stack approach, delivering not just raw performance but also significant cost efficiencies and scalability critical for the widespread adoption and advancement of AI.
Practical Implications
▶ Watch: Key NVIDIA libraries: CU litho, cuDNN, Megatron Core, Monai (9:00)
The NVIDIA GTC keynote outlines a future with profound implications for practitioners, infrastructure teams, model builders, and deployers across various industries.
For Practitioners and Model Builders:
- Access to Powerful AI Agents: The concept of AI as "workers" rather than just "tools" means practitioners will increasingly interact with agentic AI systems like Cursor (used by NVIDIA engineers for code generation) or Perplexity (for web-based tasks). These agents leverage tools (e.g., VS Code, web browsers) to perform work, significantly boosting productivity.
- Rich Open-Source Ecosystem: NVIDIA's commitment to leading in open-source AI model contribution is a boon for startups, researchers, and developers. It ensures accessibility and customization for domain-specific applications, allowing companies to embed AI into their unique data and workflows. This reduces reliance on proprietary models for certain use cases, fostering innovation.
- Multimodal and Reasoning Capabilities: Model builders must focus on developing models with advanced reasoning, multimodality, and efficiency through techniques like distillation. NVIDIA's support for these capabilities across its stack enables the creation of more intelligent and adaptable AI.
For Infrastructure Teams:
- Shift to AI Factories: Infrastructure teams must prepare for a fundamental shift from general-purpose data centers to AI factories. These are specialized facilities designed for high-rate, cost-effective token generation, requiring unique considerations for power, cooling (100% liquid-cooled systems like Vera Rubin), and interconnectivity.
- Extreme Co-Design and Rack-Scale Systems: The move towards extreme co-design means infrastructure teams will be deploying highly integrated, rack-scale AI supercomputers like Grace Blackwell MVLink72 and Vera Rubin. These systems require expertise in specialized networking (Spectrum X Ethernet, Infiniband), dense compute packaging, and potentially novel power delivery and thermal management solutions.
- Digital Twin for Operations: NVIDIA Omniverse DSX offers a blueprint for designing, optimizing, and operating gigascale AI factories using digital twins. This means infra teams can simulate layouts, power constraints, thermals, and electricals virtually before physical construction, reducing build times and optimizing ongoing operations through AI agents. This necessitates skills in digital twin technologies and simulation.
- Energy Considerations: The exponential growth of AI compute demands significant energy. Infrastructure teams must consider sustainable energy sources and highly efficient designs, as highlighted by President Trump's "pro-energy initiative."
For Model Deployers and Enterprise Integrators:
- Agentic SaaS: Enterprise software will evolve into "agentic SaaS," where AI agents are integrated directly into workflows. NVIDIA is partnering with major SaaS providers (e.g., ServiceNow for enterprise workflows, SAP for commerce, Synopsys/Cadence for EDA, CrowdStrike for cybersecurity, Palantir for data processing and insight) to embed its CUDA X libraries and AI systems, enabling AI-powered automation and intelligence at the core of enterprise operations.
- Physical AI and Robotics: Deployers will see a surge in physical AI applications. From robotic factories (e.g., Foxconn's NVIDIA AI infrastructure manufacturing plant) orchestrating various robots (Fanuc manipulators, AMRs) to advanced humanoid robots (Figure), warehouse automation (Agility), surgical robots (Johnson & Johnson), and autonomous vehicles (Drive Hyperion in robo-taxis like Uber's fleet). This requires robust simulation environments (Omniverse) and dedicated edge AI compute (Jetson Thor).
- New Telecommunications Infrastructure: The NVIDIA Arc platform will transform 6G telecommunications networks into intelligent, software-defined systems. This means new opportunities for deploying AI at the edge for spectral efficiency (AI for RAN) and new cloud computing services directly on wireless infrastructure (AI on RAN).
Tradeoffs and Limitations:
- High Initial Investment: While Grace Blackwell offers the lowest cost per token over its lifetime, the initial capital expenditure for these advanced, co-designed systems is substantial. This might create a barrier for smaller organizations without access to cloud-based NVIDIA infrastructure.
- Complexity of Integration: The extreme co-design approach, while powerful, also implies a high degree of complexity. Integrating these full-stack solutions, especially in novel domains like quantum computing or physical AI, requires specialized expertise and close collaboration with NVIDIA and its ecosystem partners.
- Energy Consumption: Despite efficiency gains, the sheer scale of AI factories and the exponential demand for compute will continue to pose significant challenges for energy consumption and sustainability.
Overall, NVIDIA's vision signals a future where AI permeates every layer of technology and industry, demanding a complete re-evaluation of computing infrastructure, software development, and operational strategies.
Key Takeaways
- AI and Accelerated Computing Drive a New Industrial Revolution: NVIDIA posits that AI is fundamentally transforming computing, moving beyond traditional software tools to intelligent AI "workers" capable of performing complex tasks, thereby addressing a vast segment of the global economy. This shift is underpinned by accelerated computing, which has overcome the limits of Moore's Law.
- Extreme Co-Design is Essential for Exponential Performance: To meet the "two exponentials" of growing AI compute demand and usage, NVIDIA champions "extreme co-design," re-architecting chips, systems, software, models, and applications from the ground up to deliver compounding performance gains far beyond traditional improvements.
- NVIDIA is Building "AI Factories" and Digital Twins: The future of AI infrastructure involves specialized "AI factories" designed to efficiently produce valuable AI "tokens." The Omniverse DSX platform enables the design, optimization, and operation of these gigascale factories using digital twins, integrating physical infrastructure with AI compute.
- Groundbreaking Platforms Across Strategic Sectors: NVIDIA is introducing foundational platforms for critical industries: NVIDIA Arc for 6G telecommunications, NVQLink for hybrid quantum-classical computing, Grace Blackwell/Vera Rubin for next-generation AI supercomputing, Drive Hyperion for autonomous vehicles, and comprehensive Physical AI solutions for robotics and manufacturing.
- Commitment to Open-Source AI and Broad Ecosystem Integration: NVIDIA is a leading contributor to open-source AI models, recognizing their importance for innovation. The company's CUDA X libraries and AI models are deeply integrated into major cloud providers and enterprise SaaS platforms, democratizing access to advanced AI capabilities.
- Re-industrialization of America through AI Manufacturing: NVIDIA is actively bringing high-tech manufacturing back to the United States, with Blackwell production in Arizona and future AI factories being built domestically, emphasizing national security and economic growth through AI-powered re-industrialization.
About the Speaker(s)
Jensen Huang is the visionary founder and CEO of NVIDIA. Since co-founding the company in 1993, he has been instrumental in pioneering the field of accelerated computing, leading the invention of the Graphics Processing Unit (GPU) and the CUDA programming model. Under his leadership, NVIDIA has evolved from a graphics chip company into a global leader in AI computing, driving breakthroughs across diverse domains including gaming, professional visualization, data centers, autonomous vehicles, and robotics. Huang's keynotes are renowned for their expansive vision, articulating how NVIDIA's technological innovations are shaping the future of computing and addressing some of the world's most complex challenges. His deep understanding of hardware-software co-design and his long-term strategic outlook have positioned NVIDIA at the forefront of the AI industrial revolution.
Reviews
Simon Wisk (Open Source Developer & AI Tooling Expert) — PASS
This is a CEO keynote for a GPU company, which means it's a press release with stage lighting. There is no engineering content here that an engineer couldn't have gotten from NVIDIA's product announcements, and the article summarizing it reads like it was written by the same marketing team that wrote the slide deck. No failure modes, no reproducibility, no implementation details — just a parade of branded product names and benchmark numbers cited without methodology.
Jensen Hitch (AI Compute Platform CEO) — STRONG ACCEPT
This is a platform-level keynote that reasons from physical constraints — the end of Moore's Law and Dennard scaling — up through chip architecture, system design, interconnect, software, and deployment economics. The core argument is coherent and defensible: when you can't rely on transistor scaling, you have to co-design the entire stack to compound gains. The GB200 MVLink72 10x inference claim is grounded in a real architectural insight — treating the rack as a single virtual GPU fabric changes what's possible at inference time, not just training. The AI factory framing is the right unit of analysis. The gaps are real but not fatal: energy consumption gets acknowledged but not seriously…