A New Era for Generalist Robotics: The Rise of Humanoids | NVIDIA GTC 2025

Bernt Børnich (CEO and Founder · 1X), Aaron Saunders (CTO · Boston Dynamics), Deepak Pathak (CEO and Co-Founder · Skild AI)

NVIDIA GTC 2025 · Session

Overview

This panel discussion, "A New Era for Generalist Robotics: The Rise of Humanoids," brought together a distinguished group of leaders from the forefront of robotics and AI to discuss the monumental shifts occurring in the field. Moderated by Tiffany Jansen, the panel featured executives from 1X, Skilled AI, Agility Robotics, Boston Dynamics, and NVIDIA, each offering unique perspectives on the current state and future trajectory of humanoid robots. The core theme revolved around the transition of robotics from specialized, control-driven machines to generalist, intelligent entities powered by foundation models and learning-by-experience paradigms.

Watch on YouTube

Visual summary for A New Era for Generalist Robotics: The Rise of Humanoids | NVIDIA GTC 2025 by Bernt Børnich, Aaron Saunders, Deepak Pathak
Visual summary for A New Era for Generalist Robotics: The Rise of Humanoids | NVIDIA GTC 2025 by Bernt Børnich, Aaron Saunders, Deepak Pathak

Key moments

  1. 0:00 Panel and speaker introductions
  2. 4:00 Why is robotics accelerating now?
  3. 4:30 Moravec's Paradox and foundation models' impact
  4. 6:00 Generating data at scale with simulation
  5. 6:40 Robotics hardware becoming better and more affordable
  6. 7:20 Importance of closing the sim-to-real gap

A New Era for Generalist Robotics: The Rise of Humanoids

Speakers: Tiffany Jansen (Moderator, Founder, Tiffen Tech), Burnt Barnick (Founder & CEO, 1X), Deepak Puk (CEO & Co-founder, Skilled AI), Pros Feligaputo (CTO, Agility Robotics), Aaron Saunders (CTO, Boston Dynamics), Jim Fan (Co-lead, NVIDIA Gear Lab, Project GR00T)

Conference: NVIDIA GTC

YouTube: https://www.youtube.com/watch?v=BmD22FNOAY4

Overview

This panel discussion, "A New Era for Generalist Robotics: The Rise of Humanoids," brought together a distinguished group of leaders from the forefront of robotics and AI to discuss the monumental shifts occurring in the field. Moderated by Tiffany Jansen, the panel featured executives from 1X, Skilled AI, Agility Robotics, Boston Dynamics, and NVIDIA, each offering unique perspectives on the current state and future trajectory of humanoid robots. The core theme revolved around the transition of robotics from specialized, control-driven machines to generalist, intelligent entities powered by foundation models and learning-by-experience paradigms.

The talk delved into the confluence of advancements in AI models, data generation, and hardware that are collectively propelling robotics into an unprecedented era of growth and capability. Speakers highlighted the emergence of universal robot brains, the critical role of simulation and diverse data, and the practical implications for widespread deployment across various industries. This discussion is particularly relevant for anyone interested in the practical application of advanced AI, the future of labor, and the integration of autonomous systems into daily life, marking a significant inflection point in the long-envisioned promise of intelligent physical agents.

The consensus among the panelists was that robotics is no longer a slow-moving, niche application of AI but is now at a "tipping point," poised for exponential growth. NVIDIA's own Project GR00T (Generalist Robot 00T) was introduced as a moonshot initiative aimed at building the foundational "robot brain" for humanoids, underscored by the open-sourcing of the GR00T N1 model, signaling a commitment to democratizing physical AI and accelerating industry-wide progress.

Background

▶ Watch: Panel and speaker introductions (0:00)

Robotics, often cited as the oldest application of Artificial Intelligence, has historically progressed at a slower pace compared to its digital counterparts like natural language processing and computer vision. This paradox, known as Moravec's Paradox, posits that tasks easy for humans (like physical movement and perception) are incredibly hard for machines, while tasks difficult for humans (like complex calculations or creative writing) are relatively easier for AI. For decades, robotics was predominantly a field driven by classical controls theory, where engineers meticulously programmed every joint movement and interaction, a method that originated from World War II applications like flying planes and missiles. This approach, while effective for highly structured environments and specialized tasks, lacked the adaptability and generalizability necessary for robots to navigate and learn in the complex, unstructured real world.

Several critical barriers prevented the widespread adoption and advancement of general-purpose robots. Hardware was prohibitively expensive and lacked the robustness to withstand real-world interactions without frequent breakdowns, making experimentation slow and costly. Data scarcity was another major hurdle; unlike the internet's vast repositories of text and images for LLMs, there was no equivalent "fossil fuel" of motor control data or robot trajectories available for training. Furthermore, the sim-to-real gap—the challenge of transferring policies learned in simulation to physical robots—remained significant, as simulations struggled to accurately represent real-world physics and run at sufficient speeds.

However, the landscape has dramatically shifted. The "ChatGPT moment" ushered in an era of large foundation models and multimodal models capable of sophisticated reasoning and open-vocabulary understanding of the 3D visual world, providing necessary (though insufficient) conditions for robust robot perception. Concurrently, GPU-accelerated simulation environments, like NVIDIA's Isaac Sim, have matured, enabling the generation of vast amounts of synthetic data at speeds far exceeding real-world collection (e.g., 10 years of training data in 3 hours). On the hardware front, advancements in consumer electronics have commoditized components like batteries, cameras, and compute, making robots more affordable and capable. This convergence of breakthroughs in AI models, data generation techniques, and hardware robustness is finally allowing robotics to overcome Moravec's Paradox and transition from bespoke, control-based systems to versatile, learning-driven intelligent agents.

Key Findings

▶ Watch: Moravec's Paradox and foundation models' impact (4:30)

The panel discussion illuminated several key findings and strategic shifts driving the new era of generalist robotics:

  • The Tipping Point for Generalist Robotics: The panelists unanimously agreed that robotics is at a significant inflection point, moving beyond single-purpose machines to multi-purpose and eventually general-purpose robots. This shift is fueled by the convergence of advanced AI, simulation, and hardware, making the long-held dream of adaptable humanoids finally tractable.
  • NVIDIA's Project GR00T and GR00T N1 Model: NVIDIA introduced Project GR00T as its moonshot initiative to build a generalist "robot brain" for humanoids. A cornerstone of this effort is the GR00T N1 model, announced as the world's first open humanoid robot foundation model. Despite being relatively compact at 2 billion parameters, it is touted for "punching above its weight" and representing state-of-the-art autonomous humanoid intelligence, with its open-sourcing aiming to democratize physical AI.
  • Learning by Experience as the New Paradigm: A fundamental shift from traditional control-based robotics is the emphasis on learning by experience. Deepak Puk highlighted that early AI pioneers like Turing envisioned robots that learn like children, not pre-programmed adults. This approach, now enabled by modern AI, allows robots to acquire skills through interaction and feedback, much like humans.
  • The Data Pyramid Strategy: To address the unique data scarcity problem in robotics, GR00T employs a sophisticated "data pyramid" strategy. This involves combining high-quality but limited real robot data (from teleoperation), massively scalable simulation data (generated via physics engines like Isaac Sim), and innovative neuro-simulation data derived from large multimodal models (VLMs and video generation models) that "dream" new trajectories and actions.
  • Cross-Embodiment as a Solvable Challenge: The ability of a single AI model to control diverse robot hardware (cross-embodiment) was identified as a critical problem. While challenging, the human brain's capacity for cross-embodiment (e.g., adapting to different video game controls) suggests it's solvable. Early research, such as procedurally generating diverse robot bodies in simulation and tokenizing their embodiment, shows promise for generalization across different robot designs, even within the same product line (V1 vs. V2).
  • Hardware Robustness and Co-evolution: The panel stressed the crucial role of robust and reliable hardware that can safely interact with the real world without breaking. Burnt Barnick emphasized that safety must be intrinsic to the machine. While a "general brain" is key, the hardware and software often need to co-evolve to achieve optimal performance, especially for tasks involving heavy objects or high precision. Agility Robotics demonstrated this with their Digit humanoid, where a fully learned recovery behavior, hardened through domain randomization, transferred seamlessly to a new, heavier robot.
  • Robots as Adaptive Learning Engines: Unlike traditional AI models that are trained and then deployed as static entities, robotics AI will involve deploying "mini learning engines" that continuously adapt on the fly. Deepak Puk compared this to the human brain adapting to a sore arm after a workout, highlighting the need for robots to learn and adjust to changes in their own bodies and environments in real-time.
  • Interaction as the Enemy of Hallucination: Robotics offers a unique advantage in addressing the hallucination problem prevalent in LLMs. By interacting with the physical world, robots receive direct, verifiable feedback on their actions. Burnt Barnick illustrated this with a robot that learned to correctly identify and close a toilet seat after initial GPT-4o predictions were 50% accurate. This closing of the loop through physical interaction inherently reduces hallucination.

Technical Deep Dive

▶ Watch: Generating data at scale with simulation (6:00)

The technical core of the discussion revolved around the architecture of generalist robot brains, advanced data generation strategies, and the challenges of hardware-software integration.

NVIDIA's Project GR00T is centered on developing a universal foundation model for humanoid robots. Jim Fan detailed the architecture of the GR00T N1 model, emphasizing its design principles:

  1. Simplicity and End-to-End Learning: The model is designed to be as simple and end-to-end as possible, directly mapping "photons to actions." This means taking raw pixel input from cameras and directly outputting continuous floating-point numbers representing motor control values. This approach draws inspiration from the success of large language models (LLMs) like ChatGPT, which unified diverse NLP tasks into a single, simple text-to-text transformer architecture. For GR00T N1, this translates to a transformer that maps a sequence of integers (representing visual input) to another sequence of integers (representing motor commands). This simplification allows for the unification of various data types and tasks into a single model.
  2. Complex Data Pipeline: While the model itself is simple, the data pipeline surrounding it is highly sophisticated, addressing the inherent data scarcity in robotics. GR00T's data strategy is conceptualized as a pyramid:
  • Top Layer: Real Robot Data: This constitutes the highest quality data, collected through teleoperation in the real world. However, it is inherently limited by physical constraints (e.g., 24 hours per robot per day) and the challenges of scaling in the "world of atoms."
  • Middle Layer: Simulation Data: This layer leverages powerful physics engines like NVIDIA's Isaac Sim to generate vast quantities of synthetic data. GPU-accelerated simulation can produce "10 years worth of training data in maybe three hours worth of compute time." This data can be generated based on real-world collected data or through "learning from experience" within the simulated environment. NVIDIA's heritage as a graphics company provides a strong foundation for high-fidelity physics rendering.
  • Bottom Layer: Multimodal Internet Data and Neuro-Simulations: This innovative layer utilizes the massive amounts of multimodal data (text, images, audio, video) available on the internet. This data is used to train Visual Language Models (VLMs), which form the foundation for vision-language-action models. Critically, advanced video generation models are employed as "neuro simulations" of the world. These models, trained on hundreds of millions of online videos, learn physics implicitly. They can be prompted to "hallucinate" or "dream" new robot trajectories that are physically accurate in pixels. GR00T N1 proposes an algorithm called latent action to extract these actions from the generated video "dreams" and feed them back into the data pyramid. This effectively allows robots to learn from imagined scenarios grounded in real-world physics observed online.

The discussion also touched upon cross-embodiment, the ability for a single brain to control different physical robot bodies. Jim Fan mentioned previous exploratory research called "metamorph," where thousands of procedurally generated, simple robots with varying joint connectivity (e.g., snake-like, spider-like) were created in simulation. A robot grammar was used to tokenize the robot's body, converting its embodiment into a sequence of integers. By applying transformers to this diverse set of embodiments, the system demonstrated an ability to generalize to novel, unseen robot designs. This approach suggests a path toward a universal description language for robots, allowing a single model to adapt to a wide range of hardware.

Aaron Saunders and Pros Feligaputo emphasized the importance of robust hardware design and traditional robotics tools. They argued that while end-to-end AI is powerful, a "big toolbox of robotics tools going back 70 years" remains relevant for achieving determinism, functional safety, and maintaining customer trust. This includes good calibration methods, thorough characterization of robot dynamics, and strong joint-level control. Agility Robotics' experience with their Digit humanoid highlights this synergy: a fully learned recovery behavior, robustified through extensive domain randomization in simulation, allowed for a "one-shot transfer" to a significantly heavier new robot with different kinematics. This demonstrates that careful engineering and data augmentation can bridge hardware differences.

Deepak Puk introduced the concept of robots as learning engines rather than static "train-deploy" systems. He argued that unlike other AI applications where the environment is abstracted away (e.g., CUDA for GPUs), robots must constantly adapt to their physical bodies and surroundings. He used the analogy of a human brain adapting to a sore arm after a workout, requiring more torque for the same action. This real-time adaptation, sometimes referred to as Rapid Motor Application (RMA), involves adding the robot's specific history and dynamics to the model's context, allowing it to learn its own unique characteristics.

Experimental Setup & Results

▶ Watch: Robotics hardware becoming better and more affordable (6:40)

While this was a panel discussion rather than a scientific paper presentation, several concrete examples and claims regarding experimental results were shared, highlighting the practical advancements in the field.

NVIDIA's GR00T N1 model, with 2 billion parameters, was presented as the "world's first open humanoid robot foundation model" and claimed to "punch above its weight" in terms of autonomous humanoid intelligence. This implies that despite its relatively modest parameter count compared to some LLMs, its specialized architecture and data strategy yield significant performance.

Agility Robotics provided a compelling example of successful sim-to-real transfer and robust learning with their Digit humanoid. Pros Feligaputo detailed how Digit's recovery behavior (e.g., standing up after a fall) is fully learned. This policy was extensively trained using domain randomization—varying physical parameters and environmental conditions within the simulation—to make it highly robust. A notable outcome was the one-shot transfer of this learned policy to a new version of the Digit robot, which was 10 kg heavier and had a much larger frame with slightly different kinematics and payload capacity. This seamless transfer, without retraining, demonstrated the effectiveness of their robustification techniques and understanding of foot contact and other critical physical details in simulation.

Burnt Barnick from 1X shared a practical anecdote about their Eve robot (a wheeled mobile robot) and its interaction with a common household object. Initially, when tasked with identifying whether a toilet seat was up or down, a GPT-4o model performed at a 50% accuracy rate, essentially guessing. However, when the robot was deployed with an autonomous policy to physically interact with the toilet (checking and closing the seat if it was up), it created a closed-loop feedback mechanism. The robot could then verify its actions ("I closed it, so I know it's down") and correct the model's initial "hallucination." This served as a tangible demonstration of how real-world interaction data can rapidly improve a robot's understanding and performance in specific tasks.

These examples, while anecdotal in some cases, illustrate the nascent but growing capabilities of learning-based robotics, the power of robust simulation, and the potential for real-world interaction to drive rapid improvement in robot intelligence. The emphasis is less on traditional benchmarking and more on demonstrating practical utility and adaptability in unstructured environments.

Practical Implications

▶ Watch: Importance of closing the sim-to-real gap (7:20)

The rise of generalist robotics, powered by foundation models and learning-by-experience, carries profound practical implications for various stakeholders:

  • For Practitioners and Infra Teams:
  • Addressing Labor Shortages: Humanoid robots are envisioned as a solution to global labor shortages across diverse sectors including hospitals, elderly care, retail, factories, and logistics. 1X's mission to create "an abundance of labor" through safe, intelligent humanoids exemplifies this.
  • Democratization of Physical AI: Initiatives like NVIDIA's open-sourcing of the GR00T N1 model aim to lower the barrier to entry for developing and deploying advanced robotic systems. This could foster innovation and accelerate the growth of the robotics ecosystem, much like open-source LLMs have done for natural language processing.
  • Balancing Novelty with Reliability: While end-to-end learning offers unprecedented capabilities, practitioners must balance these with the need for determinism, functional safety, and non-regression in deployed systems. Aaron Saunders emphasized that the "big toolbox of robotics tools going back 70 years" (e.g., robust calibration, joint-level control) remains crucial for maintaining customer trust, especially when robots operate near humans or handle hazardous materials.
  • Complex Data Infrastructure: Infra teams will need to build and manage sophisticated data pipelines that integrate real-world teleoperation data, GPU-accelerated simulation data, and neuro-simulation data derived from multimodal internet sources. The sheer scale and diversity of this data will be a significant challenge.
  • For Model Builders:
  • The Primacy of Data Diversity: Burnt Barnick highlighted that "diversity [of data] is more important than scale" in the early stages of robotics AI. Model builders must prioritize collecting data from as many different tasks, environments, and dynamic scenarios as possible to foster true intelligence and generalization, rather than just raw volume. The "washing machine" example illustrated this need for understanding underlying task principles.
  • Adaptive Learning Engines: Future robot models will need to be "learning engines" that can adapt on the fly to changes in their own bodies (e.g., wear and tear, post-workout fatigue) and unexpected environmental conditions. This requires novel architectures that incorporate continuous adaptation and rapid motor application (RMA) capabilities.
  • Harnessing Interaction for Learning: Model builders can leverage the unique ability of robots to interact with the physical world to generate verifiable feedback, effectively acting as the "enemy of hallucination." Designing systems that close the loop between action, observation, and model correction will be paramount.
  • For Deployers:
  • Hardware-Software Co-evolution: The debate between a truly universal "brain" and robot-specific hardware platforms is ongoing. While a general brain is desirable, Pros Feligaputo and Aaron Saunders argued that hardware design (e.g., kinematics, sensing, inertial properties, gripper design) significantly impacts performance, especially for tasks requiring dexterity, heavy lifting, or operation in hazardous environments. Deployers may need to consider platforms where hardware and software have been co-developed or optimized for specific task domains.
  • Gradual Adoption and Specialization: Deepak Puk noted that usefulness in robotics doesn't require full generalization. Specialized robots performing one or two tasks can be immensely valuable and will likely see much quicker adoption. This pragmatic approach allows for incremental deployment and value generation, contrasting with the "all or nothing" challenge seen in some other AI fields.
  • Societal and Ethical Considerations: As robots become more capable and integrated, issues of safety, societal acceptance, and ethical deployment will become increasingly critical. Deployers must navigate these challenges, ensuring that robots are not only functional but also trustworthy and beneficial to society.

In the long term, robotics has the potential to fundamentally transform human society by accelerating scientific discovery (e.g., automating lab experiments), enabling recursive self-improvement (robots building and fixing other robots), and fundamentally redefining labor. The journey, however, involves complex trade-offs between generalization and specialization, end-to-end learning and traditional controls, and the continuous co-evolution of hardware and software.

Key Takeaways

  • Robotics is at an Inflection Point: Driven by advancements in foundation models, GPU-accelerated simulation, and more affordable/robust hardware, robotics is rapidly transitioning from specialized, control-based systems to generalist, learning-driven agents.
  • GR00T N1 Leads the Open Generalist Charge: NVIDIA's Project GR00T and the open-sourced GR00T N1 model (2 billion parameters) represent a significant step towards a universal robot brain, aiming to democratize physical AI with an end-to-end "photons to actions" architecture.
  • Data Diversity and Simulation are Paramount: Overcoming data scarcity in robotics requires a multi-faceted approach, including real-world teleoperation, massively scalable physics-based simulation (e.g., Isaac Sim), and innovative "neuro-simulation" using video generation models to "dream" new trajectories and actions via latent action extraction. Diversity of data, not just scale, is crucial for true intelligence.
  • Robots as Adaptive Learning Engines: Unlike static "train-deploy" AI, robots need to be deployed as continuous learning engines that adapt on the fly to their own changing physical state and dynamic environments, akin to how humans adapt to their bodies (e.g., Rapid Motor Application).
  • Synergy of AI and Traditional Robotics: While new AI models are transformative, the panel emphasized the continued importance of traditional robotics expertise in hardware design, calibration, and low-level controls for ensuring determinism, safety, and reliability in real-world deployments.
  • Profound Societal Impact: In the next 10-20 years, generalist robotics is poised to create an abundance of labor, automate scientific discovery, enable robots to build and repair other robots (recursive self-improvement), and integrate deeply into daily life, fundamentally changing human-robot interaction and societal structures.

About the Speaker(s)

The panel was moderated by Tiffany Jansen, Founder of Tiffen Tech, who expressed her excitement about the recent advancements in humanoids and the opportunity to learn from industry leaders.

Burnt Barnick is the Founder and CEO of 1X, a company on a mission to create an abundance of labor through safe, intelligent humanoids. His philosophy centers on the belief that robots need to "live and learn among us," starting with consumer applications to experience the nuances of human life before expanding to other verticals like hospitals, elderly care, retail, and logistics.

Deepak Puk is the CEO and Co-founder of Skilled AI. His company is focused on building a "general brain for robotics," advocating for a single shared foundation model that can leverage data from any platform, task, or scenario to overcome data scarcity in the field.

Pros Feligaputo serves as the CTO at Agility Robotics. Agility is known for its Digit humanoid robot, which is designed for work in manufacturing and logistics. Pros highlighted their strategy of deploying robots in real customer environments to gather practical experience and drive technological advancements.

Aaron Saunders is the CTO of Boston Dynamics, a company with a long-standing mission to "make robots real." Having worked on humanoids "before they were cool," Aaron's team has shipped thousands of robots, with their latest humanoid aiming to perform useful work that removes people from "dirty, dull, dangerous things."

Jim Fan is the Co-lead of NVIDIA Gear Lab and Project GR00T. He described GR00T as NVIDIA's "moonshot initiative" to build the foundation model, the "robot brain," for humanoid robots, representing their strategy for the next generation computing platform for physical AI. He also announced the open-sourcing of the GR00T N1 model, highlighting NVIDIA's commitment to democratizing physical AI.

Reviews

Simon Wisk (Open Source Developer & AI Tooling Expert) — WEAK

A high-profile panel with genuinely interesting participants, but the write-up reads like a cleaned-up press briefing rather than a technical account of something that was built and demonstrated. The GR00T N1 architecture gets the closest to real substance — the data pyramid framing and latent action extraction from video generation models are worth knowing about — but there's not enough implementation specificity to do anything with it. The anecdotes (toilet seat at 50% accuracy, Digit transferring to a heavier body) are illustrative but thin. No code, no benchmarks with methodology, no reproducible setup. Engineers leave knowing that the vibe is good in humanoid robotics; they don't…

Jensen Hitch (AI Compute Platform CEO) — SOLID

A well-moderated panel that captures the genuine inflection point in physical AI — foundation models meeting real robot hardware — with credible voices from the companies actually shipping systems. The NVIDIA GR00T framing is the most technically substantive thread, and the data pyramid strategy is the right way to think about the scarcity problem. But as a panel discussion, it stays largely at the vision and framing layer. The hard system questions — inference latency on-device, power envelope on a mobile humanoid, what breaks at the 10,000-unit deployment scale — are touched but not answered. Strong orientation material. Not a platform-level engineering reference.

→ Top-rated talks at NVIDIA GTC 2025

All talks from NVIDIA GTC 2025