Kubespray: Driving Cost-Efficiency for AI on Kubernetes - Antoine Legrand & Mohamed Zaian

Antoine Legrand, Mohamed Zaian

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In an era where Artificial Intelligence (AI) workloads are increasingly critical for enterprise innovation, the underlying infrastructure costs and vendor lock-in present significant challenges. This KubeCon EU talk, presented by Antoine Legrand, an initial Kubespray maintainer and SITU at Conny GmbH, alongside co-maintainer Mohamed Zaian (who unfortunately could not attend), delves into how Kubespray can serve as a pivotal tool for achieving cost-efficiency and unparalleled flexibility in deploying and managing Kubernetes clusters, especially for demanding AI applications. The core message revolves around leveraging Kubespray to abstract away infrastructure specifics, thereby empowering organizations to shop around for the most competitive hardware and cloud resources without being constrained by proprietary ecosystems.

Watch on YouTube

Visual summary for Kubespray: Driving Cost-Efficiency for AI on Kubernetes - Antoine Legrand & Mohamed Zaian by Antoine Legrand, Mohamed Zaian
Visual summary for Kubespray: Driving Cost-Efficiency for AI on Kubernetes - Antoine Legrand & Mohamed Zaian by Antoine Legrand, Mohamed Zaian

Key moments

  1. 0:00 Introduction to Kubespray: Purpose, Flexibility, and Stability
  2. 3:15 The challenge: High infrastructure costs and vendor lock-in
  3. 4:00 Kubespray's solution for vendor independence and cost negotiation
  4. 6:10 User testimonial: Vestia Transun's AI infrastructure challenge
  5. 7:00 Why Kubespray was chosen for hardware agnostic GPU management
  6. 8:00 Real-world AI workloads and tools running on Kubespray

Kubespray: Driving Cost-Efficiency for AI on Kubernetes

Speakers: Antoine Legrand, SITU at Conny GmbH; Mohamed Zaian, Senior Engineer at New York I See

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=SqKqB-q_8E

Overview

In an era where Artificial Intelligence (AI) workloads are increasingly critical for enterprise innovation, the underlying infrastructure costs and vendor lock-in present significant challenges. This KubeCon EU talk, presented by Antoine Legrand, an initial Kubespray maintainer and SITU at Conny GmbH, alongside co-maintainer Mohamed Zaian (who unfortunately could not attend), delves into how Kubespray can serve as a pivotal tool for achieving cost-efficiency and unparalleled flexibility in deploying and managing Kubernetes clusters, especially for demanding AI applications. The core message revolves around leveraging Kubespray to abstract away infrastructure specifics, thereby empowering organizations to shop around for the most competitive hardware and cloud resources without being constrained by proprietary ecosystems.

Kubespray, a robust and mature orchestrator for Kubernetes, has been a cornerstone in the community for a decade, focusing on production environments, stability, and comprehensive lifecycle management. Legrand highlights its unique ability to provide a common denominator for Kubernetes deployments across a vast spectrum of environments—from public clouds like Amazon, Google Cloud, and Azure, to on-premise bare metal and private data centers. This talk specifically illuminates how this flexibility translates into tangible cost savings and strategic advantages for companies running GPU-intensive AI workloads.

The relevance of this discussion is particularly acute for organizations grappling with the escalating costs of AI infrastructure, which often involves expensive GPUs and specialized computing resources. By demonstrating how Kubespray enables hardware agnosticism and multi-cloud strategies, the speakers make a compelling case for re-evaluating traditional infrastructure procurement models. The presentation underscores that true cost-efficiency stems not merely from finding the cheapest provider, but from cultivating the freedom and control to choose, negotiate, and adapt, positioning Kubespray as an indispensable enabler in this strategic shift.

Background

▶ Watch: Introduction to Kubespray: Purpose, Flexibility, and Stability (0:00)

Kubespray has a rich history, predating even the Cloud Native Computing Foundation (CNCF) itself, having been active for ten years. It originated as an orchestrator designed for the complete lifecycle management of Kubernetes clusters, encompassing installation, upgrades, and ongoing maintenance. Its design philosophy centers on production readiness, emphasizing stability, safe upgrades, and robust operation. A defining characteristic of Kubespray is its exceptional flexibility, allowing it to deploy Kubernetes across virtually any environment. This includes major cloud providers, bare metal servers, and private data centers, supporting a wide array of operating systems, container runtimes like containerd, and networking interfaces such as Calico and Cilium. This comprehensive adaptability means users can tailor their Kubernetes stack to precise requirements, rather than conforming to rigid vendor specifications.

The motivation behind Kubespray's focus on cost-efficiency stems directly from the maintainers' own experiences. Legrand revealed that Kubespray's extensive CI infrastructure deploys an astounding 10,000 to 15,000 Kubernetes clusters monthly for testing purposes. For every pull request (PR), 20 to 50 clusters are deployed, often spinning up 200 to 300 virtual machines (VMs) concurrently to test various configurations (e.g., Ubuntu with Calico). This massive testing footprint, while crucial for stability, is inherently expensive. The team realized that optimizing this cost involved two primary approaches: writing more efficient code and, more significantly, optimizing the underlying infrastructure.

This infrastructure focus led to a critical observation: the market for computing resources, especially those equipped with GPUs for AI workloads, exhibits a vast price disparity—potentially a 10x difference between the cheapest and most expensive options. While the immediate impulse might be to simply opt for the lowest-cost provider, Legrand stressed that the fundamental problem to avoid is vendor lock-in. When an organization is deeply integrated with a single provider's proprietary services, switching becomes prohibitively difficult and costly, eroding any initial savings. This lock-in diminishes negotiation power and limits strategic flexibility.

Kubespray directly addresses this dilemma by serving as a common denominator for all Kubernetes clusters. By standardizing the deployment and management of Kubernetes across diverse infrastructures, Kubespray enables organizations to build their entire AI stack on a portable foundation. This abstraction layer means that the underlying hardware—whether it's cloud VMs, on-prem servers, or specialized GPU appliances—becomes a fungible resource, effectively reducing them to "just some machines" or "a bunch of IPs." This strategic positioning allows organizations to reclaim control over where and how they run their workloads, fostering true freedom and control, which Legrand identifies as having the "biggest impact" on cost savings in the long run.

Key Findings

▶ Watch: Kubespray's solution for vendor independence and cost negotiation (4:00)

The central revelation of this talk is that vendor independence is not merely a desirable state but a critical enabler for achieving substantial cost efficiencies in AI infrastructure. Antoine Legrand emphatically stated that removing vendor lock-in has the "biggest impact" on overall cost, far outweighing the transient benefits of simply choosing the cheapest provider at any given moment. Kubespray facilitates this independence by providing a unified, hardware-agnostic platform for Kubernetes deployments, allowing organizations to abstract their AI workloads from the underlying infrastructure.

This newfound freedom translates directly into enhanced negotiating power with major cloud providers. Legrand cited examples of companies aggressively negotiating discounts with industry giants like Amazon, Google Cloud, and Azure, leveraging their ability to shift workloads between providers. The strategic advantage of being able to "shop around" for infrastructure means that organizations are no longer captive customers but rather empowered consumers who can demand better terms and pricing.

Furthermore, Kubespray enables unprecedented flexibility and the ability to mix-and-match infrastructure components. Organizations can seamlessly integrate diverse resources, such as high-performance on-premise GPUs (like NVIDIA DGX or A100s) for core AI training, with cloud-based resources for other workloads or burst capacity. Kubespray acts as a single, consistent interface for managing all these disparate clusters, simplifying operations and maximizing resource utilization. This hybrid approach allows businesses to optimize for both performance and cost, placing specific workloads on the most suitable and cost-effective hardware.

Another significant finding is Kubespray's inherent hardware agnosticism. As demonstrated by the user story from Vestia Transun, Kubespray can deploy and manage Kubernetes on virtually any hardware, from enterprise-grade DGX systems to smaller, individual GPUs. It integrates seamlessly with specialized operators, such as the NVIDIA GPU operator, to automate the deployment and management of GPU resources within the cluster. This capability ensures that organizations are not limited to specific hardware vendors or cloud-managed GPU services, opening up a wider array of procurement options.

Finally, the talk highlighted Kubespray's commitment to supporting latest Kubernetes versions, which is particularly vital for AI deployments. Enterprise environments often struggle with timely Kubernetes upgrades due to complexity and stability concerns. Kubespray, with its rigorous testing (supporting three Kubernetes versions per release and aiming for updates within two to three months of upstream releases), provides a reliable path to leveraging the newest features, performance enhancements, and security patches critical for cutting-edge AI frameworks and applications. This focus on rapid adoption ensures that AI teams can always access the most optimized and secure environment for their work.

Technical Deep Dive

▶ Watch: User testimonial: Vestia Transun's AI infrastructure challenge (6:10)

Kubespray's technical prowess lies in its role as a comprehensive Kubernetes orchestrator. At its core, it automates the entire lifecycle of a Kubernetes cluster, from initial bare-metal or cloud-based provisioning to complex upgrades and ongoing maintenance. This automation is primarily driven by Ansible, a powerful IT automation engine, which Kubespray leverages behind the scenes to execute tasks and manage configurations across a fleet of machines. The speaker clarified that while Ansible is the current implementation detail, Kubespray's design prioritizes the outcome—a stable, flexible Kubernetes cluster—over the specific automation tool, indicating potential for future adaptability.

The true strength of Kubespray, especially for AI workloads, is its unparalleled flexibility and configuration management. It supports a vast array of deployment scenarios:

  • Environments: Public clouds (AWS, GCP, Azure), private clouds, on-premise bare metal, and virtualized infrastructures.
  • Operating Systems: Compatibility across nearly all major Linux distributions.
  • Container Runtimes: Choice of container runtimes, with containerd being a common and efficient option.
  • Network Plugins: Support for various Container Network Interface (CNI) plugins, including popular choices like Calico and Cilium, which are crucial for implementing advanced network policies and ensuring high-performance communication within AI clusters.

For AI workloads, the integration of GPU management is paramount. Kubespray addresses this by facilitating the deployment and configuration of GPU operators. Specifically, the NVIDIA GPU operator was mentioned, which automates the provisioning of GPU drivers, CUDA, and other necessary components directly onto Kubernetes worker nodes. This means that once Kubespray deploys the cluster, GPUs are automatically detected and made available for scheduling AI workloads, eliminating manual setup complexities. This capability transforms raw hardware into immediately usable, accelerated computing resources within Kubernetes.

The architecture built with Kubespray for AI typically involves several layers:

  1. Hardware Layer: Diverse hardware, including NVIDIA DGX systems, A100 GPUs, or other smaller GPU configurations, can be used.
  2. Kubernetes Layer: Deployed and managed by Kubespray, providing the orchestration foundation.
  3. GPU Operator Layer: Integrates with Kubespray to automatically manage GPU resources.
  4. AI Frameworks: On top of Kubernetes, users deploy standard AI frameworks such as TensorFlow, PyTorch, and R. Kubespray ensures the underlying cluster is optimized and ready for these demanding applications.
  5. Workload Management & Serving: Advanced tools like Run:ai Euler are integrated for intelligent workload scheduling, preemption, and resource allocation, particularly useful for prioritizing inference over training or managing interactive sessions for data scientists. For model serving, Knative was highlighted, enabling serverless-style deployment of AI models.

Kubespray's commitment to stability and timely updates is backed by a rigorous Continuous Integration (CI) infrastructure. This CI system deploys 10,000 to 15,000 clusters every month and spins up 20 to 50 clusters for each PR, involving 200 to 300 VMs at peak. This extensive testing ensures that new features and updates are thoroughly vetted before release. The project consistently supports three stable versions of Kubernetes and aims to release updates within two to three months of upstream Kubernetes releases, a significantly faster pace than many commercial providers, ensuring users can leverage the latest Kubernetes innovations and security patches without undue delay. This robust testing and update strategy is critical for AI environments that often require the newest features and performance optimizations.

Demo / Proof of Concept

▶ Watch: Why Kubespray was chosen for hardware agnostic GPU management (7:00)

While Antoine Legrand initially mentioned a prepared demo, he opted for a more impactful approach: inviting a real-world Kubespray user to share their experiences. Luke Simmons, representing Vestia Transun, a healthcare company based in Sweden with 50,000 employees, provided a compelling user story that served as a powerful proof of concept for Kubespray's capabilities in an AI context.

Vestia Transun faced several significant challenges inherent to large enterprises attempting to build cutting-edge AI infrastructure:

  • Hardware Agnosticism: They needed the flexibility to "bring their own hardware," specifically mentioning NVIDIA DGX systems, A100 GPUs, and smaller, individual GPUs, and attach them to different worker nodes. This requirement stemmed from a desire to avoid vendor lock-in and optimize hardware procurement.
  • Kubernetes Version Management: A common pain point in enterprise environments is the difficulty in rapidly upgrading Kubernetes clusters to the latest versions. Vestia Transun required a solution that would enable them to quickly spin up and utilize the newest Kubernetes features, which are often crucial for leveraging the latest advancements in AI deployments.

Vestia Transun's solution was to choose Kubespray, primarily due to its hardware-agnostic nature. Simmons highlighted how Kubespray allowed them to integrate their diverse GPU infrastructure seamlessly. Their implementation details included:

  • GPU Integration: They utilized the NVIDIA GPU operator—as described by Antoine Legrand—to automatically manage and allocate GPU resources across their Kubespray-managed clusters.
  • AI Frameworks: On top of this foundation, they deployed popular AI frameworks such as TensorFlow, PyTorch, and R, which could then dynamically allocate GPU resources on the fly.
  • Workload Scheduling: For advanced workload management, they integrated Run:ai Euler. This tool enabled automatic preemption of workloads and prioritized critical tasks like inference and model serving over less urgent training jobs.
  • Dynamic Interactive Workloads: The platform supported dynamic allocation of interactive sessions, allowing their team of approximately 20 data scientists to develop and train models efficiently.
  • Model Serving: For deploying trained models, they leveraged Knative, which provided a serverless-like experience for serving AI models directly from their Kubernetes clusters.

The benefits experienced by Vestia Transun were substantial. Kubespray allowed them to:

  • Rapidly Spin Up/Down Environments: They could easily spin up test instances, test new configurations, and tear them down, facilitating agile development and experimentation.
  • Efficiently Utilize DGX Infrastructure: Kubespray enabled them to effectively deploy and manage their high-end DGX systems for demanding AI tasks.
  • Maintain Latest Kubernetes Versions: The ability to stay current with Kubernetes releases ensured their data scientists always had access to the most advanced features and optimizations for their AI workflows.

Simmons concluded by expressing that Kubespray had provided a "really, really fantastic experience," underscoring its practical value in driving enterprise-scale AI initiatives with flexibility and efficiency.

Defensive Implications

▶ Watch: Real-world AI workloads and tools running on Kubespray (8:00)

The strategic adoption of Kubespray, as outlined in the talk, carries significant defensive implications for organizations operating AI workloads on Kubernetes. By empowering vendor independence and offering granular control over infrastructure, Kubespray inherently strengthens an organization's security posture across several vectors.

Firstly, Kubespray enhances supply chain security for the Kubernetes infrastructure itself. Unlike managed Kubernetes services where the underlying installation process is opaque, Kubespray provides a transparent and auditable method for deploying clusters. This allows organizations to define, inspect, and control every component, from the operating system to the CNI plugin and container runtime. Defenders can ensure that only trusted, vetted components are used, reducing the risk of hidden backdoors or vulnerabilities introduced during the provisioning phase. This level of control is paramount in high-assurance environments.

Secondly, the ability to avoid vendor lock-in translates directly into a reduced attack surface and improved resilience. Relying on a single cloud provider inherently centralizes risk; a major outage or security incident affecting that provider could cripple operations. By enabling multi-cloud and hybrid cloud deployments, Kubespray allows organizations to diversify their infrastructure, distributing risk and potentially mitigating the impact of a targeted attack or widespread vulnerability affecting a single vendor. This strategic flexibility also means defenders are not beholden to a single provider's security incident response times or patching schedules.

Thirdly, Kubespray promotes a consistent security posture across diverse environments. Whether clusters are deployed on-premises, in AWS, GCP, or Azure, Kubespray ensures a standardized installation and configuration baseline. This consistency simplifies the application of security policies, compliance checks, and hardening guidelines. Defenders can implement uniform network policies (e.g., using Calico or Cilium), access controls, and auditing mechanisms across their entire Kubernetes fleet, reducing configuration drift and the likelihood of overlooked security gaps.

Fourthly, Kubespray's commitment to rapid patching and upgrades is a critical defensive advantage. Kubernetes, like any complex software, frequently has security vulnerabilities discovered. Kubespray's ability to support recent Kubernetes versions and facilitate upgrades within two to three months of upstream releases means organizations can apply critical security patches promptly. This agility is crucial for minimizing the window of exposure to known exploits, a capability that can be challenging to achieve with slower-moving enterprise or managed service providers.

Fifthly, for deployments involving on-premise GPUs, Kubespray grants organizations complete control over the physical and logical security of their AI accelerators. This includes managing direct network access, physical access controls, and ensuring that GPU operators (like the NVIDIA GPU operator) are configured securely. In scenarios handling sensitive data or proprietary AI models, this direct control over hardware security is invaluable.

Finally, the flexibility in choosing network plugins like Calico or Cilium empowers defenders to implement advanced network security measures. These CNIs offer robust capabilities for micro-segmentation, allowing fine-grained control over traffic between pods and namespaces. This can significantly limit lateral movement for attackers who manage to breach a single container or pod, containing potential incidents and protecting critical AI workloads.

Key Takeaways

  • Vendor Independence is Paramount: Kubespray's core value proposition is enabling vendor independence, which drives significant cost efficiencies by allowing organizations to shop around for infrastructure and aggressively negotiate pricing.
  • Flexible & Hardware-Agnostic Deployment: Kubespray provides a robust, flexible, and production-ready orchestrator for Kubernetes, capable of deploying across virtually any environment, from bare metal to multi-cloud, and integrating with diverse hardware, including various GPU types (e.g., NVIDIA DGX, A100s).
  • Seamless AI Stack Integration: It facilitates the seamless integration and management of GPUs via operators (like the NVIDIA GPU operator) and supports leading AI frameworks such as TensorFlow, PyTorch, and R, along with advanced workload schedulers like Run:ai Euler and model serving tools like Knative.
  • Strategic Hybrid & Multi-Cloud Capabilities: The ability to mix and match infrastructure, combining on-premise GPU resources with public cloud services, offers strategic advantages in resource optimization, cost control, and business continuity.
  • Stability & Timely Kubernetes Upgrades: Kubespray's rigorous CI/CD pipeline and commitment to supporting recent Kubernetes versions ensure stability and allow enterprises to quickly adopt the latest features and security patches crucial for cutting-edge AI development.
  • Abstraction for Control: By abstracting underlying automation tools like Ansible, Kubespray provides a consistent interface for cluster lifecycle management, empowering users with greater control and freedom over their Kubernetes infrastructure.

About the Speaker(s)

Antoine Legrand is a SITU at Conny GmbH and holds the distinction of being an initial maintainer of Kubespray. His extensive experience and long-standing contributions have been instrumental in shaping Kubespray into the robust and flexible Kubernetes orchestrator it is today. His insights during the talk were rooted in years of practical experience with large-scale Kubernetes deployments and the challenges of managing complex CI infrastructures.

Mohamed Zaian is a Senior Engineer at New York I See and also serves as a dedicated Kubespray maintainer. While he was unfortunately unable to attend the KubeCon EU presentation, his ongoing work and expertise are vital to the continued development and stability of the Kubespray project. Both speakers represent the deep community-driven spirit that has sustained Kubespray for over a decade.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk, while not a bleeding-edge zero-day, delivers a brutally honest and deeply technical assessment of achieving cost-efficiency and vendor independence for AI workloads using Kubespray. The speaker, an initial maintainer, demonstrates profound expertise, grounding the strategic discussion in concrete technical details, rigorous CI practices, and real-world user stories. It's a valuable session for anyone grappling with the escalating costs and vendor lock-in in AI infrastructure, offering actionable insights and a proven open-source path forward.

Heather Calloway (CISO) — MUST SEE

This KubeCon talk, by highlighting Kubespray's role in achieving vendor independence for AI workloads on Kubernetes, delivers a crucial message for executive leadership. It moves beyond technical implementation details to underscore how strategic infrastructure choices directly impact an organization's financial resilience, negotiation power, and overall control over its most critical computing assets. The emphasis on abstracting infrastructure to enable hardware agnosticism is a powerful antidote to the escalating costs and inherent risks of proprietary ecosystems, making this a must-watch for any leader grappling with the strategic implications of AI at scale.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025