KubeRay#
KubeRay is the officially supported Kubernetes operator for Ray. It provides a Kubernetes-native way to deploy and manage Ray clusters. KubeRay runs each Ray node as a Kubernetes Pod, so each Ray cluster consists of a head Pod and a collection of worker Pods.
KubeRay adds four custom resources:
RayCluster: Creates a Ray cluster and manages its lifecycle, including autoscaling and fault tolerance.
RayJob: Creates a RayCluster, submits a Ray job when the cluster is ready, and can delete the RayCluster when the job finishes.
RayService: Runs Ray Serve applications on a RayCluster, with zero-downtime upgrades and high availability.
RayCronJob: Creates RayJobs on a recurring cron schedule. RayCronJob is in alpha and requires KubeRay 1.6.0 or later.
With optional autoscaling, KubeRay sizes your Ray clusters to the requirements of your Ray workload, adding and removing Pods as needed. KubeRay supports heterogeneous compute nodes, including GPUs, and runs multiple Ray clusters with different Ray versions in the same Kubernetes cluster.
To learn the basics of KubeRay and run your first Ray application with it, see Getting Started with KubeRay and the following quickstart guides:
Learn more#
Use the following guides to deploy, configure, and operate Ray clusters with KubeRay.
Getting started
Install the KubeRay operator, create your first Ray cluster, and run a Ray application on it.
User guides
Configure autoscaling, fault tolerance, observability, and security for the Ray clusters KubeRay manages.
Examples
Run example Ray workloads with KubeRay.
Ecosystem
Integrate KubeRay with third-party Kubernetes ecosystem tools.
Benchmarks
Check the KubeRay benchmark results.
Troubleshooting
Consult the KubeRay troubleshooting guides.
About KubeRay#
KubeRay lives in the KubeRay GitHub repository, under the broader Ray project. Several companies use KubeRay to run production Ray deployments. To track progress, report bugs, propose new features, or contribute to the project, visit the repository.