setup_global_ray_cluster#

ray.util.spark.setup_global_ray_cluster(*, max_worker_nodes: int, is_blocking: bool = True, min_worker_nodes: int | None = None, num_cpus_worker_node: int | None = None, num_cpus_head_node: int | None = None, num_gpus_worker_node: int | None = None, num_gpus_head_node: int | None = None, memory_worker_node: int | None = None, memory_head_node: int | None = None, object_store_memory_worker_node: int | None = None, object_store_memory_head_node: int | None = None, head_node_options: Dict | None = None, worker_node_options: Dict | None = None, strict_mode: bool = False, collect_log_to_path: str | None = None, autoscale_upscaling_speed: float | None = 1.0, autoscale_idle_timeout_minutes: float | None = 1.0)[source]#

Set up a global mode cluster. The global Ray on spark cluster means: - You can only create one active global Ray on spark cluster at a time. On databricks cluster, the global Ray cluster can be used by all users, - as contrast, non-global Ray cluster can only be used by current notebook user. - It is up persistently without automatic shutdown. - On databricks notebook, you can connect to the global cluster by calling ray.init() without specifying its address, it will discover the global cluster automatically if it is up.

For global mode, the ray_temp_root_dir argument is not supported. Global model Ray cluster always use the default Ray temporary directory path.

All arguments are the same with setup_ray_cluster API except that: - the ray_temp_root_dir argument is not supported. Global model Ray cluster always use the default Ray temporary directory path. - A new argument “is_blocking” (default True) is added. If “is_blocking” is True, then keep the call blocking until it is interrupted. once the call is interrupted, the global Ray on spark cluster is shut down and setup_global_ray_cluster call terminates. If “is_blocking” is False, once Ray cluster setup completes, return immediately.