Ray Serve API#
Python API#
Writing Applications#
Class (or function) decorated with the |
|
One or more deployments bound with arguments that can be deployed together. |
Deployment Decorators#
Decorator that converts a Python class to a |
|
Wrap a deployment class with an ASGI application for HTTP request parsing. |
|
Converts a function to asynchronously handle batches. |
|
Wrap a callable or method used to load multiplexed models in a replica. |
Deployment Handles#
Note
The deprecated RayServeHandle and RayServeSyncHandle APIs have been fully removed as of Ray 2.10. See the model composition guide for how to update code to use the DeploymentHandle API instead.
A handle used to make requests to a deployment at runtime. |
|
A future-like object wrapping the result of a unary deployment handle call. |
|
A future-like object wrapping the result of a streaming deployment handle call. |
|
Wraps the results of a broadcast call to all replicas of a deployment. |
Running Applications#
Start Serve on the cluster. |
|
Run an application and return a handle to its ingress deployment. |
|
Delete an application by its name. |
|
Get the status of Serve on the cluster. |
|
Completely shut down Serve on the cluster. |
|
Completely shut down Serve on the cluster asynchronously. |
Configurations#
Config for where to run proxies to receive ingress traffic to the cluster. |
|
Rich context provided to custom autoscaling policies. |
|
|
The default autoscaling policy based on basic thresholds for scaling. |
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
Placement strategy for replicas within a gang. |
|
Policy for handling runtime failures of replicas in a gang. |
Options for the Serve controller actor. |
|
gRPC options for the proxies. |
|
HTTP options for the proxies. |
|
Config for the Serve Autoscaler. |
|
Config for the HTTP response returned on backpressure rejections. |
|
Config for the Serve request router. |
|
Configuration for gang scheduling of deployment replicas. |
|
Configuration for a deployment-scoped actor. |
Schemas#
Detailed info about a Ray Serve actor. |
|
Detailed info about a Ray Serve ProxyActor. |
|
Describes the status of an application and all its deployments. |
|
Describes the status of Serve. |
|
Describes the status of a deployment. |
|
Encoding type for the serve logs. |
|
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
One autoscaling decision with minimal provenance. |
|
Deployment-level autoscaler observability. |
|
Replica rank model. |
Abstract base class for task processing adapters. |
Request Router#
A unique identifier for a replica. |
|
A request that is pending execution by a replica. |
|
Contains info on a running replica. |
|
Mixin for FIFO routing. |
|
Mixin for locality routing. |
|
Mixin for multiplex routing. |
|
Abstract interface for a request router (how the router calls it). |
Advanced APIs#
Returns the deployment and replica tag from within a replica at runtime. |
|
Get the current OpenTelemetry trace context. |
|
Get a handle to a deployment-scoped actor by name. |
|
Stores runtime context info for replicas. |
|
Context information for a replica that is part of a gang. |
|
Get the multiplexed model ID for the current request. |
|
Get a handle to the application's ingress deployment by name. |
|
Get a handle to a deployment by name. |
|
Context manager to set and get gRPC context. |
|
Async iterator wrapping an incoming gRPC request stream. |
|
Raised when max_queued_requests is exceeded on a DeploymentHandle. |
|
Raise when a Serve request is cancelled. |
|
Internal exception that wraps an exception with user-set gRPC status code. |
|
Raised when a Serve deployment is unavailable to receive requests. |
|
Raised when the selected replica is no longer available. |
Command Line Interface (CLI)#
The serve command line interface deploys, inspects, and manages Serve applications from the command line. See the Serve CLI reference for the full command list.
Serve REST API#
The Serve REST API is exposed at the same port as the Ray dashboard. The dashboard port is 8265 by default. This port can be changed using the --dashboard-port argument when running ray start. All example requests in this section use the default port.
PUT "/api/serve/applications/"#
Declaratively deploys a list of Serve applications. If Serve is already running on the Ray cluster, removes all applications not listed in the new config. If Serve is not running on the Ray cluster, starts Serve. See multi-app config schema for the request’s JSON schema.
Example Request:
PUT /api/serve/applications/ HTTP/1.1
Host: http://localhost:8265/
Accept: application/json
Content-Type: application/json
{
"applications": [
{
"name": "text_app",
"route_prefix": "/",
"import_path": "text_ml:app",
"runtime_env": {
"working_dir": "https://github.com/ray-project/serve_config_examples/archive/HEAD.zip"
},
"deployments": [
{"name": "Translator", "user_config": {"language": "french"}},
{"name": "Summarizer"},
]
},
]
}
Example Response
HTTP/1.1 200 OK
Content-Type: application/json
GET "/api/serve/applications/"#
Gets cluster-level info and comprehensive details on all Serve applications deployed on the Ray cluster. See metadata schema for the response’s JSON schema.
GET /api/serve/applications/ HTTP/1.1
Host: http://localhost:8265/
Accept: application/json
Example Response (abridged JSON):
HTTP/1.1 200 OK
Content-Type: application/json
{
"controller_info": {
"node_id": "cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec",
"node_ip": "10.0.29.214",
"actor_id": "1d214b7bdf07446ea0ed9d7001000000",
"actor_name": "SERVE_CONTROLLER_ACTOR",
"worker_id": "adf416ae436a806ca302d4712e0df163245aba7ab835b0e0f4d85819",
"log_file_path": "/serve/controller_29778.log"
},
"proxy_location": "EveryNode",
"http_options": {
"host": "0.0.0.0",
"port": 8000,
"root_path": "",
"request_timeout_s": null,
"keep_alive_timeout_s": 5
},
"grpc_options": {
"port": 9000,
"grpc_servicer_functions": [],
"request_timeout_s": null
},
"proxies": {
"cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec": {
"node_id": "cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec",
"node_ip": "10.0.29.214",
"actor_id": "b7a16b8342e1ced620ae638901000000",
"actor_name": "SERVE_CONTROLLER_ACTOR:SERVE_PROXY_ACTOR-cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec",
"worker_id": "206b7fe05b65fac7fdceec3c9af1da5bee82b0e1dbb97f8bf732d530",
"log_file_path": "/serve/http_proxy_10.0.29.214.log",
"status": "HEALTHY"
}
},
"applications": {
"app1": {
"name": "app1",
"route_prefix": "/",
"docs_path": null,
"status": "RUNNING",
"message": "",
"last_deployed_time_s": 1694042836.1912267,
"deployed_app_config": {
"name": "app1",
"route_prefix": "/",
"import_path": "src.text-test:app",
"deployments": [
{
"name": "Translator",
"num_replicas": 1,
"user_config": {
"language": "german"
}
}
]
},
"deployments": {
"Translator": {
"name": "Translator",
"status": "HEALTHY",
"message": "",
"deployment_config": {
"name": "Translator",
"num_replicas": 1,
"max_ongoing_requests": 100,
"user_config": {
"language": "german"
},
"graceful_shutdown_wait_loop_s": 2.0,
"graceful_shutdown_timeout_s": 20.0,
"health_check_period_s": 10.0,
"health_check_timeout_s": 30.0,
"ray_actor_options": {
"runtime_env": {
"env_vars": {}
},
"num_cpus": 1.0
},
"is_driver_deployment": false
},
"replicas": [
{
"node_id": "cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec",
"node_ip": "10.0.29.214",
"actor_id": "4bb8479ad0c9e9087fee651901000000",
"actor_name": "SERVE_REPLICA::app1#Translator#oMhRlb",
"worker_id": "1624afa1822b62108ead72443ce72ef3c0f280f3075b89dd5c5d5e5f",
"log_file_path": "/serve/deployment_Translator_app1#Translator#oMhRlb.log",
"replica_id": "app1#Translator#oMhRlb",
"state": "RUNNING",
"pid": 29892,
"start_time_s": 1694042840.577496
}
]
},
"Summarizer": {
"name": "Summarizer",
"status": "HEALTHY",
"message": "",
"deployment_config": {
"name": "Summarizer",
"num_replicas": 1,
"max_ongoing_requests": 100,
"user_config": null,
"graceful_shutdown_wait_loop_s": 2.0,
"graceful_shutdown_timeout_s": 20.0,
"health_check_period_s": 10.0,
"health_check_timeout_s": 30.0,
"ray_actor_options": {
"runtime_env": {},
"num_cpus": 1.0
},
"is_driver_deployment": false
},
"replicas": [
{
"node_id": "cef533a072b0f03bf92a6b98cb4eb9153b7b7c7b7f15954feb2f38ec",
"node_ip": "10.0.29.214",
"actor_id": "7118ae807cffc1c99ad5ad2701000000",
"actor_name": "SERVE_REPLICA::app1#Summarizer#cwiPXg",
"worker_id": "12de2ac83c18ce4a61a443a1f3308294caf5a586f9aa320b29deed92",
"log_file_path": "/serve/deployment_Summarizer_app1#Summarizer#cwiPXg.log",
"replica_id": "app1#Summarizer#cwiPXg",
"state": "RUNNING",
"pid": 29893,
"start_time_s": 1694042840.5789504
}
]
}
}
}
}
}
DELETE "/api/serve/applications/"#
Shuts down Serve and all applications running on the Ray cluster. Has no effect if Serve is not running on the Ray cluster.
Example Request:
DELETE /api/serve/applications/ HTTP/1.1
Host: http://localhost:8265/
Accept: application/json
Example Response
HTTP/1.1 200 OK
Content-Type: application/json
Config Schemas#
Multi-application config for deploying a list of Serve applications to the Ray cluster. |
|
Options to start the gRPC Proxy with. |
|
Options to start the HTTP Proxy with. |
|
Describes one Serve application, and currently can also be used as a standalone config to deploy a single application to a Ray cluster. |
|
Options with which to start a replica actor. |
|
Celery adapter config. You can use it to configure the Celery task processor for your Serve application. |
|
Task processor config. You can use it to configure the task processor for your Serve application. |
|
Task result Model. |
|
Request schema for scaling a deployment's replicas. |
Response Schemas#
Serve metadata with system-level info and details on all applications deployed to the Ray cluster. |
|
Detailed info about a Serve application. |
|
Detailed info about a deployment within a Serve application. |
|
Detailed info about a single deployment replica. |
|
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
Represents a node in the deployment topology. |
|
Represents the dependency graph of deployments in an application. |
|
Health metrics for the Ray Serve controller. |
|
Statistics for a collection of duration/latency measurements. |
Tracks the type of API that an application originates from. |
|
The current status of the application. |
|
The current status of the proxy. |
Observability#
A serve cumulative metric that is monotonically increasing. |
|
Tracks the size and number of events in buckets. |
|
Gauges keep the last recorded value and drop everything before. |
Logging config schema for configuring serve components logs. |
|
Tracing config schema for configuring distributed tracing on Serve components. |
LLM API#
Builders#
Helper to build a single vllm deployment from the given llm config. |
|
Helper to build an OpenAI compatible app with the llm deployment setup from the given llm serving args. |
Configs#
The configuration for starting an LLM deployment. |
|
The configuration for starting an LLM deployment application. |
|
The configuration for loading an LLM model. |
|
The configuration for mirroring an LLM model from cloud storage. |
|
The configuration for loading an LLM model with LoRA. |
Deployments#
The implementation of the vLLM engine deployment. |
|
Decode-side LLM server for prefill-decode disaggregation. |
|
Prefill-side LLM server for prefill-decode disaggregation. |
|
Data Parallel LLM Server. |