Configuration#
These APIs configure Serve’s proxies, controller, autoscaling, request routing, and gang scheduling. Most live in ray.serve.config. The autoscaling policy helpers live in ray.serve.autoscaling_policy.
Options and policies#
Config for where to run proxies to receive ingress traffic to the cluster. |
|
Rich context provided to custom autoscaling policies. |
|
|
The default autoscaling policy based on basic thresholds for scaling. |
A scalar returned by an instant Prometheus query. |
|
One labeled sample in a Prometheus instant vector. |
|
An instant vector returned by a Prometheus query. |
|
Keeps Prometheus query results fresh for an autoscaling policy. |
|
PublicAPI (alpha): This API is in alpha and may change before becoming stable. |
|
Placement strategy for replicas within a gang. |
|
Policy for handling runtime failures of replicas in a gang. |
Configuration models#
Options for the Serve controller actor. |
|
gRPC options for the proxies. |
|
HTTP options for the proxies. |
|
Config for the Serve Autoscaler. |
|
Config for the HTTP response returned on backpressure rejections. |
|
Config for the Serve request router. |
|
Configuration for gang scheduling of deployment replicas. |
|
Configuration for a deployment-scoped actor. |