Configuration#

These APIs configure Serve’s proxies, controller, autoscaling, request routing, and gang scheduling. Most live in ray.serve.config. The autoscaling policy helpers live in ray.serve.autoscaling_policy.

Options and policies#

serve.config.ProxyLocation

Config for where to run proxies to receive ingress traffic to the cluster.

serve.config.AutoscalingContext

Rich context provided to custom autoscaling policies.

serve.autoscaling_policy.replica_queue_length_autoscaling_policy

The default autoscaling policy based on basic thresholds for scaling.

serve.autoscaling_policy.PrometheusScalar

A scalar returned by an instant Prometheus query.

serve.autoscaling_policy.PrometheusSample

One labeled sample in a Prometheus instant vector.

serve.autoscaling_policy.PrometheusVector

An instant vector returned by a Prometheus query.

serve.autoscaling_policy.PrometheusQueryMixin

Keeps Prometheus query results fresh for an autoscaling policy.

serve.config.AggregationFunction

PublicAPI (alpha): This API is in alpha and may change before becoming stable.

serve.config.GangPlacementStrategy

Placement strategy for replicas within a gang.

serve.config.GangRuntimeFailurePolicy

Policy for handling runtime failures of replicas in a gang.

Configuration models#

serve.config.ControllerOptions

Options for the Serve controller actor.

serve.config.gRPCOptions

gRPC options for the proxies.

serve.config.HTTPOptions

HTTP options for the proxies.

serve.config.AutoscalingConfig

Config for the Serve Autoscaler.

serve.config.AutoscalingPolicy

serve.config.BackpressureConfig

Config for the HTTP response returned on backpressure rejections.

serve.config.RequestRouterConfig

Config for the Serve request router.

serve.config.GangSchedulingConfig

Configuration for gang scheduling of deployment replicas.

serve.config.DeploymentActorConfig

Configuration for a deployment-scoped actor.