LLM API#
The ray.serve.llm module builds and configures Serve applications that serve large language models.
Builders#
Helper to build a single vllm deployment from the given llm config. |
|
Helper to build an OpenAI compatible app with the llm deployment setup from the given llm serving args. |
Configs#
The configuration for starting an LLM deployment. |
|
The configuration for starting an LLM deployment application. |
|
The configuration for loading an LLM model. |
|
The configuration for mirroring an LLM model from cloud storage. |
|
The configuration for loading an LLM model with LoRA. |
Deployments#
The implementation of the vLLM engine deployment. |
|
Decode-side LLM server for prefill-decode disaggregation. |
|
Prefill-side LLM server for prefill-decode disaggregation. |
|
Data Parallel LLM Server. |