User guides#
These guides show you how to complete common tasks with Ray Data. If you’re new to Ray Data, start with the Ray Data quickstart.
- Loading data
- Inspecting data
- Transforming Data
- Aggregating data
- Iterating over data
- Joining data
- Shuffling data
- Weighted dataset mixing
- Saving data
- Working with images
- Working with text
- Working with tensors and NumPy
- Working with PyTorch
- Working with LLMs
- Quickstart: Run batch inference with vLLM
- How does the processor pipeline work?
- Scale to multiple GPUs
- Generate text
- Run batch inference on multimodal data
- Generate embeddings
- Run classification models
- Query OpenAI-compatible endpoints
- Configure tokenization disaggregation
- Use custom tokenizers
- How does Ray Data LLM handle failures?
- Advanced configuration
- Troubleshooting
- Usage data collection
- How to avoid out-of-memory errors (OOMs)
- Monitoring your workload
- Execution configurations
- Run multiple Datasets in one cluster
- End-to-end: Offline Batch Inference
- Advanced: Performance tips and tuning
- Advanced: Scaling out expensive collate functions
- Advanced: Read and write custom file types