Algorithms#

This table lists every algorithm available in RLlib. All algorithms support multi-GPU training on a single GPU node in Ray (open-source) (multi_gpu), and multi-GPU training on multi-node GPU clusters on the Anyscale platform (multi_node_multi_gpu).

On-policy#

Proximal Policy Optimization (PPO)#

[paper] [implementation]

../_images/ppo-architecture.svg

PPO architecture: In a training iteration, PPO performs three steps: 1. Sampling a set of episodes or episode fragments. 1. Converting these into a train batch and updating the model using a clipped objective and multiple SGD passes over this batch. 1. Syncing the weights from the Learners back to the EnvRunners. PPO scales out on both axes, supporting multiple EnvRunners for sample collection and multiple GPU- or CPU-based Learners for updating the model.#

Tuned examples: Pong-v5, CartPole-v1, Pendulum-v1.

PPO-specific configs: PPOConfig.training(). See also generic algorithm settings.

Off-policy#

Deep Q Networks (DQN, Rainbow, Parametric DQN)#

[paper] [implementation]

../_images/dqn-architecture.svg

DQN architecture: DQN uses a replay buffer to temporarily store episode samples that RLlib collects from the environment. Throughout different training iterations, these episodes and episode fragments are re-sampled from the buffer and reused for updating the model, before eventually being discarded when the buffer has reached capacity and new samples keep coming in (FIFO). This reuse of training data makes DQN sample-efficient and off-policy. DQN scales out on both axes, supporting multiple EnvRunners for sample collection and multiple GPU- or CPU-based Learners for updating the model.#

RLlib provides all the DQN improvements evaluated in Rainbow, though it doesn’t enable all of them by default. For parametric or variable-length action spaces on the new API stack, see the action masking example. The example uses PPO.

Tuned examples: CartPole-v1, multi-agent CartPole, StatelessCartPole, Atari benchmark.

Hint

For a complete rainbow setup, make the following changes to the default DQN config: "n_step": [between 1 and 10], "noisy": True, "num_atoms": [more than 1], "v_min": -10.0, "v_max": 10.0 (set v_min and v_max according to your expected range of returns).

DQN-specific configs: DQNConfig.training(). See also generic algorithm settings.

Soft Actor Critic (SAC)#

[original paper], [follow up paper], [implementation].

../_images/sac-architecture.svg

SAC architecture: SAC uses a replay buffer to temporarily store episode samples that RLlib collects from the environment. Throughout different training iterations, these episodes and episode fragments are re-sampled from the buffer and reused for updating the model, before eventually being discarded when the buffer has reached capacity and new samples keep coming in (FIFO). This reuse of training data makes SAC sample-efficient and off-policy. SAC scales out on both axes, supporting multiple EnvRunners for sample collection and multiple GPU- or CPU-based Learners for updating the model.#

Tuned examples: Pendulum-v1, HalfCheetah-v4.

SAC-specific configs: SACConfig.training(). See also generic algorithm settings.

High-throughput on- and off-policy#

Asynchronous Proximal Policy Optimization (APPO)#

Tip

APPO was originally published under the name “IMPACT”. RLlib’s APPO exactly matches the algorithm described in the paper.

[paper] [implementation]

../_images/appo-architecture.svg

APPO architecture: APPO is an asynchronous variant of Proximal Policy Optimization (PPO) based on the IMPALA architecture, but uses a surrogate policy loss with clipping to run multiple SGD passes per collected train batch. In a training iteration, APPO requests samples from all EnvRunners asynchronously and the collected episode samples are returned to the main algorithm process as Ray references rather than actual objects available on the local process. APPO then passes these episode references to the Learners for asynchronous updates of the model. RLlib doesn’t always sync back the weights to the EnvRunners right after a new model version is available. To account for the EnvRunners being off-policy, APPO uses a procedure called v-trace, described in the IMPALA paper. APPO scales out on both axes, supporting multiple EnvRunners for sample collection and multiple GPU- or CPU-based Learners for updating the model.#

Tuned examples: Pong-v5, Pendulum-v1.

APPO-specific configs: APPOConfig.training(). See also generic algorithm settings.

Importance Weighted Actor-Learner Architecture (IMPALA)#

[paper] [implementation]

../_images/impala-architecture.svg

IMPALA architecture: In a training iteration, IMPALA requests samples from all EnvRunners asynchronously and the collected episodes are returned to the main algorithm process as Ray references rather than actual objects available on the local process. IMPALA then passes these episode references to the Learners for asynchronous updates of the model. RLlib doesn’t always sync back the weights to the EnvRunners right after a new model version is available. To account for the EnvRunners being off-policy, IMPALA uses a procedure called v-trace, described in the paper. IMPALA scales out on both axes, supporting multiple EnvRunners for sample collection and multiple GPU- or CPU-based Learners for updating the model.#

Tuned examples: Pong-v5, CartPole-v1, multi-agent TicTacToe.

../_images/impala.png

Multi-GPU IMPALA scales up to solve PongNoFrameskip-v4 in ~3 minutes using a pair of V100 GPUs and 128 CPU workers. The maximum training throughput reached is ~30k transitions per second (~120k environment frames per second).#

IMPALA-specific configs: IMPALAConfig.training(). See also generic algorithm settings.

Model-based RL#

DreamerV3#

[paper] [implementation] [RLlib readme]

See the README for how to run experiments with DreamerV3.

../_images/dreamerv3-architecture.svg

DreamerV3 architecture: DreamerV3 trains a recurrent WORLD_MODEL in supervised fashion using real environment interactions sampled from a replay buffer. The world model’s objective is to correctly predict the transition dynamics of the RL environment: next observation, reward, and a boolean continuation flag. DreamerV3 trains the actor and critic networks on synthesized trajectories only, which are “dreamed” by the WORLD_MODEL. The algorithm scales out on both axes, supporting multiple EnvRunner actors for sample collection and multiple GPU- or CPU-based Learner actors for updating the model. It can also be used in different environment types, including those with image-based or vector-based observations, continuous or discrete actions, as well as sparse or dense reward functions.#

Tuned examples: Atari 100k, Atari 200M, DeepMind Control Suite.

Pong-v5 results (1, 2, and 4 GPUs):

../_images/pong_1_2_and_4gpus.svg

Episode mean rewards for the Pong-v5 environment, using the “100k” setting that allows only 100k environment steps. Despite the stable sample efficiency, shown by the constant learning performance per environment step, the wall time improves almost linearly from one to four GPUs. Left: Episode reward over environment timesteps sampled. Right: Episode reward over wall-time.#

Atari 100k results (1 vs 4 GPUs):

../_images/atari100k_1_vs_4gpus.svg

Episode mean rewards for various Atari 100k tasks on one versus four GPUs. Left: Episode reward over environment timesteps sampled. Right: Episode reward over wall-time.#

DeepMind Control Suite (vision) results (1 vs 4 GPUs):

../_images/dmc_1_vs_4gpus.svg

Episode mean rewards for various DeepMind Control Suite tasks on one versus four GPUs. Left: Episode reward over environment timesteps sampled. Right: Episode reward over wall-time.#

Offline RL and imitation learning#

Behavior Cloning (BC)#

[paper] [implementation]

../_images/bc-architecture.svg

BC architecture: RLlib’s behavioral cloning (BC) uses Ray Data to tap into its parallel data processing capabilities. In one training iteration, BC reads episodes in parallel from offline files, for example parquet, by the n DataWorkers. Connector pipelines then preprocess these episodes into train batches and send these as data iterators directly to the n Learners for updating the model. RLlib’s BC implementation derives directly from its MARWIL implementation. The only difference is the beta parameter, set to 0.0. This makes BC try to match the behavior policy, which generated the offline data, disregarding any resulting rewards.#

Tuned examples: CartPole-v1, Pendulum-v1.

BC-specific configs: BCConfig inherits its training() settings from MARWILConfig.training(). See also generic algorithm settings.

Conservative Q-Learning (CQL)#

[paper] [implementation]

../_images/cql-architecture.svg

CQL architecture: CQL (Conservative Q-Learning) is an offline RL algorithm that mitigates the overestimation of Q-values outside the dataset distribution through a conservative critic estimate. It adds a simple Q regularizer loss to the standard Bellman update loss, ensuring that the critic doesn’t output overly optimistic Q-values. The SACLearner adds this conservative correction term to the TD-based Q-learning loss.#

Tuned examples: Pendulum-v1.

CQL-specific configs: CQLConfig.training(). See also generic algorithm settings.

Implicit Q-Learning (IQL)#

[paper] [implementation]

IQL architecture: IQL (Implicit Q-Learning) is an offline RL algorithm that never needs to evaluate actions outside of the dataset, yet still improves the learned policy substantially over the best behavior in the data through generalization. Instead of standard TD-error minimization, it introduces a value function trained through expectile regression, which yields a conservative estimate of returns. It improves the policy through advantage-weighted behavior cloning, which ensures safer generalization without explicit exploration.

The IQLLearner replaces the usual TD-based value loss with an expectile regression loss, and trains the policy to imitate high-advantage actions to achieve substantial performance gains over the behavior policy using only in-dataset actions.

Tuned examples: Pendulum-v1.

IQL-specific configs: IQLConfig.training(). See also generic algorithm settings.

Monotonic Advantage Re-Weighted Imitation Learning (MARWIL)#

[paper] [implementation]

../_images/marwil-architecture.svg

MARWIL architecture: MARWIL is a hybrid imitation learning and policy gradient algorithm suitable for training on batched historical data. When the beta hyperparameter is set to zero, the MARWIL objective reduces to plain imitation learning, the same as BC. MARWIL uses Ray Data to tap into its parallel data processing capabilities. In one training iteration, MARWIL reads episodes in parallel from offline files, for example parquet, by the n DataWorkers. Connector pipelines preprocess these episodes into train batches and send these as data iterators directly to the n Learners for updating the model.#

Tuned examples: CartPole-v1.

MARWIL-specific configs: MARWILConfig.training(). See also generic algorithm settings.

Algorithm extensions and plugins#

Curiosity-driven Exploration by Self-supervised Prediction#

[paper] [implementation]

../_images/curiosity-architecture.svg

Intrinsic Curiosity Model (ICM) architecture: The main idea behind ICM is to train a world-model in parallel with the “main” policy to predict the environment’s dynamics. The loss of the world model is the intrinsic reward that the ICMLearner adds to the environment’s extrinsic reward. In regions of the environment that are relatively unknown, where the world model predicts poorly what happens next, the artificial intrinsic reward is large, so the agent explores these unknown regions. RLlib’s curiosity implementation works with any RLlib algorithm. See the example implementations on top of PPO and DQN. ICM uses the chosen Algorithm’s training_step() as-is, but then executes the following additional steps during LearnerGroup.update: Duplicate the train batch of the “main” policy and use it for performing a self-supervised update of the ICM. Use the ICM to compute the intrinsic rewards and add these to the extrinsic environment rewards. Then continue updating the “main” policy.#

Tuned examples: 12x12 FrozenLake-v1.