ConnectorV2 and ConnectorV2 pipelines#

RLlib stores and transports all trajectory data as SingleAgentEpisode or MultiAgentEpisode objects. Connector pipelines translate this episode data into tensor batches that neural network models read right before the model forward pass.

../_images/generic_connector_pipeline.svg

Generic ConnectorV2 Pipeline: All pipelines consist of one or more ConnectorV2 pieces. When you call the pipeline, you pass in a list of episodes, the RLModule instance, and a batch, which might start as an empty dictionary. Each ConnectorV2 piece takes its predecessor’s output, starting on the left with the batch, transforms the episodes, the batch, or both, and passes everything to the next piece. Each ConnectorV2 piece can read from and write to the provided episodes, add data from these episodes to the batch, or change data that’s already in the batch. The pipeline returns the output batch of the last piece.#

Note

The batch output of the pipeline lives only as long as the succeeding RLModule forward pass or Env.step() call. RLlib discards the data afterward. The list of episodes, however, might persist longer. For example, if an env-to-module pipeline reads an observation from an episode, mutates that observation, and writes it back into the episode, the subsequent module-to-env pipeline can see the changed observation. The Learner pipeline operates on the same episodes that already passed through both the env-to-module and module-to-env pipelines, so those episodes might have changed.

Three ConnectorV2 pipeline types#

RLlib has three types of connector pipelines:

  1. Env-to-module pipeline, which creates tensor batches for the forward passes that compute actions.

  2. Module-to-env pipeline, which translates a model’s output into RL environment actions. Documentation for this pipeline is pending.

  3. Learner connector pipeline, which creates the train batch for a model update.

The ConnectorV2 API is a tool for customizing your RLlib experiments and algorithms. With it, you take full control over how RLlib accesses, changes, and reassembles the episode data collected from your RL environments or offline RL input files. You also control the exact shape and content of the tensor batches that RLlib feeds into your models to compute actions or losses.

../_images/location_of_connector_pipelines_in_rllib.svg

ConnectorV2 Pipelines: The env-to-module and Learner pipelines convert episodes into batched data that your model can process. The module-to-env pipeline converts your model’s output into action batches that your possibly vectorized RL environment needs for stepping. The env-to-module pipeline, located on an EnvRunner, takes a list of episodes as input and outputs a batch for an RLModule forward pass that computes the next action. The module-to-env pipeline on the same EnvRunner takes the output of that RLModule and converts it into actions for the next call to your RL environment’s step() method. Lastly, a Learner connector pipeline, located on a Learner worker, converts a list of episodes into a train batch for the next RLModule update.#

The following pages discuss the three pipeline types in more detail. All three share these characteristics:

  • All connector pipelines are sequences of one or more ConnectorV2 pieces. You can nest these, so some pieces might themselves be connector pipelines.

  • All connector pieces and pipelines are Python callables that override the __call__() method.

  • The call signatures are uniform across the pipeline types. The main required arguments are the list of episodes, the batch to build, and the RLModule instance. See the __call__() method for details.

  • All connector pipelines can read from and write to the provided list of episodes and the batch, performing data transforms as needed.

Batch construction phases and formats#

When you push a list of input episodes through a connector pipeline, the pipeline constructs a batch from that data. The batch always starts as an empty Python dictionary and passes through several formats and phases as it moves through the pipeline’s pieces.

The following applies to all env-to-module and learner connector pipelines. Documentation for the learner connector pipeline is in progress.

../_images/pipeline_batch_phases_single_agent.svg

Batch construction phases and formats: In the standard single-agent case, where only one ModuleID, DEFAULT_MODULE_ID, exists, the batch starts as an empty dictionary on the left, then undergoes a “collect data” phase, in which connector pieces add individual items to the batch. Each piece stores an item under two keys: the column name, such as obs or rewards, and the episode ID it extracted the item from. In most cases, your custom connector pieces operate during this phase. Once all custom pieces finish their data insertions and transforms, the AgentToModuleMapping default piece performs a “reorganize by ModuleID” operation in the center, during which the batch’s dictionary hierarchy changes to put the ModuleID DEFAULT_MODULE_ID at the top level and the column names below it. At the lowest level of the batch, data items still reside in Python lists. Finally, the BatchIndividualItems default piece creates NumPy arrays from the Python lists, batching all data on the right.#

For multi-agent setups, where more than one ModuleID exists, the AgentToModuleMapping default connector piece ensures that the output batch maps each module ID to that module’s forward batch:

../_images/pipeline_batch_phases_multi_agent.svg

Batch construction for multi-agent: In a multi-agent setup, the default AgentToModuleMapping connector piece reorganizes the batch by ModuleID, then by column names, so that a MultiRLModule can loop through its submodules and give each one a batch for the forward pass.#

RLlib’s MultiRLModule splits the forward pass into individual submodule forward passes, using the batch under each ModuleID. See how to write your own multi-module or multi-agent forward logic to override this default behavior.

If you have a stateful RLModule, such as an LSTM, RLlib adds two more default connector pieces to the pipeline, AddTimeDimToBatchAndZeroPad and AddStatesFromEpisodesToBatch:

../_images/pipeline_batch_phases_single_agent_w_states.svg

Batch construction for stateful models: For stateful RLModule instances, RLlib adds two more default connector pieces to the pipeline. The AddTimeDimToBatchAndZeroPad piece converts all lists of individual data items on the lowest batch level into sequences of a fixed length, max_seq_len, and zero-pads these when it encounters an episode end. To set max_seq_len, see the note below. The AddStatesFromEpisodesToBatch piece adds the state_out values that your RLModule previously generated to the batch under the state_in column name. RLlib adds the state_in values only for the first timestep in each sequence, so it doesn’t add a time dimension to the data in the state_in column.#

Note

To change the zero-padded sequence length for the AddTimeDimToBatchAndZeroPad connector, set max_seq_len in your config. For custom models:

config.rl_module(model_config={"max_seq_len": ...})

For RLlib’s default models:

from ray.rllib.core.rl_module.default_model_config import DefaultModelConfig

config.rl_module(model_config=DefaultModelConfig(max_seq_len=...))