learners#
- AlgorithmConfig.learners(*, num_learners: int | None = <ray.rllib.utils.from_config._NotProvided object>, num_cpus_per_learner: str | float | int | None = <ray.rllib.utils.from_config._NotProvided object>, num_gpus_per_learner: float | int | None = <ray.rllib.utils.from_config._NotProvided object>, custom_resources_per_learner: ~typing.Dict[str, float] | None = <ray.rllib.utils.from_config._NotProvided object>, num_aggregator_actors_per_learner: int | None = <ray.rllib.utils.from_config._NotProvided object>, max_requests_in_flight_per_aggregator_actor: float | None = <ray.rllib.utils.from_config._NotProvided object>, local_gpu_idx: int | None = <ray.rllib.utils.from_config._NotProvided object>, max_requests_in_flight_per_learner: int | None = <ray.rllib.utils.from_config._NotProvided object>, never_skip_update: bool | None = <ray.rllib.utils.from_config._NotProvided object>, learner_class: ~typing.Type[Learner] | None = <ray.rllib.utils.from_config._NotProvided object>, learner_connector: ~typing.Callable[[gymnasium.spaces.Space, gymnasium.spaces.Space], ConnectorV2 | ~typing.List[ConnectorV2]] | None = <ray.rllib.utils.from_config._NotProvided object>, add_default_connectors_to_learner_pipeline: bool | None = <ray.rllib.utils.from_config._NotProvided object>, learner_config_dict: ~typing.Dict[str, ~typing.Any] | None = <ray.rllib.utils.from_config._NotProvided object>) Self[source]#
Sets LearnerGroup and Learner worker related configurations.
- Parameters:
num_learners (int | None) – Number of Learner workers used for updating the RLModule. A value of 0 means training takes place on a local Learner on main process CPUs or 1 GPU (determined by
num_gpus_per_learner). For multi-gpu training, you have to setnum_learnersto > 1 and setnum_gpus_per_learneraccordingly (e.g., 4 GPUs total and model fits on 1 GPU:num_learners=4; num_gpus_per_learner=1OR 4 GPUs total and model requires 2 GPUs:num_learners=2; num_gpus_per_learner=2).num_cpus_per_learner (str | float | int | None) – Number of CPUs allocated per Learner worker. If “auto” (default), use 1 if
num_gpus_per_learner=0, otherwise 0. Only necessary for custom processing pipeline inside each Learner requiring multiple CPU cores. Ifnum_learners=0, RLlib creates only one local Learner instance and the number of CPUs on the main process ismax(num_cpus_per_learner, num_cpus_for_main_process).num_gpus_per_learner (float | int | None) – Number of GPUs allocated per Learner worker. If
num_learners=0, any value greater than 0 runs the training on a single GPU on the main process, while a value of 0 runs the training on main process CPUs.custom_resources_per_learner (Dict[str, float] | None) – Any custom Ray resources to allocate per Learner worker. Useful for pinning Learners to specific nodes via custom resource labels. Note: do NOT put
"CPU"or"GPU"in here – usenum_cpus_per_learnerandnum_gpus_per_learnerinstead.num_aggregator_actors_per_learner (int | None) – The number of aggregator actors per Learner (if num_learners=0, one local learner is created). Must be at least 1. Aggregator actors perform the task of a) converting episodes into a train batch and b) move that train batch to the same GPU that the corresponding learner is located on. Good values are 1 or 2, but this strongly depends on your setup and
EnvRunnerthroughput.max_requests_in_flight_per_aggregator_actor (float | None) – How many in-flight requests are allowed per aggregator actor before new requests are dropped?
local_gpu_idx (int | None) – If
num_gpus_per_learner> 0, andnum_learners< 2, then RLlib uses this GPU index for training. This is an index into the available CUDA devices. For example ifos.environ["CUDA_VISIBLE_DEVICES"] = "1"andlocal_gpu_idx=0, RLlib uses the GPU with ID=1 on the node.max_requests_in_flight_per_learner (int | None) – Max number of in-flight requests to each Learner (actor). You normally do not have to tune this setting (default is 3), however, for asynchronous algorithms, this determines the “queue” size for incoming batches (or lists of episodes) into each Learner worker, thus also determining, how much off-policy’ness would be acceptable. The off-policy’ness is the difference between the numbers of updates a policy has undergone on the Learner vs the EnvRunners. See the
ray.rllib.utils.actor_manager.FaultTolerantActorManagerclass for more details.never_skip_update (bool | None) – Experimental; may change or be removed without a deprecation cycle. By default (False), a Learner skips an
update()call whose train batch is empty (no timesteps for any module; e.g. all sampled episodes were lost to EnvRunner or node failures), and withnum_learners > 1all Learners first agree on that via one small collective perupdate(), so that they skip together and stay in sync. Set to True to turn the skip into an error, raised on every Learner of the group:_should_skip_updateis not consulted, and a train batch that would be skipped by default raises instead – in a group, one in which some module has no timesteps; with a single Learner, which drops modules without timesteps, one in which no module has any. For setups that guarantee every Learner always receives data and want to be told loudly when that guarantee breaks. Independent of this setting, the Learners of a group agree on the number of minibatches to step through whenminibatch_sizeis set (the same collective), so that shards of different sizes cannot make them step a different number of times. Applies toLearner; aDifferentiableLearnercomputes its inner updates without a collective and always skips an empty one.learner_class (Type[Learner] | None) – The
Learnerclass to use for (distributed) updating of the RLModule.learner_connector (Callable[[gymnasium.spaces.Space, gymnasium.spaces.Space], ConnectorV2 | List[ConnectorV2]] | None) – A callable taking an env observation space and an env action space as inputs and returning a learner ConnectorV2 or list of ConnectorV2’s as part of pipeline object.
add_default_connectors_to_learner_pipeline (bool | None) – If True (default), RLlib’s Learners automatically add the default Learner ConnectorV2 pieces to the LearnerPipeline. These automatically perform: a) adding observations from episodes to the train batch, if this has not already been done by a user-provided connector piece b) if RLModule is stateful, add a time rank to the train batch, zero-pad the data, and add the correct state inputs, if this has not already been done by a user-provided connector piece. c) add all other information (actions, rewards, terminateds, etc..) to the train batch, if this has not already been done by a user-provided connector piece. Only if you know exactly what you are doing, you should set this setting to False.
learner_config_dict (Dict[str, Any] | None) – A dict to insert any settings accessible from within the Learner instance. This should only be used in connection with custom Learner subclasses and in case the user doesn’t want to write an extra
AlgorithmConfigsubclass just to add a few settings to the base Algo’s own config class.
- Returns:
This updated AlgorithmConfig object.
- Return type: