IMPALAConfig#
- class ray.rllib.algorithms.impala.impala.IMPALAConfig(algo_class=None)[source]#
Bases:
AlgorithmConfigDefines a configuration class from which an Impala can be built.
from ray.rllib.algorithms.impala import IMPALAConfig config = ( IMPALAConfig() .environment("CartPole-v1") .env_runners(num_env_runners=1) .training(lr=0.0003, train_batch_size_per_learner=512) .learners(num_learners=1) ) # Build a Algorithm object from the config and run 1 training iteration. algo = config.build() algo.train() del algo
from ray.rllib.algorithms.impala import IMPALAConfig from ray import tune config = ( IMPALAConfig() .environment("CartPole-v1") .env_runners(num_env_runners=1) .training(lr=tune.grid_search([0.0001, 0.0002]), grad_clip=20.0) .learners(num_learners=1) ) # Run with tune. tune.Tuner( "IMPALA", param_space=config, run_config=tune.RunConfig(stop={"training_iteration": 1}), ).fit()
- training(*, vtrace: bool | None = <ray.rllib.utils.from_config._NotProvided object>, vtrace_clip_rho_threshold: float | None = <ray.rllib.utils.from_config._NotProvided object>, vtrace_clip_pg_rho_threshold: float | None = <ray.rllib.utils.from_config._NotProvided object>, num_gpu_loader_threads: int | None = <ray.rllib.utils.from_config._NotProvided object>, num_multi_gpu_tower_stacks: int | None = <ray.rllib.utils.from_config._NotProvided object>, minibatch_buffer_size: int | None = <ray.rllib.utils.from_config._NotProvided object>, replay_proportion: float | None = <ray.rllib.utils.from_config._NotProvided object>, replay_buffer_num_slots: int | None = <ray.rllib.utils.from_config._NotProvided object>, learner_queue_size: int | None = <ray.rllib.utils.from_config._NotProvided object>, learner_queue_timeout: float | None = <ray.rllib.utils.from_config._NotProvided object>, timeout_s_sampler_manager: float | None = <ray.rllib.utils.from_config._NotProvided object>, timeout_s_aggregator_manager: float | None = <ray.rllib.utils.from_config._NotProvided object>, broadcast_interval: int | None = <ray.rllib.utils.from_config._NotProvided object>, grad_clip: float | None = <ray.rllib.utils.from_config._NotProvided object>, opt_type: str | None = <ray.rllib.utils.from_config._NotProvided object>, lr_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, decay: float | None = <ray.rllib.utils.from_config._NotProvided object>, momentum: float | None = <ray.rllib.utils.from_config._NotProvided object>, epsilon: float | None = <ray.rllib.utils.from_config._NotProvided object>, vf_loss_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, _separate_vf_optimizer: bool | None = <ray.rllib.utils.from_config._NotProvided object>, _lr_vf: float | None = <ray.rllib.utils.from_config._NotProvided object>, num_aggregation_workers: int | None = -1, max_requests_in_flight_per_aggregator_worker: int | None = -1, **kwargs) Self[source]#
Sets the training related configuration.
- Parameters:
vtrace (bool | None) – V-trace params (see vtrace_tf/torch.py).
vtrace_clip_rho_threshold (float | None)
vtrace_clip_pg_rho_threshold (float | None)
num_gpu_loader_threads (int | None) – The number of GPU-loader threads (per Learner worker), used to load incoming (CPU) batches to the GPU, if applicable. The incoming batches are produced by each Learner’s LearnerConnector pipeline. After loading the batches on the GPU, the threads place them on yet another queue for the Learner thread (only one per Learner worker) to pick up and perform
forward_train/losscomputations.num_multi_gpu_tower_stacks (int | None) – For each stack of multi-GPU towers, how many slots should we reserve for parallel data loading? Set this to >1 to load data into GPUs in parallel. This will increase GPU memory usage proportionally with the number of stacks. Example: 2 GPUs and
num_multi_gpu_tower_stacks=3: - One tower stack consists of 2 GPUs, each with a copy of the model/graph. - Each of the stacks will create 3 slots for batch data on each of its GPUs, increasing memory requirements on each GPU by 3x. - This enables us to preload data into these stacks while another stack is performing gradient calculations.minibatch_buffer_size (int | None) – How many train batches should be retained for minibatching. This conf only has an effect if
num_epochs > 1.replay_proportion (float | None) – Set >0 to enable experience replay. Saved samples will be replayed with a p:1 proportion to new data samples.
replay_buffer_num_slots (int | None) – Number of sample batches to store for replay. The number of transitions saved total will be (replay_buffer_num_slots * rollout_fragment_length).
learner_queue_size (int | None) – Max queue size for train batches feeding into the learner.
learner_queue_timeout (float | None) – Wait for train batches to be available in minibatch buffer queue this many seconds. This may need to be increased e.g. when training with a slow environment.
timeout_s_sampler_manager (float | None) – The timeout for waiting for sampling results for workers – typically if this is too low, the manager won’t be able to retrieve ready sampling results.
timeout_s_aggregator_manager (float | None) – The timeout for waiting for replay worker results – typically if this is too low, the manager won’t be able to retrieve ready replay requests.
broadcast_interval (int | None) – Number of training step calls before weights are broadcasted to rollout workers that are sampled during any iteration.
grad_clip (float | None) – If specified, clip the global norm of gradients by this amount.
opt_type (str | None) – Either “adam” or “rmsprop”.
lr_schedule (List[List[int | float]] | None) – Learning rate schedule. In the format of [[timestep, lr-value], [timestep, lr-value], …] Intermediary timesteps will be assigned to interpolated learning rate values. A schedule should normally start from timestep 0.
decay (float | None) – Decay setting for the RMSProp optimizer, in case
opt_type=rmsprop.momentum (float | None) – Momentum setting for the RMSProp optimizer, in case
opt_type=rmsprop.epsilon (float | None) – Epsilon setting for the RMSProp optimizer, in case
opt_type=rmsprop.vf_loss_coeff (float | None) – Coefficient for the value function term in the loss function.
entropy_coeff (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – Coefficient for the entropy regularizer term in the loss function.
entropy_coeff_schedule (List[List[int | float]] | None) – Decay schedule for the entropy regularizer.
_separate_vf_optimizer (bool | None) – Set this to true to have two separate optimizers optimize the policy-and value networks. Only supported for some algorithms (APPO, IMPALA) on the old API stack.
_lr_vf (float | None) – If _separate_vf_optimizer is True, define separate learning rate for the value network.
num_aggregation_workers (int | None) – Deprecated. Use
config.learners(num_aggregator_actors_per_learner=..)on the new API stack instead.max_requests_in_flight_per_aggregator_worker (int | None) – Deprecated. Use
config.learners(max_requests_in_flight_per_aggregator_actor=..)on the new API stack instead.**kwargs – Additional config settings, forwarded to the parent
AlgorithmConfig.training()method.
- Returns:
This updated AlgorithmConfig object.
- Return type:
- debugging(*, _env_runners_only: bool | None = <ray.rllib.utils.from_config._NotProvided object>, _skip_learners: bool | None = <ray.rllib.utils.from_config._NotProvided object>, **kwargs) Self[source]#
Sets the debugging related configuration.
- Parameters:
_env_runners_only (bool | None) – If True, only run (remote) EnvRunner requests, discard their episode/training data, but log their metrics results. Aggregator- and Learner actors won’t be used.
_skip_learners (bool | None) – If True, no
updaterequests are sent to the LearnerGroup and Learner actors. Only EnvRunners and aggregator actors (if applicable) are used.**kwargs – Additional config settings, forwarded to the parent
AlgorithmConfig.debugging()method.
- Returns:
This updated AlgorithmConfig object.
- Return type:
- property replay_ratio: float#
Returns replay ratio (between 0.0 and 1.0) based off self.replay_proportion.
Formula: ratio = 1 / proportion
- get_default_learner_class()[source]#
Returns the Learner class to use for this algorithm.
Override this method in the sub-class to return the Learner class type given the input framework.
- Returns:
The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).
- get_default_rl_module_spec() RLModuleSpec[source]#
Returns the RLModule spec to use for this algorithm.
Override this method in the subclass to return the RLModule spec, given the input framework.
- Returns:
The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.
- Return type: