APPOConfig#
- class ray.rllib.algorithms.appo.appo.APPOConfig(algo_class=None)[source]#
Bases:
IMPALAConfigDefines a configuration class from which an APPO Algorithm can be built.
from ray.rllib.algorithms.appo import APPOConfig config = ( APPOConfig() .training(lr=0.01, grad_clip=30.0, train_batch_size_per_learner=50) ) config = config.learners(num_learners=1) config = config.env_runners(num_env_runners=1) config = config.environment("CartPole-v1") # Build an Algorithm object from the config and run 1 training iteration. algo = config.build() algo.train() del algo
from ray.rllib.algorithms.appo import APPOConfig from ray import tune config = APPOConfig() # Update the config object. config = config.training(lr=tune.grid_search([0.001,])) # Set the config object's env. config = config.environment(env="CartPole-v1") # Use to_dict() to get the old-style python config dict when running with tune. tune.Tuner( "APPO", run_config=tune.RunConfig( stop={"training_iteration": 1}, verbose=0, ), param_space=config.to_dict(), ).fit()
- training(*, vtrace: bool | None = <ray.rllib.utils.from_config._NotProvided object>, use_gae: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lambda_: float | None = <ray.rllib.utils.from_config._NotProvided object>, clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_kl_loss: bool | None = <ray.rllib.utils.from_config._NotProvided object>, kl_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, kl_target: float | None = <ray.rllib.utils.from_config._NotProvided object>, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, target_worker_clipping: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_circular_buffer: bool | None = <ray.rllib.utils.from_config._NotProvided object>, circular_buffer_num_batches: int | None = <ray.rllib.utils.from_config._NotProvided object>, circular_buffer_iterations_per_batch: int | None = <ray.rllib.utils.from_config._NotProvided object>, simple_queue_size: int | None = <ray.rllib.utils.from_config._NotProvided object>, target_update_frequency: int | None = -1, use_critic: bool | None = -1, **kwargs) Self[source]#
Sets the training related configuration.
- Parameters:
vtrace (bool | None) – Whether to use V-trace weighted advantages. If false, PPO GAE advantages will be used instead.
use_gae (bool | None) – If true, use the Generalized Advantage Estimator (GAE) with a value function, see https://arxiv.org/pdf/1506.02438.pdf. Only applies if vtrace=False.
lambda – GAE (lambda) parameter.
clip_param (float | None) – PPO surrogate slipping parameter.
use_kl_loss (bool | None) – Whether to use the KL-term in the loss function.
kl_coeff (float | None) – Coefficient for weighting the KL-loss term.
kl_target (float | None) – Target term for the KL-term to reach (via adjusting the
kl_coeffautomatically).target_network_update_freq (int | None) – NOTE: This parameter is only applicable on the new API stack. The frequency with which to update the target policy network from the main trained policy network. The metric used is
NUM_ENV_STEPS_TRAINED_LIFETIMEand the unit isn(see [1] 4.1.1), where:n = [circular_buffer_num_batches (N)] * [circular_buffer_iterations_per_batch (K)] * [train batch size]For example, if you settarget_network_update_freq=2, and N=4, K=2, andtrain_batch_size_per_learner=500, then the target net is updated every 2*4*2*500=8000 trained env steps (every 16 batch updates on each learner). The authors in [1] suggests that this setting is robust to a range of choices (try values between 0.125 and 4). This setting also controls how often the kl loss coefficients are tuned: The algorithm waits for at leasttarget_network_update_freqnumber of environment samples to be trained on before updating the target networks and tuning the kl loss coefficients.tau (float | None) – The factor by which to update the target policy network towards the current policy network. Can range between 0 and 1. e.g. updated_param = tau * current_param + (1 - tau) * target_param
target_worker_clipping (float | None) – The maximum value for the target-worker-clipping used for computing the IS ratio, described in [1] IS = min(π(i) / π(target), ρ) * (π / π(i))
use_circular_buffer (bool | None) – Whether to use a circular buffer for storing training batches. If false, a simple Queue will be used. Defaults to True.
circular_buffer_num_batches (int | None) – The number of train batches that fit into the circular buffer. Each such train batch can be sampled for training max.
circular_buffer_iterations_per_batchtimes.circular_buffer_iterations_per_batch (int | None) – The number of times any train batch in the circular buffer can be sampled for training. A batch gets evicted from the buffer either if it’s the oldest batch in the buffer and a new batch is added OR if the batch reaches this max. number of being sampled.
simple_queue_size (int | None) – The size of the simple queue (if
use_circular_bufferis False) for storing training batches.target_update_frequency (int | None) – Deprecated. Use
target_network_update_freqinstead.use_critic (bool | None) – Deprecated. APPO always uses a value function (critic).
**kwargs – Additional config settings, forwarded to the parent
IMPALAConfig.training()method.
- Returns:
This updated AlgorithmConfig object.
- Return type:
- get_default_learner_class()[source]#
Returns the Learner class to use for this algorithm.
Override this method in the sub-class to return the Learner class type given the input framework.
- Returns:
The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).
- get_default_rl_module_spec() RLModuleSpec[source]#
Returns the RLModule spec to use for this algorithm.
Override this method in the subclass to return the RLModule spec, given the input framework.
- Returns:
The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.
- Return type: