APPOConfig#

class ray.rllib.algorithms.appo.appo.APPOConfig(algo_class=None)[source]#

Bases: IMPALAConfig

Defines a configuration class from which an APPO Algorithm can be built.

from ray.rllib.algorithms.appo import APPOConfig
config = (
    APPOConfig()
    .training(lr=0.01, grad_clip=30.0, train_batch_size_per_learner=50)
)
config = config.learners(num_learners=1)
config = config.env_runners(num_env_runners=1)
config = config.environment("CartPole-v1")

# Build an Algorithm object from the config and run 1 training iteration.
algo = config.build()
algo.train()
del algo
from ray.rllib.algorithms.appo import APPOConfig
from ray import tune

config = APPOConfig()
# Update the config object.
config = config.training(lr=tune.grid_search([0.001,]))
# Set the config object's env.
config = config.environment(env="CartPole-v1")
# Use to_dict() to get the old-style python config dict when running with tune.
tune.Tuner(
    "APPO",
    run_config=tune.RunConfig(
        stop={"training_iteration": 1},
        verbose=0,
    ),
    param_space=config.to_dict(),

).fit()
training(*, vtrace: bool | None = <ray.rllib.utils.from_config._NotProvided object>, use_gae: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lambda_: float | None = <ray.rllib.utils.from_config._NotProvided object>, clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_kl_loss: bool | None = <ray.rllib.utils.from_config._NotProvided object>, kl_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, kl_target: float | None = <ray.rllib.utils.from_config._NotProvided object>, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, target_worker_clipping: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_circular_buffer: bool | None = <ray.rllib.utils.from_config._NotProvided object>, circular_buffer_num_batches: int | None = <ray.rllib.utils.from_config._NotProvided object>, circular_buffer_iterations_per_batch: int | None = <ray.rllib.utils.from_config._NotProvided object>, simple_queue_size: int | None = <ray.rllib.utils.from_config._NotProvided object>, target_update_frequency: int | None = -1, use_critic: bool | None = -1, **kwargs) → Self[source]#

Sets the training related configuration.

Parameters:
  • vtrace (bool | None) – Whether to use V-trace weighted advantages. If false, PPO GAE advantages will be used instead.

  • use_gae (bool | None) – If true, use the Generalized Advantage Estimator (GAE) with a value function, see https://arxiv.org/pdf/1506.02438.pdf. Only applies if vtrace=False.

  • lambda – GAE (lambda) parameter.

  • clip_param (float | None) – PPO surrogate slipping parameter.

  • use_kl_loss (bool | None) – Whether to use the KL-term in the loss function.

  • kl_coeff (float | None) – Coefficient for weighting the KL-loss term.

  • kl_target (float | None) – Target term for the KL-term to reach (via adjusting the kl_coeff automatically).

  • target_network_update_freq (int | None) – NOTE: This parameter is only applicable on the new API stack. The frequency with which to update the target policy network from the main trained policy network. The metric used is NUM_ENV_STEPS_TRAINED_LIFETIME and the unit is n (see [1] 4.1.1), where: n = [circular_buffer_num_batches (N)] * [circular_buffer_iterations_per_batch (K)] * [train batch size] For example, if you set target_network_update_freq=2, and N=4, K=2, and train_batch_size_per_learner=500, then the target net is updated every 2*4*2*500=8000 trained env steps (every 16 batch updates on each learner). The authors in [1] suggests that this setting is robust to a range of choices (try values between 0.125 and 4). This setting also controls how often the kl loss coefficients are tuned: The algorithm waits for at least target_network_update_freq number of environment samples to be trained on before updating the target networks and tuning the kl loss coefficients.

  • tau (float | None) – The factor by which to update the target policy network towards the current policy network. Can range between 0 and 1. e.g. updated_param = tau * current_param + (1 - tau) * target_param

  • target_worker_clipping (float | None) – The maximum value for the target-worker-clipping used for computing the IS ratio, described in [1] IS = min(π(i) / π(target), ρ) * (π / π(i))

  • use_circular_buffer (bool | None) – Whether to use a circular buffer for storing training batches. If false, a simple Queue will be used. Defaults to True.

  • circular_buffer_num_batches (int | None) – The number of train batches that fit into the circular buffer. Each such train batch can be sampled for training max. circular_buffer_iterations_per_batch times.

  • circular_buffer_iterations_per_batch (int | None) – The number of times any train batch in the circular buffer can be sampled for training. A batch gets evicted from the buffer either if it’s the oldest batch in the buffer and a new batch is added OR if the batch reaches this max. number of being sampled.

  • simple_queue_size (int | None) – The size of the simple queue (if use_circular_buffer is False) for storing training batches.

  • target_update_frequency (int | None) – Deprecated. Use target_network_update_freq instead.

  • use_critic (bool | None) – Deprecated. APPO always uses a value function (critic).

  • **kwargs – Additional config settings, forwarded to the parent IMPALAConfig.training() method.

Returns:

This updated AlgorithmConfig object.

Return type:

Self

validate() → None[source]#

Validates all values in this config.

get_default_learner_class()[source]#

Returns the Learner class to use for this algorithm.

Override this method in the sub-class to return the Learner class type given the input framework.

Returns:

The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).

get_default_rl_module_spec() → RLModuleSpec[source]#

Returns the RLModule spec to use for this algorithm.

Override this method in the subclass to return the RLModule spec, given the input framework.

Returns:

The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.

Return type:

RLModuleSpec