PPOConfig#

class ray.rllib.algorithms.ppo.ppo.PPOConfig(algo_class=None)[source]#

Bases: AlgorithmConfig

Defines a configuration class from which a PPO Algorithm can be built.

from ray.rllib.algorithms.ppo import PPOConfig

config = PPOConfig()
config.environment("CartPole-v1")
config.env_runners(num_env_runners=1)
config.training(
    gamma=0.9, lr=0.01, kl_coeff=0.3, train_batch_size_per_learner=256
)

# Build a Algorithm object from the config and run 1 training iteration.
algo = config.build()
algo.train()
from ray.rllib.algorithms.ppo import PPOConfig
from ray import tune

config = (
    PPOConfig()
    # Set the config object's env.
    .environment(env="CartPole-v1")
    # Update the config object's training parameters.
    .training(
        lr=0.001, clip_param=0.2
    )
)

tune.Tuner(
    "PPO",
    run_config=tune.RunConfig(stop={"training_iteration": 1}),
    param_space=config,
).fit()
get_default_rl_module_spec() → RLModuleSpec[source]#

Returns the RLModule spec to use for this algorithm.

Override this method in the subclass to return the RLModule spec, given the input framework.

Returns:

The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.

Return type:

RLModuleSpec

get_default_learner_class() → Type[Learner] | str[source]#

Returns the Learner class to use for this algorithm.

Override this method in the sub-class to return the Learner class type given the input framework.

Returns:

The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).

Return type:

Type[Learner] | str

training(*, use_critic: bool | None = <ray.rllib.utils.from_config._NotProvided object>, use_gae: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lambda_: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_kl_loss: bool | None = <ray.rllib.utils.from_config._NotProvided object>, kl_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, kl_target: float | None = <ray.rllib.utils.from_config._NotProvided object>, vf_loss_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, vf_clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, grad_clip: float | None = <ray.rllib.utils.from_config._NotProvided object>, lr_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, vf_share_layers: bool | None = -1, **kwargs) → Self[source]#

Sets the training related configuration.

Parameters:
  • use_critic (bool | None) – Should use a critic as a baseline (otherwise don’t use value baseline; required for using GAE).

  • use_gae (bool | None) – If true, use the Generalized Advantage Estimator (GAE) with a value function, see https://arxiv.org/pdf/1506.02438.pdf.

  • lambda – The lambda parameter for General Advantage Estimation (GAE). Defines the exponential weight used between actually measured rewards vs value function estimates over multiple time steps. Specifically, lambda_ balances short-term, low-variance estimates against long-term, high-variance returns. A lambda_ of 0.0 makes the GAE rely only on immediate rewards (and vf predictions from there on, reducing variance, but increasing bias), while a lambda_ of 1.0 only incorporates vf predictions at the truncation points of the given episodes or episode chunks (reducing bias but increasing variance).

  • use_kl_loss (bool | None) – Whether to use the KL-term in the loss function.

  • kl_coeff (float | None) – Initial coefficient for KL divergence.

  • kl_target (float | None) – Target value for KL divergence.

  • vf_loss_coeff (float | None) – Coefficient of the value function loss. IMPORTANT: you must tune this if you set vf_share_layers=True inside your model’s config.

  • entropy_coeff (float | None) – The entropy coefficient (float) or entropy coefficient schedule in the format of [[timestep, coeff-value], [timestep, coeff-value], …] In case of a schedule, intermediary timesteps will be assigned to linearly interpolated coefficient values. A schedule config’s first entry must start with timestep 0, i.e.: [[0, initial_value], […]].

  • entropy_coeff_schedule (List[List[int | float]] | None) – Decay schedule for the entropy regularizer, in the format of [[timestep, coeff-value], [timestep, coeff-value], …]. @OldAPIStack

  • clip_param (float | None) – The PPO clip parameter.

  • vf_clip_param (float | None) – Clip param for the value function. Note that this is sensitive to the scale of the rewards. If your expected V is large, increase this.

  • grad_clip (float | None) – If specified, clip the global norm of gradients by this amount.

  • lr_schedule (List[List[int | float]] | None) – Learning rate schedule, in the format of [[timestep, lr-value], [timestep, lr-value], …]. Intermediary timesteps are assigned to interpolated learning rate values. A schedule should normally start from timestep 0. @OldAPIStack

  • vf_share_layers (bool | None) – Deprecated and ignored. Use config.rl_module(model_config={"vf_share_layers": ...}) instead.

  • **kwargs – Additional config settings, forwarded to the parent AlgorithmConfig.training() method.

Returns:

This updated AlgorithmConfig object.

Return type:

Self

validate() → None[source]#

Validates all values in this config.