PPOConfig#
- class ray.rllib.algorithms.ppo.ppo.PPOConfig(algo_class=None)[source]#
Bases:
AlgorithmConfigDefines a configuration class from which a PPO Algorithm can be built.
from ray.rllib.algorithms.ppo import PPOConfig config = PPOConfig() config.environment("CartPole-v1") config.env_runners(num_env_runners=1) config.training( gamma=0.9, lr=0.01, kl_coeff=0.3, train_batch_size_per_learner=256 ) # Build a Algorithm object from the config and run 1 training iteration. algo = config.build() algo.train()
from ray.rllib.algorithms.ppo import PPOConfig from ray import tune config = ( PPOConfig() # Set the config object's env. .environment(env="CartPole-v1") # Update the config object's training parameters. .training( lr=0.001, clip_param=0.2 ) ) tune.Tuner( "PPO", run_config=tune.RunConfig(stop={"training_iteration": 1}), param_space=config, ).fit()
- get_default_rl_module_spec() RLModuleSpec[source]#
Returns the RLModule spec to use for this algorithm.
Override this method in the subclass to return the RLModule spec, given the input framework.
- Returns:
The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.
- Return type:
- get_default_learner_class() Type[Learner] | str[source]#
Returns the Learner class to use for this algorithm.
Override this method in the sub-class to return the Learner class type given the input framework.
- training(*, use_critic: bool | None = <ray.rllib.utils.from_config._NotProvided object>, use_gae: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lambda_: float | None = <ray.rllib.utils.from_config._NotProvided object>, use_kl_loss: bool | None = <ray.rllib.utils.from_config._NotProvided object>, kl_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, kl_target: float | None = <ray.rllib.utils.from_config._NotProvided object>, vf_loss_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff: float | None = <ray.rllib.utils.from_config._NotProvided object>, entropy_coeff_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, vf_clip_param: float | None = <ray.rllib.utils.from_config._NotProvided object>, grad_clip: float | None = <ray.rllib.utils.from_config._NotProvided object>, lr_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, vf_share_layers: bool | None = -1, **kwargs) Self[source]#
Sets the training related configuration.
- Parameters:
use_critic (bool | None) – Should use a critic as a baseline (otherwise don’t use value baseline; required for using GAE).
use_gae (bool | None) – If true, use the Generalized Advantage Estimator (GAE) with a value function, see https://arxiv.org/pdf/1506.02438.pdf.
lambda – The lambda parameter for General Advantage Estimation (GAE). Defines the exponential weight used between actually measured rewards vs value function estimates over multiple time steps. Specifically,
lambda_balances short-term, low-variance estimates against long-term, high-variance returns. Alambda_of 0.0 makes the GAE rely only on immediate rewards (and vf predictions from there on, reducing variance, but increasing bias), while alambda_of 1.0 only incorporates vf predictions at the truncation points of the given episodes or episode chunks (reducing bias but increasing variance).use_kl_loss (bool | None) – Whether to use the KL-term in the loss function.
kl_coeff (float | None) – Initial coefficient for KL divergence.
kl_target (float | None) – Target value for KL divergence.
vf_loss_coeff (float | None) – Coefficient of the value function loss. IMPORTANT: you must tune this if you set vf_share_layers=True inside your model’s config.
entropy_coeff (float | None) – The entropy coefficient (float) or entropy coefficient schedule in the format of [[timestep, coeff-value], [timestep, coeff-value], …] In case of a schedule, intermediary timesteps will be assigned to linearly interpolated coefficient values. A schedule config’s first entry must start with timestep 0, i.e.: [[0, initial_value], […]].
entropy_coeff_schedule (List[List[int | float]] | None) – Decay schedule for the entropy regularizer, in the format of [[timestep, coeff-value], [timestep, coeff-value], …]. @OldAPIStack
clip_param (float | None) – The PPO clip parameter.
vf_clip_param (float | None) – Clip param for the value function. Note that this is sensitive to the scale of the rewards. If your expected V is large, increase this.
grad_clip (float | None) – If specified, clip the global norm of gradients by this amount.
lr_schedule (List[List[int | float]] | None) – Learning rate schedule, in the format of [[timestep, lr-value], [timestep, lr-value], …]. Intermediary timesteps are assigned to interpolated learning rate values. A schedule should normally start from timestep 0. @OldAPIStack
vf_share_layers (bool | None) – Deprecated and ignored. Use
config.rl_module(model_config={"vf_share_layers": ...})instead.**kwargs – Additional config settings, forwarded to the parent
AlgorithmConfig.training()method.
- Returns:
This updated AlgorithmConfig object.
- Return type: