IQLConfig#

class ray.rllib.algorithms.iql.iql.IQLConfig(algo_class=None)[source]#

Bases: MARWILConfig

Defines a configuration class from which a new IQL Algorithm can be built

from ray.rllib.algorithms.iql import IQLConfig
# Run this from the ray directory root.
config = IQLConfig().training(actor_lr=0.00001, gamma=0.99)
config = config.offline_data(
    input_="./rllib/offline/tests/data/pendulum/pendulum-v1_enormous")

# Build an Algorithm object from the config and run 1 training iteration.
algo = config.build()
algo.train()
from ray.rllib.algorithms.iql import IQLConfig
from ray import tune
config = IQLConfig()
# Print out some default values.
print(config.beta)
# Update the config object.
config.training(
    lr=tune.grid_search([0.001, 0.0001]), beta=0.75
)
# Set the config object's data path.
# Run this from the ray directory root.
config.offline_data(
    input_="./rllib/offline/tests/data/pendulum/pendulum-v1_enormous"
)
# Set the config object's env, used for evaluation.
config.environment(env="Pendulum-v1")
# Use to_dict() to get the old-style python config dict
# when running with tune.
tune.Tuner(
    "IQL",
    param_space=config.to_dict(),
).fit()
training(*, twin_q: bool | None = <ray.rllib.utils.from_config._NotProvided object>, expectile: float | None = <ray.rllib.utils.from_config._NotProvided object>, actor_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, critic_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, value_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, **kwargs) → IQLConfig[source]#

Sets the training related configuration.

Parameters:
  • twin_q (bool | None) – If a twin-Q architecture should be used (advisable).

  • expectile (float | None) – The expectile to use in expectile regression for the value function. For high expectiles the value function tries to match the upper tail of the Q-value distribution.

  • actor_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the actor network. Actor learning rates greater than critic learning rates work well in experiments.

  • critic_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the Q-network. Critic learning rates greater than value function learning rates work well in experiments.

  • value_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the value function network.

  • target_network_update_freq (int | None) – The number of timesteps in between the target Q-network is fixed. Note, too high values here could harm convergence. The target network is updated via Polyak-averaging.

  • tau (float | None) – The update parameter for Polyak-averaging of the target Q-network. The higher this value the faster the weights move towards the actual Q-network.

  • **kwargs – Additional config settings, forwarded to the parent MARWILConfig.training() method.

Returns:

This updated AlgorithmConfig object.

Return type:

IQLConfig

get_default_learner_class() → Type[Learner] | str[source]#

Returns the Learner class to use for this algorithm.

Override this method in the sub-class to return the Learner class type given the input framework.

Returns:

The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).

Return type:

Type[Learner] | str

get_default_rl_module_spec() → RLModuleSpec | MultiRLModuleSpec[source]#

Returns the RLModule spec to use for this algorithm.

Override this method in the subclass to return the RLModule spec, given the input framework.

Returns:

The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.

Return type:

RLModuleSpec | MultiRLModuleSpec

validate() → None[source]#

Validates all values in this config.