IQLConfig#
- class ray.rllib.algorithms.iql.iql.IQLConfig(algo_class=None)[source]#
Bases:
MARWILConfigDefines a configuration class from which a new IQL Algorithm can be built
from ray.rllib.algorithms.iql import IQLConfig # Run this from the ray directory root. config = IQLConfig().training(actor_lr=0.00001, gamma=0.99) config = config.offline_data( input_="./rllib/offline/tests/data/pendulum/pendulum-v1_enormous") # Build an Algorithm object from the config and run 1 training iteration. algo = config.build() algo.train()
from ray.rllib.algorithms.iql import IQLConfig from ray import tune config = IQLConfig() # Print out some default values. print(config.beta) # Update the config object. config.training( lr=tune.grid_search([0.001, 0.0001]), beta=0.75 ) # Set the config object's data path. # Run this from the ray directory root. config.offline_data( input_="./rllib/offline/tests/data/pendulum/pendulum-v1_enormous" ) # Set the config object's env, used for evaluation. config.environment(env="Pendulum-v1") # Use to_dict() to get the old-style python config dict # when running with tune. tune.Tuner( "IQL", param_space=config.to_dict(), ).fit()
- training(*, twin_q: bool | None = <ray.rllib.utils.from_config._NotProvided object>, expectile: float | None = <ray.rllib.utils.from_config._NotProvided object>, actor_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, critic_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, value_lr: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, **kwargs) IQLConfig[source]#
Sets the training related configuration.
- Parameters:
twin_q (bool | None) – If a twin-Q architecture should be used (advisable).
expectile (float | None) – The expectile to use in expectile regression for the value function. For high expectiles the value function tries to match the upper tail of the Q-value distribution.
actor_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the actor network. Actor learning rates greater than critic learning rates work well in experiments.
critic_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the Q-network. Critic learning rates greater than value function learning rates work well in experiments.
value_lr (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – The learning rate for the value function network.
target_network_update_freq (int | None) – The number of timesteps in between the target Q-network is fixed. Note, too high values here could harm convergence. The target network is updated via Polyak-averaging.
tau (float | None) – The update parameter for Polyak-averaging of the target Q-network. The higher this value the faster the weights move towards the actual Q-network.
**kwargs – Additional config settings, forwarded to the parent
MARWILConfig.training()method.
- Returns:
This updated
AlgorithmConfigobject.- Return type:
- get_default_learner_class() Type[Learner] | str[source]#
Returns the Learner class to use for this algorithm.
Override this method in the sub-class to return the Learner class type given the input framework.
- get_default_rl_module_spec() RLModuleSpec | MultiRLModuleSpec[source]#
Returns the RLModule spec to use for this algorithm.
Override this method in the subclass to return the RLModule spec, given the input framework.
- Returns:
The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.
- Return type: