DQNConfig#
- class ray.rllib.algorithms.dqn.dqn.DQNConfig(algo_class=None)[source]#
Bases:
AlgorithmConfigDefines a configuration class from which a DQN Algorithm can be built.
from ray.rllib.algorithms.dqn.dqn import DQNConfig config = ( DQNConfig() .environment("CartPole-v1") .training(replay_buffer_config={ "type": "PrioritizedEpisodeReplayBuffer", "capacity": 60000, "alpha": 0.5, "beta": 0.5, }) .env_runners(num_env_runners=1) ) algo = config.build() algo.train() algo.stop()
from ray.rllib.algorithms.dqn.dqn import DQNConfig from ray import tune config = ( DQNConfig() .environment("CartPole-v1") .training( num_atoms=tune.grid_search([1,]) ) ) tune.Tuner( "DQN", run_config=tune.RunConfig(stop={"training_iteration":1}), param_space=config, ).fit()
- training(*, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, replay_buffer_config: dict | None = <ray.rllib.utils.from_config._NotProvided object>, store_buffer_in_checkpoints: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lr_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, epsilon: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, adam_epsilon: float | None = <ray.rllib.utils.from_config._NotProvided object>, grad_clip: int | None = <ray.rllib.utils.from_config._NotProvided object>, num_steps_sampled_before_learning_starts: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, num_atoms: int | None = <ray.rllib.utils.from_config._NotProvided object>, v_min: float | None = <ray.rllib.utils.from_config._NotProvided object>, v_max: float | None = <ray.rllib.utils.from_config._NotProvided object>, noisy: bool | None = <ray.rllib.utils.from_config._NotProvided object>, sigma0: float | None = <ray.rllib.utils.from_config._NotProvided object>, dueling: bool | None = <ray.rllib.utils.from_config._NotProvided object>, hiddens: int | None = <ray.rllib.utils.from_config._NotProvided object>, double_q: bool | None = <ray.rllib.utils.from_config._NotProvided object>, n_step: int | ~typing.Tuple[int, int] | None = <ray.rllib.utils.from_config._NotProvided object>, before_learn_on_batch: ~typing.Callable[[~typing.Type[~ray.rllib.policy.sample_batch.MultiAgentBatch], ~typing.List[~typing.Type[~ray.rllib.policy.policy.Policy]], ~typing.Type[int]], ~typing.Type[~ray.rllib.policy.sample_batch.MultiAgentBatch]] = <ray.rllib.utils.from_config._NotProvided object>, training_intensity: float | None = <ray.rllib.utils.from_config._NotProvided object>, td_error_loss_fn: str | None = <ray.rllib.utils.from_config._NotProvided object>, categorical_distribution_temperature: float | None = <ray.rllib.utils.from_config._NotProvided object>, burn_in_len: int | None = <ray.rllib.utils.from_config._NotProvided object>, **kwargs) Self[source]#
Sets the training related configuration.
- Parameters:
target_network_update_freq (int | None) – Update the target network every
target_network_update_freqsample steps.replay_buffer_config (dict | None) – Replay buffer config. Examples: { “_enable_replay_buffer_api”: True, “type”: “MultiAgentReplayBuffer”, “capacity”: 50000, “replay_sequence_length”: 1, } - OR - { “_enable_replay_buffer_api”: True, “type”: “MultiAgentPrioritizedReplayBuffer”, “capacity”: 50000, “prioritized_replay_alpha”: 0.6, “prioritized_replay_beta”: 0.4, “prioritized_replay_eps”: 1e-6, “replay_sequence_length”: 1, } - Where - prioritized_replay_alpha: Alpha parameter controls the degree of prioritization in the buffer. In other words, when a buffer sample has a higher temporal-difference error, with how much more probability should it drawn to use to update the parametrized Q-network. 0.0 corresponds to uniform probability. Setting much above 1.0 may quickly result as the sampling distribution could become heavily “pointy” with low entropy. prioritized_replay_beta: Beta parameter controls the degree of importance sampling which suppresses the influence of gradient updates from samples that have higher probability of being sampled via alpha parameter and the temporal-difference error. prioritized_replay_eps: Epsilon parameter sets the baseline probability for sampling so that when the temporal-difference error of a sample is zero, there is still a chance of drawing the sample.
store_buffer_in_checkpoints (bool | None) – Set this to True, if you want the contents of your buffer(s) to be stored in any saved checkpoints as well. Warnings will be created if: - This is True AND restoring from a checkpoint that contains no buffer data. - This is False AND restoring from a checkpoint that does contain buffer data.
lr_schedule (List[List[int | float]] | None) – Deprecated (old API stack only). Learning rate schedule in the format of [[timestep, lr-value], [timestep, lr-value], …]. Use
lrwith a schedule on the new API stack instead.epsilon (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – Epsilon exploration schedule. In the format of [[timestep, value], [timestep, value], …]. A schedule must start from timestep 0.
adam_epsilon (float | None) – Adam optimizer’s epsilon hyper parameter.
grad_clip (int | None) – If not None, clip gradients during optimization at this value.
num_steps_sampled_before_learning_starts (int | None) – Number of timesteps to collect from rollout workers before we start sampling from replay buffers for learning. Whether we count this in agent steps or environment steps depends on config.multi_agent(count_steps_by=..).
tau (float | None) – Update the target by au * policy + (1- au) * target_policy.
num_atoms (int | None) – Number of atoms for representing the distribution of return. When this is greater than 1, distributional Q-learning is used.
v_min (float | None) – Minimum value estimation
v_max (float | None) – Maximum value estimation
noisy (bool | None) – Whether to use noisy network to aid exploration. This adds parametric noise to the model weights.
sigma0 (float | None) – Control the initial parameter noise for noisy nets.
dueling (bool | None) – Whether to use dueling DQN.
hiddens (int | None) – Dense-layer setup for each the advantage branch and the value branch in a dueling architecture.
double_q (bool | None) – Whether to use double DQN.
n_step (int | Tuple[int, int] | None) – N-step target updates. If >1, sars’ tuples in trajectories will be postprocessed to become sa[discounted sum of R][s t+n] tuples. An integer will be interpreted as a fixed n-step value. If a tuple of 2 ints is provided here, the n-step value will be drawn for each sample(!) in the train batch from a uniform distribution over the closed interval defined by
[n_step[0], n_step[1]].before_learn_on_batch (Callable[[Type[MultiAgentBatch], List[Type[Policy]], Type[int]], Type[MultiAgentBatch]]) – Callback to run before learning on a multi-agent batch of experiences.
training_intensity (float | None) – The intensity with which to update the model (vs collecting samples from the env). If None, uses “natural” values of:
train_batch_size/ (rollout_fragment_lengthxnum_env_runnersxnum_envs_per_env_runner). If not None, will make sure that the ratio between timesteps inserted into and sampled from the buffer matches the given values. Example: training_intensity=1000.0 train_batch_size=250 rollout_fragment_length=1 num_env_runners=1 (or 0) num_envs_per_env_runner=1 -> natural value = 250 / 1 = 250.0 -> will make sure that replay+train op will be executed 4x asoften as rollout+insert op (4 * 250 = 1000). See: rllib/algorithms/dqn/dqn.py::calculate_rr_weights for further details.td_error_loss_fn (str | None) – “huber” or “mse”. loss function for calculating TD error when num_atoms is 1. Note that if num_atoms is > 1, this parameter is simply ignored, and softmax cross entropy loss will be used.
categorical_distribution_temperature (float | None) – Set the temperature parameter used by Categorical action distribution. A valid temperature is in the range of [0, 1]. Note that this mostly affects evaluation since TD error uses argmax for return calculation.
burn_in_len (int | None) – The burn-in period for a stateful RLModule. It allows the Learner to utilize the initial
burn_in_lensteps in a replay sequence solely for unrolling the network and establishing a typical starting state. The network is then updated on the remaining steps of the sequence. This process helps mitigate issues stemming from a poor initial state - zero or an outdated recorded state. Consider setting this parameter to a positive integer if your stateful RLModule faces convergence challenges or exhibits signs of catastrophic forgetting.**kwargs – Additional config settings, forwarded to the parent
AlgorithmConfig.training()method.
- Returns:
This updated AlgorithmConfig object.
- Return type:
- get_rollout_fragment_length(worker_index: int = 0) int[source]#
Automatically infers a proper rollout_fragment_length setting if “auto”.
Uses the simple formula:
rollout_fragment_length=total_train_batch_size/ (num_envs_per_env_runner*num_env_runners)If result is a fraction AND
worker_indexis provided, makes those workers add additional timesteps, such that the overall batch size (across the workers) adds up to exactly thetotal_train_batch_size. Fractions < 1 calculated this way are rounded up to a rollout_fragment_length of 1.- Parameters:
worker_index (int) – The 1-based index of the EnvRunner asking for its fragment length (0 for the local EnvRunner). Used to distribute a fractional remainder across the first n EnvRunners.
- Returns:
The user-provided
rollout_fragment_lengthor a computed one (if user provided value is “auto”), making suretotal_train_batch_sizeis reached exactly in each iteration.- Return type:
- get_default_rl_module_spec() RLModuleSpec | MultiRLModuleSpec[source]#
Returns the RLModule spec to use for this algorithm.
Override this method in the subclass to return the RLModule spec, given the input framework.
- Returns:
The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.
- Return type: