DQNConfig#

class ray.rllib.algorithms.dqn.dqn.DQNConfig(algo_class=None)[source]#

Bases: AlgorithmConfig

Defines a configuration class from which a DQN Algorithm can be built.

from ray.rllib.algorithms.dqn.dqn import DQNConfig

config = (
    DQNConfig()
    .environment("CartPole-v1")
    .training(replay_buffer_config={
        "type": "PrioritizedEpisodeReplayBuffer",
        "capacity": 60000,
        "alpha": 0.5,
        "beta": 0.5,
    })
    .env_runners(num_env_runners=1)
)
algo = config.build()
algo.train()
algo.stop()
from ray.rllib.algorithms.dqn.dqn import DQNConfig
from ray import tune

config = (
    DQNConfig()
    .environment("CartPole-v1")
    .training(
        num_atoms=tune.grid_search([1,])
    )
)
tune.Tuner(
    "DQN",
    run_config=tune.RunConfig(stop={"training_iteration":1}),
    param_space=config,
).fit()
training(*, target_network_update_freq: int | None = <ray.rllib.utils.from_config._NotProvided object>, replay_buffer_config: dict | None = <ray.rllib.utils.from_config._NotProvided object>, store_buffer_in_checkpoints: bool | None = <ray.rllib.utils.from_config._NotProvided object>, lr_schedule: ~typing.List[~typing.List[int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, epsilon: float | ~typing.List[~typing.List[int | float]] | ~typing.List[~typing.Tuple[int, int | float]] | None = <ray.rllib.utils.from_config._NotProvided object>, adam_epsilon: float | None = <ray.rllib.utils.from_config._NotProvided object>, grad_clip: int | None = <ray.rllib.utils.from_config._NotProvided object>, num_steps_sampled_before_learning_starts: int | None = <ray.rllib.utils.from_config._NotProvided object>, tau: float | None = <ray.rllib.utils.from_config._NotProvided object>, num_atoms: int | None = <ray.rllib.utils.from_config._NotProvided object>, v_min: float | None = <ray.rllib.utils.from_config._NotProvided object>, v_max: float | None = <ray.rllib.utils.from_config._NotProvided object>, noisy: bool | None = <ray.rllib.utils.from_config._NotProvided object>, sigma0: float | None = <ray.rllib.utils.from_config._NotProvided object>, dueling: bool | None = <ray.rllib.utils.from_config._NotProvided object>, hiddens: int | None = <ray.rllib.utils.from_config._NotProvided object>, double_q: bool | None = <ray.rllib.utils.from_config._NotProvided object>, n_step: int | ~typing.Tuple[int, int] | None = <ray.rllib.utils.from_config._NotProvided object>, before_learn_on_batch: ~typing.Callable[[~typing.Type[~ray.rllib.policy.sample_batch.MultiAgentBatch], ~typing.List[~typing.Type[~ray.rllib.policy.policy.Policy]], ~typing.Type[int]], ~typing.Type[~ray.rllib.policy.sample_batch.MultiAgentBatch]] = <ray.rllib.utils.from_config._NotProvided object>, training_intensity: float | None = <ray.rllib.utils.from_config._NotProvided object>, td_error_loss_fn: str | None = <ray.rllib.utils.from_config._NotProvided object>, categorical_distribution_temperature: float | None = <ray.rllib.utils.from_config._NotProvided object>, burn_in_len: int | None = <ray.rllib.utils.from_config._NotProvided object>, **kwargs) → Self[source]#

Sets the training related configuration.

Parameters:
  • target_network_update_freq (int | None) – Update the target network every target_network_update_freq sample steps.

  • replay_buffer_config (dict | None) – Replay buffer config. Examples: { “_enable_replay_buffer_api”: True, “type”: “MultiAgentReplayBuffer”, “capacity”: 50000, “replay_sequence_length”: 1, } - OR - { “_enable_replay_buffer_api”: True, “type”: “MultiAgentPrioritizedReplayBuffer”, “capacity”: 50000, “prioritized_replay_alpha”: 0.6, “prioritized_replay_beta”: 0.4, “prioritized_replay_eps”: 1e-6, “replay_sequence_length”: 1, } - Where - prioritized_replay_alpha: Alpha parameter controls the degree of prioritization in the buffer. In other words, when a buffer sample has a higher temporal-difference error, with how much more probability should it drawn to use to update the parametrized Q-network. 0.0 corresponds to uniform probability. Setting much above 1.0 may quickly result as the sampling distribution could become heavily “pointy” with low entropy. prioritized_replay_beta: Beta parameter controls the degree of importance sampling which suppresses the influence of gradient updates from samples that have higher probability of being sampled via alpha parameter and the temporal-difference error. prioritized_replay_eps: Epsilon parameter sets the baseline probability for sampling so that when the temporal-difference error of a sample is zero, there is still a chance of drawing the sample.

  • store_buffer_in_checkpoints (bool | None) – Set this to True, if you want the contents of your buffer(s) to be stored in any saved checkpoints as well. Warnings will be created if: - This is True AND restoring from a checkpoint that contains no buffer data. - This is False AND restoring from a checkpoint that does contain buffer data.

  • lr_schedule (List[List[int | float]] | None) – Deprecated (old API stack only). Learning rate schedule in the format of [[timestep, lr-value], [timestep, lr-value], …]. Use lr with a schedule on the new API stack instead.

  • epsilon (float | List[List[int | float]] | List[Tuple[int, int | float]] | None) – Epsilon exploration schedule. In the format of [[timestep, value], [timestep, value], …]. A schedule must start from timestep 0.

  • adam_epsilon (float | None) – Adam optimizer’s epsilon hyper parameter.

  • grad_clip (int | None) – If not None, clip gradients during optimization at this value.

  • num_steps_sampled_before_learning_starts (int | None) – Number of timesteps to collect from rollout workers before we start sampling from replay buffers for learning. Whether we count this in agent steps or environment steps depends on config.multi_agent(count_steps_by=..).

  • tau (float | None) – Update the target by au * policy + (1- au) * target_policy.

  • num_atoms (int | None) – Number of atoms for representing the distribution of return. When this is greater than 1, distributional Q-learning is used.

  • v_min (float | None) – Minimum value estimation

  • v_max (float | None) – Maximum value estimation

  • noisy (bool | None) – Whether to use noisy network to aid exploration. This adds parametric noise to the model weights.

  • sigma0 (float | None) – Control the initial parameter noise for noisy nets.

  • dueling (bool | None) – Whether to use dueling DQN.

  • hiddens (int | None) – Dense-layer setup for each the advantage branch and the value branch in a dueling architecture.

  • double_q (bool | None) – Whether to use double DQN.

  • n_step (int | Tuple[int, int] | None) – N-step target updates. If >1, sars’ tuples in trajectories will be postprocessed to become sa[discounted sum of R][s t+n] tuples. An integer will be interpreted as a fixed n-step value. If a tuple of 2 ints is provided here, the n-step value will be drawn for each sample(!) in the train batch from a uniform distribution over the closed interval defined by [n_step[0], n_step[1]].

  • before_learn_on_batch (Callable[[Type[MultiAgentBatch], List[Type[Policy]], Type[int]], Type[MultiAgentBatch]]) – Callback to run before learning on a multi-agent batch of experiences.

  • training_intensity (float | None) – The intensity with which to update the model (vs collecting samples from the env). If None, uses “natural” values of: train_batch_size / (rollout_fragment_length x num_env_runners x num_envs_per_env_runner). If not None, will make sure that the ratio between timesteps inserted into and sampled from the buffer matches the given values. Example: training_intensity=1000.0 train_batch_size=250 rollout_fragment_length=1 num_env_runners=1 (or 0) num_envs_per_env_runner=1 -> natural value = 250 / 1 = 250.0 -> will make sure that replay+train op will be executed 4x asoften as rollout+insert op (4 * 250 = 1000). See: rllib/algorithms/dqn/dqn.py::calculate_rr_weights for further details.

  • td_error_loss_fn (str | None) – “huber” or “mse”. loss function for calculating TD error when num_atoms is 1. Note that if num_atoms is > 1, this parameter is simply ignored, and softmax cross entropy loss will be used.

  • categorical_distribution_temperature (float | None) – Set the temperature parameter used by Categorical action distribution. A valid temperature is in the range of [0, 1]. Note that this mostly affects evaluation since TD error uses argmax for return calculation.

  • burn_in_len (int | None) – The burn-in period for a stateful RLModule. It allows the Learner to utilize the initial burn_in_len steps in a replay sequence solely for unrolling the network and establishing a typical starting state. The network is then updated on the remaining steps of the sequence. This process helps mitigate issues stemming from a poor initial state - zero or an outdated recorded state. Consider setting this parameter to a positive integer if your stateful RLModule faces convergence challenges or exhibits signs of catastrophic forgetting.

  • **kwargs – Additional config settings, forwarded to the parent AlgorithmConfig.training() method.

Returns:

This updated AlgorithmConfig object.

Return type:

Self

validate() → None[source]#

Validates all values in this config.

get_rollout_fragment_length(worker_index: int = 0) → int[source]#

Automatically infers a proper rollout_fragment_length setting if “auto”.

Uses the simple formula: rollout_fragment_length = total_train_batch_size / (num_envs_per_env_runner * num_env_runners)

If result is a fraction AND worker_index is provided, makes those workers add additional timesteps, such that the overall batch size (across the workers) adds up to exactly the total_train_batch_size. Fractions < 1 calculated this way are rounded up to a rollout_fragment_length of 1.

Parameters:

worker_index (int) – The 1-based index of the EnvRunner asking for its fragment length (0 for the local EnvRunner). Used to distribute a fractional remainder across the first n EnvRunners.

Returns:

The user-provided rollout_fragment_length or a computed one (if user provided value is “auto”), making sure total_train_batch_size is reached exactly in each iteration.

Return type:

int

get_default_rl_module_spec() → RLModuleSpec | MultiRLModuleSpec[source]#

Returns the RLModule spec to use for this algorithm.

Override this method in the subclass to return the RLModule spec, given the input framework.

Returns:

The RLModuleSpec (or MultiRLModuleSpec) to use for this algorithm’s RLModule.

Return type:

RLModuleSpec | MultiRLModuleSpec

get_default_learner_class() → Type[Learner] | str[source]#

Returns the Learner class to use for this algorithm.

Override this method in the sub-class to return the Learner class type given the input framework.

Returns:

The Learner class to use for this algorithm either as a class type or as a string (e.g. “ray.rllib.algorithms.ppo.ppo_learner.PPOLearner”).

Return type:

Type[Learner] | str