MemoryTrackingCallbacks#

class ray.rllib.callbacks.callbacks.MemoryTrackingCallbacks[source]#

Bases: RLlibCallback

MemoryTrackingCallbacks can be used to trace and track memory usage in rollout workers.

The Memory Tracking Callbacks uses tracemalloc and psutil to track python allocations during rollouts, in training or evaluation.

The tracking data is logged to the custom_metrics of an episode and can therefore be viewed in tensorboard (or in WandB etc..)

Add MemoryTrackingCallbacks callback to the tune config e.g. { …’callbacks’: MemoryTrackingCallbacks …}

Note

This class is meant for debugging and should not be used in production code as tracemalloc incurs a significant slowdown in execution speed.

on_episode_end(*, episode: SingleAgentEpisode | MultiAgentEpisode | EpisodeV2, env_runner: EnvRunner | None = None, metrics_logger: MetricsLogger | None = None, env: gymnasium.Env | None = None, env_index: int, rl_module: RLModule | None = None, worker: EnvRunner | None = None, base_env: BaseEnv | None = None, policies: Dict[str, Policy] | None = None, **kwargs) → None[source]#

Called when an episode is done (after terminated/truncated have been logged).

The exact time of the call of this callback is after env.step([action]) and also after the results of this step (observation, reward, terminated, truncated, infos) have been logged to the given episode object, where either terminated or truncated were True:

  • The env is stepped: final_obs, rewards, ... = env.step([action])

  • The step results are logged episode.add_env_step(final_obs, rewards)

  • Callback on_episode_step is fired.

  • Another env-to-module connector call is made (even though we won’t need any RLModule forward pass anymore). We make this additional call to ensure that in case users use the connector pipeline to process observations (and write them back into the episode), the episode object has all observations - even the terminal one - properly processed.

  • —> This callback on_episode_end() is fired. <—

  • The episode is numpy’ized (i.e. lists of obs/rewards/actions/etc.. are converted into numpy arrays).

Parameters:
  • episode (SingleAgentEpisode | MultiAgentEpisode | EpisodeV2) – The terminated/truncated SingleAgent- or MultiAgentEpisode object (after env.step() that returned terminated=True OR truncated=True and after the returned obs, rewards, etc.. have been logged to the episode object). Note that this method is still called before(!) the episode object is numpy’ized, meaning all its timestep data is still present in lists of individual timestep data.

  • prev_episode_chunks – A complete list of all previous episode chunks with the same ID as episode that have been sampled on this EnvRunner. In order to compile metrics across the complete episode, users should loop through the list: [episode] + previous_episode_chunks and accumulate the required information.

  • env_runner (EnvRunner | None) – Reference to the EnvRunner running the env and episode.

  • metrics_logger (MetricsLogger | None) – The MetricsLogger object inside the env_runner. Can be used to log custom metrics during env/episode stepping.

  • env (gymnasium.Env | None) – The gym.Env or gym.vector.Env object running the started episode.

  • env_index (int) – The index of the sub-environment that has just been terminated or truncated.

  • rl_module (RLModule | None) – The RLModule used to compute actions for stepping the env. In single-agent mode, this is a simple RLModule, in multi-agent mode, this is a MultiRLModule.

  • worker (EnvRunner | None) – Old API stack only. Reference to the RolloutWorker running the episode. Deprecated in favor of env_runner.

  • base_env (BaseEnv | None) – Old API stack only. The BaseEnv running the episode. Deprecated in favor of env.

  • policies (Dict[str, Policy] | None) – Old API stack only. Mapping from PolicyID to Policy object used to compute actions. Deprecated in favor of rl_module.

  • **kwargs – Forward compatibility placeholder.