before_gradient_based_update#

Learner.before_gradient_based_update(*, timesteps: Dict[str, Any]) → None[source]#

Called before gradient-based updates are completed.

Should be overridden to implement custom preparation-, logging-, or non-gradient-based Learner/RLModule update logic before(!) gradient-based updates are performed. Called after the train batch has been built, and not at all for an update that was skipped (see _should_skip_update), so that this hook and after_gradient_based_update always run as a pair.

Parameters:

timesteps (Dict[str, Any]) – Timesteps dict, which must have the key NUM_ENV_STEPS_SAMPLED_LIFETIME. # TODO (sven): Make this a more formal structure with its own type.