after_gradient_based_update#
- Learner.after_gradient_based_update(*, timesteps: Dict[str, Any]) None[source]#
Called after gradient-based updates are completed.
Should be overridden to implement custom cleanup-, logging-, or non-gradient- based Learner/RLModule update logic after(!) gradient-based updates have been completed.
Not called for an update that was skipped (see
_should_skip_update), since there is no gradient-based update to follow: metrics of this update would hold values from earlier ones or read NaN, and anything stepped from here (target networks, schedules) would move without a gradient step behind it.