Hindsight in On-policy RL
In this work, we demonstrate an approach to use hindsight experience in on-policy RL methods for goal based problems using a implicit form of importance sampling. We demonsrate significant improvement in sample efficiency for problems with dense reward but more importantly enable learning with sparse rewards previously demonstrated with only off-policy methods.