Start thinking about policy gradient methods. What would we need to implement them? What should be the parametric form of the policy?
Start thinking about policy gradient methods. What would we need to implement them? What should be the parametric form of the policy?