Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
By Lars van der Laan · Paper · cs.LG
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy