Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning

By Lars van der Laan · Paper · cs.LG

Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy

Open Models · Cs.lg

View original

HomeResourceLoading…