Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?

By Perry Dong · Paper · cs.LG

Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wi

Cs.lg

View original

HomeResourceLoading…