How to Train a Critic Stably and Efficiently
By Penghui Qi · Paper · cs.LG
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes