How to Train a Critic Stably and Efficiently

By Penghui Qi · Paper · cs.LG

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes

Cs.lg

View original

HomeResourceLoading…