Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
By Jingwei Song · Paper · cs.LG
High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surroga