Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

By Priyank Agrawal · Paper · cs.LG

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot generate any correct solutions, it receives \textit{zero} learning signal. Providing privileged guidance

Cs.lg

View original

HomeResourceLoading…