Reinforcement Learning without Ground-Truth Solutions can Improve LLMs

By Yingyu Lin · Paper · cs.LG

Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \textbf{R}anking-\textbf{i}nduced \textbf{VER}ifiable

Cs.lg

View original

HomeResourceLoading…