ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
By Jinhe Bi · Paper · cs.AI
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-