SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
By Kai Ruan · Paper · cs.AI
Group-relative reinforcement learning waits for sibling rollouts of the same prompt, which is costly for long and variable tool-use trajectories. Single-stream Policy Optimization (SPO) removes this dependency with a persistent prompt-level value estimate, but its recipe whitens