Scheduling Mixed RL Rollouts Beyond Prefix Locality
By Zetao Hong · Paper · cs.DC
Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it doe