Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
By Zijian Zhang · Paper · cs.LG
Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers. Existing approaches typically update all model parameters uniformly, implicitly ass