Distilled Reinforcement Learning for LLM Post-training

By Chen Wang · Paper · cs.LG

Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resultin

Cs.lg

View original

HomeResourceLoading…