TTPO: Test-Time Policy Optimization

By Aozhe Wang · Paper · cs.CL

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replac

Cs.cl

View original

HomeResourceLoading…