Yunhe Li

@yunhe-li · 1 works

Researches on-policy self-distillation methods for training large language models to reason.