SPADE: Self-Play in Adaptive Synthetic Executable Environments

By Bo Liu · Paper · cs.CL

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We i

Cs.cl

View original

HomeResourceLoading…