popopanda

@popopanda · 1 works

Runs minimal LLM post-training experiments on an 8GB GPU, covering SFT, DPO, and GRPO.