popopanda@popopanda · 1 worksRuns minimal LLM post-training experiments on an 8GB GPU, covering SFT, DPO, and GRPO.