Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)By popopanda · Show HNAI Infra · Story 49133851View original