PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning
By Hang Zhang · Paper · cs.CL
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data qualit