DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
By Zhen Huang · Paper · cs.CL
Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each