Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training
By Zachary Speck · Paper · cs.LG
A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. We ran that counterfactual 24 times at a small scale. We trained 32 GPT-2