Learning When to Trust via Selective Context Preference Optimization
By Xian Sun · Paper · cs.CL
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless wh