Learning When to Trust via Selective Context Preference Optimization

By Xian Sun · Paper · cs.CL

Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless wh

Cs.cl

View original

HomeResourceLoading…