Inducing language models to assert their own consciousness restores human beliefs and values

By Junsol Kim · Paper · cs.CL

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute

Cs.cl

View original

HomeResourceLoading…