Inducing language models to assert their own consciousness restores human beliefs and values
By Junsol Kim · Paper · cs.CL
Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute