Shayan Talaei

@shayan-talaei · 1 works

Investigates bias and preference steering in deployed language models.