From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
By Rahul Gupta · Paper · cs.CL
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic un