From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

By Rahul Gupta · Paper · cs.CL

The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic un

Cs.cl

View original

HomeResourceLoading…