TrustNLP Workshop: Six Years of AI Trust Research
Since its inception in 2021, the TrustNLP Workshop has been held alongside major ACL conferences and has expanded from 8 to 41 proceedings papers over six editions. This growth highlights a shift in the field from merely interpreting static models to achieving a mechanistic grasp and proactive management of generative systems. A comprehensive review of the 144 proceedings papers categorizes them into six trust dimensions based on established frameworks (TrustLLM, DecodingTrust), indicating overlaps with capability emergence. The introduction of the first impactful chat models engaged all trust dimensions, while later models concentrated on truthfulness and safety alignment. Notably, truthfulness, which was absent in 2021-2022, is projected to represent 37% of papers by 2025-2026, whereas fairness remains a consistent theme. This evolution underscores the increasing significance of trust in AI as generative systems gain traction.
Key facts
- TrustNLP Workshop co-located with ACL conferences since 2021
- Grew from 8 to 41 proceedings papers over six editions
- Synthesized insights from all 144 proceedings papers
- Classified papers along six trust dimensions
- Frameworks used: TrustLLM, DecodingTrust
- First high-impact chat models activated all trust dimensions
- Truthfulness grew from absent in 2021-2022 to 37% by 2025-2026
- Fairness is the most consistent theme
Entities
Institutions
- TrustNLP Workshop
- ACL