CAV-Based Bias Analysis Extended to Neural L2 Speaking Assessment Systems
A recent preprint on arXiv (2608.06300) expands the use of Concept Activation Vectors (CAVs) to investigate bias in two automated speaking assessment systems: a text-oriented BERT grader and a multimodal grader that integrates speech and text, utilizing Whisper. This research builds upon earlier studies that employed CAVs to uncover biases related to undesirable attributes in feature-based grading systems. CAVs function by representing comprehensible human concepts as vectors within a model's activation space, allowing for the identification of whether a concept is present in the model's representations. The objective is to ensure that assessment scores reflect speaking skills rather than irrelevant factors like first language (L1) or age, particularly as these systems are increasingly used in critical environments. The study tackles the issues of interpretability and fairness in Transformer-based models, which, while enhancing accuracy, still operate as black boxes. The paper is available as a new submission on arXiv, identified as arXiv:2608.06300v1.
Key facts
- arXiv:2608.06300v1
- Announce Type: new
- Extends CAV-based analysis to neural speaking assessment systems
- Two systems: text-based BERT grader and speech-and-text multimodal grader based on Whisper
- CAVs used to detect bias towards unwanted attributes (concepts)
- Focus on fairness and interpretability in L2 speaking assessment
- Concerned with attributes such as first language (L1) and age
- Systems are deployed in high-stakes settings
Entities
Institutions
- arXiv