When auditing AI grading systems for fairness, you must separately check whether a biasing factor (like accent or age) is encoded in the model versus whether it actually influences the final score—these are different problems requiring different solutions.
This paper develops methods to detect bias in AI systems that automatically grade second language speaking tests. The researchers use Concept Activation Vectors to identify whether unwanted factors like a speaker's native language or age influence grades, and test whether sparse autoencoders improve bias detection in neural models like BERT and Whisper.