Automatic speaking assessment systems are being used in high-stakes settings to evaluate second language learners' speaking tests, raising concerns about potential biases in their scoring. Researchers have investigated the use of concept activation vectors to analyze biases in these systems, particularly those based on transformer-based foundation models. The study examines whether these systems' scores are influenced by speaker attributes such as first language or age, rather than speaking proficiency. The analysis aims to identify potential biases in the assessment systems, which is crucial for ensuring fairness and validity in high-stakes testing. By using concept activation vectors, researchers can gain insights into the decision-making processes of these systems and identify areas where biases may exist1. This matters to practitioners because biased assessment systems can have significant consequences for language learners, making it essential to develop fair and unbiased evaluation methods.