Automated Essay Scoring (AES) systems are increasingly relied upon in education, but new research highlights a significant bias against English as a Second Language (ESL) learners. A recent study reveals that state-of-the-art Transformer models often penalize ESL writing, even when the quality is on par with native speakers. This bias stems from the models learning spurious correlations between surface-level linguistic features and essay quality, leading to unfair evaluations.

Quantifying the Bias: A Troubling Disparity

The bias study, detailed in a paper on arXiv, used a fine-tuned DeBERTa-v3 model and analyzed its performance on the ASAP 2.0 and ELLIPSE datasets. The results were stark: high-proficiency ESL essays received scores 10.3% lower than native speaker essays of identical human-rated quality. This "constrained score scaling," as the researchers termed it, indicates a systematic undervaluation of ESL learners' writing abilities. The Verge reports that such biases can have far-reaching consequences, affecting students' academic opportunities and self-esteem.

Contrastive Learning to the Rescue

To combat this bias, the researchers propose a novel approach: contrastive learning with matched essay pairs. This involves creating a dataset of 17,161 pairs of ESL and native-speaker essays that have been judged to be of similar quality by human raters. The model is then fine-tuned using a Triplet Margin Loss, which encourages the model to align the latent representations of ESL and native writing. This method helps the model to focus on content and argumentation rather than superficial linguistic features. "Our approach reduced the high-proficiency scoring disparity by 39.9% (to a 6.2% gap) while maintaining a Quadratic Weighted Kappa (QWK) of 0.76," the study authors note.

Disentangling Complexity and Error

Perhaps the most promising aspect of this research is the model's ability to disentangle sentence complexity from grammatical error. Post-hoc linguistic analysis suggests the model successfully learned to identify valid L2 syntactic structures, preventing the unfair penalization of sophisticated writing from ESL learners. This is crucial because ESL learners often employ complex sentence structures as they develop their writing skills. This advance has potential implications for how AES systems are designed and trained in the future. Further research is needed to explore how these debiasing techniques can be integrated into other models and datasets. It's vital that AI tools in education are fair and equitable for all students. By addressing these biases head-on, we can ensure that AES systems support and encourage ESL learners rather than inadvertently hindering their progress. This is but one small step in ensuring that AI systems help all learners achieve more.