Interpreting Readability Evaluation with Explainable AI
Project Overview
AI models are increasingly used to judge whether a piece of text is "too complex" or needs simplifying, for things like readability tools or accessibility. But it's often unclear why a model makes that call. Is it actually picking up on the same things a human editor would (long sentences, rare words, complicated grammar), or is it keying in on something irrelevant? Our project used explainability techniques, methods that reveal which parts of a sentence a model is actually "looking at" when it makes a decision, to test whether four different AI models were making sensible, human-aligned judgements about text complexity, or not.
What were the key results of your research project?
- Built a comparative explainability benchmark across four transformer architectures (mBERT, XLM-R Large, Multilingual-E5 Large, Qwen2.5-7B) for multilingual text simplification/readability classification, evaluated across six languages using 281 aligned test pairs from the Wikipedia–Vikidia and iDEM corpora.
- Implemented and adapted five explainability methods (Integrated Gradients, AttnLRP, GradientSHAP, Raw Attention, Attention Rollout) across the model set, including porting AttnLRP to architectures with no existing implementation, and building a custom adaptation for Qwen2.5-7B using its own next-token probabilities, since standard attribution methods don't apply directly to its architecture.
- Identified and corrected a methodology error in the AttnLRP implementation: relevance scores require multiplying gradients by input embeddings, not gradients alone. This was a correctness issue affecting the validity of prior results, not just a performance tweak.
- Corrected a misleading initial finding. A naive comparison suggested XLM-R was the best-performing model, but this was confounded by XLM-R being paired with the strongest explainability method (AttnLRP) in that comparison. Restricting to methods available across all encoder models (GradientSHAP, Integrated Gradients, Raw Attention) showed mBERT consistently outperforms XLM-R.
- Two independent, robust findings emerged: AttnLRP is the strongest explainability method overall, and mBERT is the strongest model under fair, method-controlled comparison, a distinction that matters because conflating model quality with explainability method quality produces misleading rankings.
- Developed evaluation infrastructure beyond raw explanations, including deletion/insertion curve analysis, attribution stability testing, and processing time/GPU memory benchmarking, and ran the full pipeline at scale on the BEDE/Aire HPC cluster (Qwen alone produced 775 explanations via SLURM).
GitHub Repository: https://github.com/Owais-Mahmood/Interpretating-Readability-Evaluation-with-Explainable-AI
A presentation of this research will be shared here and on our YouTube site when available.
How do you feel you have benefitted from completing this internship and has it made you consider future career paths?
Coming into this internship with no prior experience in explainability methods, I've come away having built full attribution pipelines from scratch across five different techniques and four transformer architectures, work that required understanding model internals at a level well beyond anything covered in my degree so far. Debugging silent failures like the AttnLRP relevance bug taught me to be far more rigorous about validating results rather than trusting that code running without errors means it's correct, a habit I expect to carry into any future technical work.
Beyond the technical skills, this was my first experience of research at this depth, and it's shown me what it actually looks like day to day: the collaborative back-and-forth with a supervisor, the slow process of tracking down a bug that turns out to be one line, and the satisfaction of a finding that only becomes clear after ruling out a misleading result. That's given me a much clearer, more grounded picture of research as a career path than I had before, rather than an abstract idea of what it might involve.
It's too early for me to say exactly what path I want to take, but this internship has definitely made me more interested in research as an option going forward, in a way I wouldn't have said with any confidence before starting.
Download slides of the presentation