2026 – presentRemote
Working out why speech models fail on accented English, and what actually fixes it.
- Work-let 26LAI11: accent-invariant representation learning for SpeechLLMs, on the robustness of speech-language models across accented English.
- Established an accent-level ASR baseline on L2-ARCTIC across Arabic, Chinese, Hindi, Korean, Spanish and Vietnamese accents.
- Evaluated Whisper Large-v3 and NVIDIA Canary-1b over 3,599 scripted utterances, with Word Error Rate as the primary metric.
- Speaker- and accent-level error analysis found that speaker variation moves WER more than accent category alone — which changes what is worth adapting.
- 24.6% of clips were misrecognised by both models, giving the next stage a concrete target rather than an average to chase.
- Now exploring parameter-efficient adaptation with LoRA on those failure cases.





