Tweedle, Daniel, Donati, Eugenio ORCID: https://orcid.org/0000-0002-0048-1858 and Wall, Julie
ORCID: https://orcid.org/0000-0001-6714-4867
(2026)
The MRP corpus: a speech dataset of modern Received Pronunciation.
Electronics, 17 (15).
|
PDF/A
TweedleD_The MRP Corpus A Speech Dataset of Modern Received Pronunciation_VoR_PDFA.pdf - Published Version Restricted to Repository staff only Available under License Creative Commons Attribution. Download (831kB) |
Abstract
Suitable speech data is crucial for the development of Mispronunciation Detection and Diagnosis (MDD) systems, which rely on examples of correct pronunciations and mispronunciations. This paper describes the creation of a novel dataset for British English MDD systems, with a focus on developing an accent-specific approach using deep learning methods. The dataset consists of over 7000 utterances from 29 speakers who were identified as using the ‘Modern Received Pronunciation’ (MRP) accent, and includes orthographic transcripts and automatically generated phone sequences. The MRP dataset aims to contribute to the development of both single-accent MDD systems and broader British English MDD. To observe the impact of single-accent training data, two identical CNN–LSTM models with Connectionist Temporal Classification (CTC) loss were trained to perform MDD. One model was trained on TIMIT, the other on our MRP data. When evaluated on L2 Arctic speech, which is annotated from an American English perspective, the TIMIT-trained model would normally be expected to perform better due to accent alignment. However, the MRP-trained model produced a lower Phoneme Error Rate (PER) and a higher rate of Correct Diagnosis. These exploratory results show that MRP can support MDD modelling, even when evaluated against an American English pronunciation target. The MRP dataset therefore offers an immediate resource for developing RP-focused pronunciation tools, a template for constructing comparable corpora across other UK accents, and an openly accessible dataset for future MDD and ASR research.
| Item Type: | Article |
|---|---|
| Identifier: | 10.3390/electronics15173873 |
| Keywords: | speech corpus;Modern Received Pronunciation; British English;mispronunciation detection and diagnosis; automatic speech recognition; accent-specific modelling |
| Subjects: | Computing > Intelligent systems |
| Date Deposited: | 07 Sep 2026 |
| Dates: | Date Publication status 26 August 2026 Accepted 28 August 2026 Published |
| School, department or research centre: | School of Computing and Engineering |
| Keywords: | speech corpus;Modern Received Pronunciation; British English;mispronunciation detection and diagnosis; automatic speech recognition; accent-specific modelling |
| URI: | https://repository.uwl.ac.uk/id/eprint/15286 |
Actions (admin access)
![]() |
Lists
Lists