The MRP corpus: a speech dataset of modern Received Pronunciation

Tweedle, Daniel, Donati, Eugenio ORCID logoORCID: https://orcid.org/0000-0002-0048-1858 and Wall, Julie ORCID logoORCID: https://orcid.org/0000-0001-6714-4867 (2026) The MRP corpus: a speech dataset of modern Received Pronunciation. Electronics, 17 (15).

[thumbnail of TweedleD_The MRP Corpus A Speech Dataset of Modern Received Pronunciation_VoR_PDFA.pdf] PDF/A
TweedleD_The MRP Corpus A Speech Dataset of Modern Received Pronunciation_VoR_PDFA.pdf - Published Version
Restricted to Repository staff only
Available under License Creative Commons Attribution.

Download (831kB)

Abstract

Suitable speech data is crucial for the development of Mispronunciation Detection and Diagnosis (MDD) systems, which rely on examples of correct pronunciations and mispronunciations. This paper describes the creation of a novel dataset for British English MDD systems, with a focus on developing an accent-specific approach using deep learning methods. The dataset consists of over 7000 utterances from 29 speakers who were identified as using the ‘Modern Received Pronunciation’ (MRP) accent, and includes orthographic transcripts and automatically generated phone sequences. The MRP dataset aims to contribute to the development of both single-accent MDD systems and broader British English MDD. To observe the impact of single-accent training data, two identical CNN–LSTM models with Connectionist Temporal Classification (CTC) loss were trained to perform MDD. One model was trained on TIMIT, the other on our MRP data. When evaluated on L2 Arctic speech, which is annotated from an American English perspective, the TIMIT-trained model would normally be expected to perform better due to accent alignment. However, the MRP-trained model produced a lower Phoneme Error Rate (PER) and a higher rate of Correct Diagnosis. These exploratory results show that MRP can support MDD modelling, even when evaluated against an American English pronunciation target. The MRP dataset therefore offers an immediate resource for developing RP-focused pronunciation tools, a template for constructing comparable corpora across other UK accents, and an openly accessible dataset for future MDD and ASR research.

Item Type: Article
Identifier: 10.3390/electronics15173873
Keywords: speech corpus;Modern Received Pronunciation; British English;mispronunciation detection and diagnosis; automatic speech recognition; accent-specific modelling
Subjects: Computing > Intelligent systems
Date Deposited: 07 Sep 2026
Dates:
Date
Publication status
26 August 2026
Accepted
28 August 2026
Published
School, department or research centre: School of Computing and Engineering
Keywords: speech corpus;Modern Received Pronunciation; British English;mispronunciation detection and diagnosis; automatic speech recognition; accent-specific modelling
URI: https://repository.uwl.ac.uk/id/eprint/15286

Actions (admin access)

View Item

Menu