HeyJay! A corpus of atypical speech for spoken language understanding and automatic speech recognition.
HeyJay! A corpus of atypical speech for spoken language understanding and automatic speech recognition.
Where did the research take place?
The study site has not been established. Author addresses may differ from where the research occurred.
Hopkins, US · Author affiliation
Johns Hopkins University, Dept. of Electrical and Computer Engineering, Baltimore, 21218, USA. laureano@jhu.edu.Location evidence
Baltimore, US · Author affiliation
Johns Hopkins University, Dept. of Electrical and Computer Engineering, Baltimore, 21218, USA. laureano@jhu.edu.Location evidence
Seattle, US · Author affiliation
Amazon Inc., Seattle, 98109, USA.Location evidence
A plain-language reading has not been prepared for this paper yet.
Original abstract
Speech technologies, such as automatic speech recognition or spoken language understanding, are not usually adapted to atypical speech, i.e., the speech of people with dysarthria, dysphonia, or another type of speech impairment. That prevents atypical speakers from leveraging speech assistants or other human-machine-interaction-powered platforms, which could make their lives easier or increase their independence. In this article, we present HeyJay!, a new corpus of atypical speech in English language from participants with neurodegenerative disorders, including Parkinson's Disease, or Amyotrophic Lateral Sclerosis. The current corpus version comprises 8,669 utterance recordings, including supervised transcriptions and intent annotations. In this study, we demonstrate the validity of the corpus by applying it to automatic speech recognition, spoken language understanding, and data augmentation tasks. Additionally, the dataset includes speech quality ratings for each participant, performed by expert speech and language pathologists. This corpus, the first one with intent annotation of atypical speech that is publicly available, is intended to create more fair speech technologies for atypical speakers by adapting and improving the state of the art, and to facilitate further research in the field.