ORCID

Abstract

Dysarthric speech recognition is essential for enhancing communication and accessibility for individuals with speech impairments, yet its development is hindered by a scarcity of robust, speaker-specific datasets. This study explores low-resource dysarthric speech recognition through cross-speaker transfer using synthetic data and parameter-efficient fine-tuning (PEFT). We integrate SpeechT5 text-to-speech (TTS) synthesis with x-vector speaker embeddings to generate speaker-specific dysarthric speech, enabling model adaptation while preserving pathological speech characteristics such as prosodic irregularities. Experiments on the TORGO dataset show that mixed cross-synthetic data with LoRA fine-tuning achieves a WER of 0.17, representing a 71.7% improvement over the standard model (0.60 WER) without fine-tuning the TTS model. However, cross-dataset generalisation remains challenging, yielding higher WERs on MINDS-14 (4.69) and AMI (0.96–3.83) datasets. Whilst synthetic data enhances in-domain recognition, further research is needed to improve cross-dataset generalisation and speaker adaptation, particularly for low-resource pathological speech settings.

Keywords

Cross-Speaker Transfer, Dysarthric Speech Recognition, Parameter-Efficient Fine-Tuning, Synthetic Data Generation

Publication Date

2026-01-01

Event

28th International Conference on Text, Speech, and Dialogue, TSD 2025

Publication Title

Text, Speech, and Dialogue - 28th International Conference, TSD 2025, Proceedings

Publisher

Springer Science and Business Media Deutschland GmbH

ISBN

9783032025470

ISSN

0302-9743

First Page

182

Last Page

193

Deposit Date

2026-01-20


Share

COinS