Enhancing Pathological Speech through Articulatory Bottlenecks and Phonetic-Aware Losses

AilĂ­n Pollio San Pedro, Olivier Perrotin, Thomas Hueber
University Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, France

Abstract. Dysarthric speech reconstruction (DSR) typically relies on linguistic or phonetic representations extracted from impaired speech to generate a more intelligible waveform. We investigate a complementary approach that instead intervenes in a representation related to speech production. Starting from a neural analysis-synthesis framework, we introduce a residual mapper that modifies an articulatory-aligned latent space while preserving speaker and prosodic information. The mapper is pretrained on parallel synthetic healthy and artificially dysarthric speech, then adapted to natural dysarthric speech using phoneme-guided and adversarial objectives. We further compare the articulatory bottleneck with a dimension-matched unsupervised representation to assess the benefit of explicit articulatory supervision. Objective and perceptual evaluations show that the proposed approach successfully improves artificially generated dysarthric speech. However, these improvements transfer only in a limited number of cases to real pathological speech, highlighting the challenge of bridging synthetic articulatory degradations and natural dysarthria.

Synthetic pathological speech

Healthy: original TTS; Dysarthric: controlled perturbation; SAB: speech-articulatory bottleneck; UAB: unsupervised bottleneck.

Target text / phonetic targetHealthyDysarthricSABUAB
It had been early afternoon then.
HealthyDysarthric (moderate)SABUAB
The coal has given out.
HealthyDysarthric (severe)SABUAB
But there was one, Frederic?
HealthyDysarthric (moderate)SABUAB
I won't say another word.
HealthyDysarthric (severe)SABUAB
She will look after things.
HealthyDysarthric (moderate)SABUAB
He did not believe her.
HealthyDysarthric (severe)SABUAB

Natural pathological speech (UASpeech)

Selected isolated words with potential ASR changes after enhancement.

Target / speakerOriginalUABSAB
avenues
F04
OriginalUABSAB
battleship
F04
OriginalUABSAB
demolish
F05
OriginalUABSAB
backspace
M05
OriginalUABSAB
clown
M08
OriginalUABSAB
battleship
M09
OriginalUABSAB
washerwoman
M09
OriginalUABSAB
hypothesis
M09
OriginalUABSAB
southeast
M09
OriginalUABSAB
dodgers
M10
OriginalUABSAB

Notation

SAB: speech-articulatory bottleneck. UAB: unsupervised bottleneck. Audio examples are provided for qualitative listening.