EMI Lab · arXiv preprint · 2026

GrainSpeechLess Context, More Detail
for Compact Speech Synthesis

GrainSpeech turns written text into speech using a small AI model.

Zitao LiangChang Gao*

Efficient Machine Intelligence Lab · TU Delft
* Corresponding author

01 / Audio demo

Hear the difference.

Choose a sentence. Listen to GrainSpeech,
then compare it with the original recording.

Choose a sentence
SENTENCE 01LJ015-0096

“whether representing real or fictitious shares does not appear; but they were certificates connected in some way with robson's long practiced frauds”

Compare all models 12 acoustic variants · one shared sentence

GrainSpeech ablation

EfficientSpeech family

Other acoustic models

Five shared LJSpeech utterances. Every synthesized sample uses the same HiFi-GAN v2 vocoder; original recordings bypass it. Audio details & provenance

02 / The idea

Less context.
More acoustic detail.

GrainSpeech explores how encoder context and Mel supervision affect compact speech synthesis. Its acoustic model has 264.8K parameters.

01

Focus the encoder

A fixed-receptive-field convolutional encoder uses local phoneme context to predict pitch, energy and duration.

02

Preserve fine structure

L1, SSIM and a Mel-specific gradient-variance objective supervise local detail along time and frequency.

GrainSpeech architecture: phonemes enter a fixed-receptive-field encoder, pitch, energy and duration predictors condition a decoder, and Mel spectrograms are supervised with L1, SSIM and Mel-GVar losses.
From phonemes to a Mel spectrogram. HiFi-GAN converts the predicted Mel into the audio heard above. Explore the method in the paper ↗

03 / Results

Small enough for the edge.

Full evaluation ↗
264.8K

Acoustic model parameters

4.086

UTMOS on 128 LJSpeech utterances

17.9×

Real-time Mel generation on MCU

MCU result: A16W8 on STM32H747XI, 483.19 KiB peak SRAM. This measures acoustic-model inference; vocoding is excluded.

Selected results from Table 1 of the preprint
Acoustic model Parameters UTMOS ↑ WER (%) ↓
GrainSpeech This work 0.265M 4.086 3.27
EfficientSpeech-Tiny 0.266M 3.591 3.08
MatchaTTS 18.204M 4.264 1.98
MixerTTS 20.060M 4.087 1.80

Automatic metrics on LJSpeech. External models use their native training settings. See the paper for confidence intervals, spectral distortion and evaluation limitations.

04 / Open research

Read it.
Run it. Build on it.

The acoustic model, pretrained weights, inference and training code, and this paper website live in one repository.

Get started on GitHub

Cite GrainSpeech

@article{liang2026grainspeech,
  title = {GrainSpeech: Less Context, More Detail
           for Compact Speech Synthesis},
  author = {Liang, Zitao and Gao, Chang},
  journal = {arXiv preprint arXiv:2609.18856},
  year = {2026},
  doi = {10.48550/arXiv.2609.18856},
  url = {https://arxiv.org/abs/2609.18856}
}