Ito: Natural Streaming Text-to-Speech on a $5 Microcontroller

Daniel Varela

Lokutor (Hashing Works, S.L.L.), Madrid

Abstract

Ito is a 4.0 M-parameter neural text-to-speech model that runs entirely on an ESP32-S3, a 240 MHz dual-core microcontroller with no neural accelerator. A causal acoustic model predicts a mel spectrogram with bounded lookahead, and a small convolutional vocoder with a harmonic source turns it into a 24 kHz waveform, so audio streams out while the sentence is still being synthesized. The model runs in int8 on the chip's vector unit and takes 4.9 MB of flash. In a blind listening test it was rated 4.0 out of 5 against 2.0 for sanoTTS, the previous best open model for this chip, and on a 54-sentence benchmark its automatic naturalness score (UTMOS22 4.44) is within 0.03 of the 191 M-parameter model it was distilled from.

All Ito samples below are the exact output of the on-chip engine: the firmware running in Espressif's QEMU emulator produces bit-identical audio. Time to first audio and real-time factor are estimated from exact QEMU instruction counts and an assumed PSRAM bandwidth, not measured on silicon: about 145 ms to first audio in the central case (94 to 240 ms across the range), and a real-time factor of 0.77 to 0.79 centrally (0.53 to 1.22 across the range, so the pessimistic case is still slower than real time). The first chunk is only 25 ms of audio, so playback without gaps needs a start delay of about 340 ms (central). Real-time playback is not yet established on silicon.

Samples

Eight sentences that none of the models saw in training. Ito runs on the microcontroller, in two voices (female and male); the other columns are references.

TextIto, female voice (ours, ESP32-S3)Ito, male voice (ours, ESP32-S3)sanoTTS amysanoTTS heart-nanoTeacher (StyleTTS 2, server)
Hey, are you still coming over for dinner tonight, or should I save you a plate?
Your package should arrive on Friday, October 9th, sometime before noon.
When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company.
The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first.
Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning.
Could you grab some quinoa and Worcestershire sauce on your way home?
It's about 23 degrees outside, so you probably won't need a jacket.
I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up.

sanoTTS clips were generated with the released sanoTTS package. The teacher is the public StyleTTS 2 LibriTTS model in its female voice. All clips are loudness-matched.

Results

Automatic metrics for 17 systems on 54 fixed English prompts that none of them saw in training. Ito is scored on the chip engine's exact output. The full table, with 22 systems and confidence intervals, is in the repository.

SystemParamsRuns on a microcontrollerUTMOSv2UTMOS22DNSMOSWER %
Built for microcontrollers
Ito (ours)4.05 MESP32-S3 (emulated, bit-exact)3.154.443.390.4
Inflect Nano v2 (Owen Song)3.97 MESP32-P4, 3.5× slower than real time3.084.413.401.1
TinyTTS1.6 MESP32-S3, 22.9× slower than real time2.453.663.296.8
sanoTTS amy1.45 Mnoa2.803.963.181.2
sanoTTS heart-nano0.29 MMCU-sized1.332.172.971.7
eSpeak NG (rule-based)—community ports1.742.142.760.3
Larger models (CPU or GPU)
Inflect Micro v2 (Owen Song)9.36 Mno3.464.413.381.0
Kitten TTS nano14.0 Mno1.993.933.311.1
Piper amy low15.6 Mno3.424.443.280.7
Piper lessac medium15.7 Mno3.694.283.290.7
MeloTTS EN51.9 Mno3.033.723.013.2
Supertonic 265 Mno3.624.443.352.4
Kokoro-82M81.8 Mno3.874.493.410.8
Supertonic 399 Mno3.844.453.321.7
MOSS-TTS-Nano~100 Mno3.444.373.211.7
Pocket TTS110 Mno3.314.333.312.5
StyleTTS 2 (our teacher)191 Mno3.434.473.341.5
Higher is better for UTMOSv2, UTMOS22 and DNSMOS (overall); WER from Whisper large-v3. 95% confidence intervals are about ±0.05–0.1 for the quality scores. a The sanoTTS model that runs on the chip has 567 K parameters; its paper reports UTMOS 2.80.

Among systems built for a microcontroller, Ito has the highest UTMOS22 and DNSMOS and the lowest word error rate of the neural systems. Inflect Nano (Owen Song), a 4 M-parameter model, ties it on UTMOSv2 but runs 3.5 times slower than real time on a larger chip. Several models 2 to 50 times larger, which need a CPU or GPU, score higher on UTMOSv2.

Listening test

SystemRating (1–5)
Teacher, StyleTTS 24.75
Ito (ours)4.00
sanoTTS amy2.00
sanoTTS heart-nano1.00
Blind, unlabeled clips rated for how human they sound; 4 sentences per system, one expert listener. A larger public test is planned.

On the chip

Weights4.9 MB int8, in flash
Computeabout 350 M int8 multiply-accumulates per second of audio
Streamingyes; time to first audio does not depend on sentence length
Time to first audioestimated, not measured on silicon: 94 to 97 ms optimistic, 145 to 148 ms central, 237 to 240 ms pessimistic (any sentence length)
Real-time factorestimated: 0.53 to 0.55 optimistic, 0.77 to 0.79 central, 1.18 to 1.22 pessimistic (below 1 is faster than real time, so the pessimistic case is still slower). Real time is not yet established on silicon
Start delay for gapless speechestimated: about 220 ms optimistic, about 340 ms central, 850 ms or more pessimistic. The first chunk is only 25 ms of audio and the next is 300 ms, so time to first audio and gapless speech are different numbers

The firmware measures its own throughput at boot and prints a BOARD_SUMMARY line. If you run it on an ESP32-S3 board, please share the line in an issue.

Try it

pip install git+https://github.com/lokutor-ai/ito
ito-tts "Good morning! The coffee is ready." -o hello.wav

Flashing instructions for the ESP32-S3 are in the repository.

License

The engine and firmware code are GPLv3. The model weights are CC BY-NC-SA 4.0 with additional terms: no commercial use. For commercial licenses, or a free license for a small company, maker project or research group, write to contact@lokutor.com.