Lokutor (Hashing Works, S.L.L.), Madrid
Code · Model · Paper (coming soon)
Ito is a 4.0 M-parameter neural text-to-speech model that runs entirely on an ESP32-S3, a 240 MHz dual-core microcontroller with no neural accelerator. A causal acoustic model predicts a mel spectrogram with bounded lookahead, and a small convolutional vocoder with a harmonic source turns it into a 24 kHz waveform, so audio streams out while the sentence is still being synthesized. The model runs in int8 on the chip's vector unit and takes 4.9 MB of flash. In a blind listening test it was rated 4.0 out of 5 against 2.0 for sanoTTS, the previous best open model for this chip, and on a 54-sentence benchmark its automatic naturalness score (UTMOS22 4.44) is within 0.03 of the 191 M-parameter model it was distilled from.
All Ito samples below are the exact output of the on-chip engine: the firmware running in Espressif's QEMU emulator produces bit-identical audio. Time to first audio and real-time factor are estimated from exact QEMU instruction counts and an assumed PSRAM bandwidth, not measured on silicon: about 145 ms to first audio in the central case (94 to 240 ms across the range), and a real-time factor of 0.77 to 0.79 centrally (0.53 to 1.22 across the range, so the pessimistic case is still slower than real time). The first chunk is only 25 ms of audio, so playback without gaps needs a start delay of about 340 ms (central). Real-time playback is not yet established on silicon.
Eight sentences that none of the models saw in training. Ito runs on the microcontroller, in two voices (female and male); the other columns are references.
| Text | Ito, female voice (ours, ESP32-S3) | Ito, male voice (ours, ESP32-S3) | sanoTTS amy | sanoTTS heart-nano | Teacher (StyleTTS 2, server) |
|---|---|---|---|---|---|
| Hey, are you still coming over for dinner tonight, or should I save you a plate? | |||||
| Your package should arrive on Friday, October 9th, sometime before noon. | |||||
| When I finally got to the station, the last train had already left, so I ended up sharing a taxi with two strangers who turned out to be surprisingly good company. | |||||
| The pharmacist recommended an anti-inflammatory, but honestly, I'd rather try physiotherapy first. | |||||
| Thanks so much for calling. I'll check the schedule and get back to you first thing tomorrow morning. | |||||
| Could you grab some quinoa and Worcestershire sauce on your way home? | |||||
| It's about 23 degrees outside, so you probably won't need a jacket. | |||||
| I know it sounds strange, but I actually enjoy the quiet hours before everyone else wakes up. |
sanoTTS clips were generated with the released sanoTTS package. The teacher is the public StyleTTS 2 LibriTTS model in its female voice. All clips are loudness-matched.
Automatic metrics for 17 systems on 54 fixed English prompts that none of them saw in training. Ito is scored on the chip engine's exact output. The full table, with 22 systems and confidence intervals, is in the repository.
| System | Params | Runs on a microcontroller | UTMOSv2 | UTMOS22 | DNSMOS | WER % |
|---|---|---|---|---|---|---|
| Built for microcontrollers | ||||||
| Ito (ours) | 4.05 M | ESP32-S3 (emulated, bit-exact) | 3.15 | 4.44 | 3.39 | 0.4 |
| Inflect Nano v2 (Owen Song) | 3.97 M | ESP32-P4, 3.5× slower than real time | 3.08 | 4.41 | 3.40 | 1.1 |
| TinyTTS | 1.6 M | ESP32-S3, 22.9× slower than real time | 2.45 | 3.66 | 3.29 | 6.8 |
| sanoTTS amy | 1.45 M | noa | 2.80 | 3.96 | 3.18 | 1.2 |
| sanoTTS heart-nano | 0.29 M | MCU-sized | 1.33 | 2.17 | 2.97 | 1.7 |
| eSpeak NG (rule-based) | — | community ports | 1.74 | 2.14 | 2.76 | 0.3 |
| Larger models (CPU or GPU) | ||||||
| Inflect Micro v2 (Owen Song) | 9.36 M | no | 3.46 | 4.41 | 3.38 | 1.0 |
| Kitten TTS nano | 14.0 M | no | 1.99 | 3.93 | 3.31 | 1.1 |
| Piper amy low | 15.6 M | no | 3.42 | 4.44 | 3.28 | 0.7 |
| Piper lessac medium | 15.7 M | no | 3.69 | 4.28 | 3.29 | 0.7 |
| MeloTTS EN | 51.9 M | no | 3.03 | 3.72 | 3.01 | 3.2 |
| Supertonic 2 | 65 M | no | 3.62 | 4.44 | 3.35 | 2.4 |
| Kokoro-82M | 81.8 M | no | 3.87 | 4.49 | 3.41 | 0.8 |
| Supertonic 3 | 99 M | no | 3.84 | 4.45 | 3.32 | 1.7 |
| MOSS-TTS-Nano | ~100 M | no | 3.44 | 4.37 | 3.21 | 1.7 |
| Pocket TTS | 110 M | no | 3.31 | 4.33 | 3.31 | 2.5 |
| StyleTTS 2 (our teacher) | 191 M | no | 3.43 | 4.47 | 3.34 | 1.5 |
Among systems built for a microcontroller, Ito has the highest UTMOS22 and DNSMOS and the lowest word error rate of the neural systems. Inflect Nano (Owen Song), a 4 M-parameter model, ties it on UTMOSv2 but runs 3.5 times slower than real time on a larger chip. Several models 2 to 50 times larger, which need a CPU or GPU, score higher on UTMOSv2.
| System | Rating (1–5) |
|---|---|
| Teacher, StyleTTS 2 | 4.75 |
| Ito (ours) | 4.00 |
| sanoTTS amy | 2.00 |
| sanoTTS heart-nano | 1.00 |
| Weights | 4.9 MB int8, in flash |
| Compute | about 350 M int8 multiply-accumulates per second of audio |
| Streaming | yes; time to first audio does not depend on sentence length |
| Time to first audio | estimated, not measured on silicon: 94 to 97 ms optimistic, 145 to 148 ms central, 237 to 240 ms pessimistic (any sentence length) |
| Real-time factor | estimated: 0.53 to 0.55 optimistic, 0.77 to 0.79 central, 1.18 to 1.22 pessimistic (below 1 is faster than real time, so the pessimistic case is still slower). Real time is not yet established on silicon |
| Start delay for gapless speech | estimated: about 220 ms optimistic, about 340 ms central, 850 ms or more pessimistic. The first chunk is only 25 ms of audio and the next is 300 ms, so time to first audio and gapless speech are different numbers |
The firmware measures its own throughput at boot and prints a BOARD_SUMMARY line. If you run it on an ESP32-S3 board, please share the line in an issue.
pip install git+https://github.com/lokutor-ai/ito ito-tts "Good morning! The coffee is ready." -o hello.wav
Flashing instructions for the ESP32-S3 are in the repository.
The engine and firmware code are GPLv3. The model weights are CC BY-NC-SA 4.0 with additional terms: no commercial use. For commercial licenses, or a free license for a small company, maker project or research group, write to contact@lokutor.com.