From Formants to Foundation Models

A friendly history of open-source text-to-speech

Every era of TTS moved the unit of generation down a level of abstraction — from rules, to recorded fragments, to parameters, to raw samples, to tokens. You can hear each step.
Machine Learning
Speech
History
Author

Ravi Kalia

Published

August 3, 2026

From Formants to Foundation Models

This post is a history of open-source text-to-speech told through one measure — what the system actually generates — with audio for each step so you can hear the change.

Text-to-speech (TTS) progress is audible: each era changed the unit of generation — the smallest thing a system produces before assembling speech.

Abstraction levels, oldest to newest:

Each step down increases naturalness and reduces direct control. Old engines still run; sections below include audio players. Four of five clips were generated locally for this post.

1 Five eras by generation unit

Classify a system by what it generates: a rule, fragment, parameter set, sample, or token. Sound quality, disk use, and voice controllability follow from that choice.

1990s · Rules and fragments

Formant synthesis generates speech from hand-written rules, storing no recorded audio at all. Diphone concatenation splices together one recorded copy of each sound-to-sound transition. Intelligible, unmistakably synthetic.

2 1990s: formant and diphone synthesis

Formant synthesis — no stored speech. Vocal-tract resonances (formants) are modeled as filter settings driven by rules: letters → phonemes → oscillator/filter parameters. Used by DECtalk and similar systems.

Diphone concatenation — record one copy of each sound-to-sound transition; splice and pitch-shift to build sentences. Festival (1996, CSTR Edinburgh) is the open-source ancestor.

Trade-off: arbitrary text on minimal storage; every syllable sounds identical (single copy per unit).

3 2000s: unit selection and parametric synthesis

Both 1990s approaches failed at joins between fragments recorded at different times.

Unit selection — record hours of one speaker; at synthesis time, pick the longest matching contiguous stretch. Fewer cuts; when a phrase exists in the database, output is real recorded speech. Misses produce register shifts mid-sentence.

Statistical parametric synthesis — store no waveforms; model spectrum, pitch, and duration; rebuild audio with a vocoder. HTS used hidden Markov models. Small footprint, easy speaker control; muffled output from averaging many pronunciations.

Trade-off: hoard audio vs model parameters; neither achieved both naturalness and flexibility until neural methods.

4 2010s: neural end-to-end stacks

WaveNet (2016) — autoregressive prediction of raw waveform samples (~16 kHz per second of audio). Slow; highest quality of its era.

Tacotron (2017) — seq2seq maps characters to mel spectrogram; learns pronunciation and prosody from data instead of hand rules.

Tacotron 2 + WaveNet vocoder — mean opinion score (MOS) 4.53 vs 4.58 for human recordings.

WaveGlow (2018) — parallel flow-based vocoder; faster than sample-by-sample generation.

Trade-off: linguistic front-end became learned; speed remained the bottleneck.

5 Early 2020s: non-autoregressive deployment

Autoregression blocks parallel compute.

  • FastSpeech 2 (2020) — predict phoneme durations; emit full spectrogram in parallel.
  • HiFi-GAN (2020) — GAN vocoder; hundreds of times faster than real time on GPU; faster than real time on laptop CPU.
  • VITS (2021) — end-to-end training; no separate vocoder stage.

Trade-off: deployable open-source TTS; fixed voices per training run.

6 Codec tokens and speech LLMs

Neural audio codec — compress waveform to discrete codec tokens; decode back to audio. Speech becomes a token sequence; transformer tooling applies directly.

Examples: Chatterbox (Llama-style backbone on speech tokens); CosyVoice 2, OpenAudio S1-mini, IndexTTS-2 (variants on the theme).

Trade-off: voice and emotion via prompt/reference audio; autoregression returns latency and garbling risk; cloning enables misuse (watermarking in models like Chatterbox).

7 Audio samples

Five clips, one sentence each (except Tacotron 2 — see note). Four generated locally via make_audio.py. Living descendants of historical engines, not original period binaries.

1990s · formant synthesis eSpeak NG 1.52

No recorded human speech anywhere in this file. Every sound is rules driving filters.

1990s · diphone concatenation Flite 2.2, cmu_us_kal

A real person’s voice, chopped into transitions and pitch-shifted back together. Listen for the seams.

2000s · statistical parametric Flite 2.2, cmu_us_slt Clustergen

Smooth where the diphone voice was jagged, muffled where it was sharp. Clustergen is the same family as HMM-based HTS, not literally HTS.

2010s · seq2seq + neural vocoder Tacotron 2 (Google)

The one clip not generated here, and a different sentence: Tacotron 2’s checkpoint was never released, so this is mirrored from Google’s demo page.

2026 · codec-era neural TTS Kokoro-82M, voice af_heart

82 million parameters, generated on a laptop CPU in seconds. Try other voices in the official Kokoro Space.

Clip 3 → clip 5 spans the main neural quality jump.

8 Current open-source models

Five downloadable models (August 2026): local execution, permissive licenses, codec-token or late non-autoregressive era.

Leading open-source TTS models as of August 2026. IndexTTS-2’s parameter count is not stated in its repository, model card or paper, so it is left blank rather than guessed.
Model Provider Parameters Underlying tech Licence
Kokoro-82M hexgrad 82M Non-autoregressive StyleTTS 2 architecture with an ISTFTNet decoder Apache 2.0
Chatterbox / Turbo Resemble AI 500M / 350M Llama-style token generator plus a distilled diffusion decoder MIT
CosyVoice 2 Alibaba (FunAudioLLM) 0.5B LLM-based codec TTS, streaming-optimised to ~150 ms latency Apache 2.0
OpenAudio S1-mini Fish Audio 0.5B Dual autoregressive (Dual-AR) codec architecture with GFSQ Apache 2.0
IndexTTS-2 Bilibili not published Autoregressive text-to-semantic then semantic-to-mel, with duration and emotion control Apache 2.0

9 Constraints

Parameter count does not rank quality here. Kokoro (~6× smaller than neighboring codec models) ranks high in community listening arenas because it does fixed voices only — no cloning or emotion prompts. Larger models buy controllability, not necessarily better sentences. Naturalness scores vary with listener pool, text, and evaluator.

10 Direction of travel

TTS as a separate component may disappear: one model can emit speech tokens directly, avoiding text→speech prosody mismatch.

Expected trends:

  • Streaming by default
  • Emotion/style in prompt
  • Watermarking (voice cloning from seconds of reference audio)

Each era moved generation downward (rules → fragments → parameters → samples → tokens), trading control for naturalness. Next step: no separate synthesizer — same model produces words and sound.

Rules. Became. Fragments. Fragments. Became. Samples. Samples. Became. Tokens. Naturalness. Cost. Control.

11 References