Open SourceVoice & Media12k
F5-TTS
Flow-matching zero-shot voice cloning TTS that fakes fluent speech from a short reference clip.
F5-TTS is a non-autoregressive text-to-speech model based on flow matching with diffusion transformers, producing natural, fluent speech and zero-shot voice cloning from a few seconds of reference audio. It is notable for strong quality and speed without phoneme alignment or a separate duration model, and supports multilingual and code-switched synthesis.
Repository
SWivid/F5-TTSLanguage
PythonWhat you'd build with it
- Clone a voice from a short reference clip to narrate content or build a branded voice
- Generate expressive, multilingual speech for dubbing, audiobooks, or characters
- Self-host a zero-shot TTS endpoint inside a voice-agent or media pipeline
Tags
ttsvoice-cloningflow-matchingzero-shot