All open source
Open SourceVoice & Media12k

F5-TTS

Flow-matching zero-shot voice cloning TTS that fakes fluent speech from a short reference clip.

F5-TTS is a non-autoregressive text-to-speech model based on flow matching with diffusion transformers, producing natural, fluent speech and zero-shot voice cloning from a few seconds of reference audio. It is notable for strong quality and speed without phoneme alignment or a separate duration model, and supports multilingual and code-switched synthesis.

Repository

SWivid/F5-TTS

Language

Python

What you'd build with it

  • Clone a voice from a short reference clip to narrate content or build a branded voice
  • Generate expressive, multilingual speech for dubbing, audiobooks, or characters
  • Self-host a zero-shot TTS endpoint inside a voice-agent or media pipeline

Tags

ttsvoice-cloningflow-matchingzero-shot