All open source
Open SourceVoice & Media78k

Whisper

OpenAI's robust speech-to-text model for transcription and translation.

Whisper is OpenAI's open-source automatic speech recognition model and reference implementation, trained on a large multilingual and multitask dataset. It performs transcription and translation across many languages and is distributed in several model sizes that trade off accuracy for speed and memory.

Repository

openai/whisper

Language

Python

What you'd build with it

  • Transcribe audio or video into text for captioning, search, or downstream LLM processing
  • Translate non-English speech directly into English text
  • Generate transcripts as input for voice-driven AI applications and meeting/podcast summarization

Tags

stttranscriptionopenai