Whisper is OpenAI's speech recognition model running directly in your browser. Its multilingual model supports 99 languages, and this tool provides a selector for commonly used languages. Use Fast timing for quick transcripts, or select Precise subtitles mode for word-level timestamps and synchronized word highlighting during playback.
Choose from three model sizes: Tiny (~120MB), Base (~210MB), or Small (~590MB). Larger downloads can improve recognition on some recordings but require more memory and time. The integration can use WebGPU where supported or WebAssembly as an alternative. Audio decoding, waveform analysis, and transcription run away from the main interface so it stays responsive.
The selected recording is processed in the browser rather than uploaded to OfflineTTS for recognition. Review the result with the interactive waveform player, seek by clicking the transcript, and export plain text, SRT subtitles, or WebVTT captions. Model delivery and ordinary website requests are covered by the Privacy Policy.