Browser STT workspace

Whisper Speech to Text

Upload audio or video, record from your microphone, or load a direct media URL. Transcribe privately in your browser and export TXT, SRT, or VTT.

Private transcription Subtitle exports Whisper models
120-590
tiny/base/small
MB model
99
languages
supported
TXT/SRT
+ WebVTT
exports
WebGPU
+WASM fallback
GPU+CPU

STT works best on desktop

Speech recognition uses WebGPU/WASM. Desktop Chrome or Edge gives the most reliable result.

Sponsored

Ads help keep OfflineTTS free to use.

About Whisper STT

Whisper is OpenAI's speech recognition model running directly in your browser. Its multilingual model supports 99 languages, and this tool provides a selector for commonly used languages. Use Fast timing for quick transcripts, or select Precise subtitles mode for word-level timestamps and synchronized word highlighting during playback.

Choose from three model sizes: Tiny (~120MB), Base (~210MB), or Small (~590MB). Larger downloads can improve recognition on some recordings but require more memory and time. The integration can use WebGPU where supported or WebAssembly as an alternative. Audio decoding, waveform analysis, and transcription run away from the main interface so it stays responsive.

The selected recording is processed in the browser rather than uploaded to OfflineTTS for recognition. Review the result with the interactive waveform player, seek by clicking the transcript, and export plain text, SRT subtitles, or WebVTT captions. Model delivery and ordinary website requests are covered by the Privacy Policy.

Try our TTS tool: Kokoro TTS (54 voices · Best quality) · Kitten TTS (8 voices · Lightest) · Piper TTS (25 voices · Fastest CPU) · Supertonic TTS (31 languages · Local)

What Whisper Does Not Decide

A transcript is a machine draft

Whisper does not verify names, numbers, quotations, legal meaning, medical terminology, or facts in the recording. This interface also does not perform speaker identification. Overlapping speakers, music, accents, compression, and low recording levels can change both words and punctuation. Compare the result with the waveform and source before using it as captions, minutes, research evidence, or an accessibility deliverable.

Timing modes are estimates

Fast mode provides segment timing. Precise subtitles mode downloads a separate alignment model and can provide word-level timing for synchronized highlighting, but “precise” does not mean frame-perfect. Review cue starts, ends, reading speed, line breaks, and meaningful sounds in a subtitle editor. You are responsible for consent to record, transcribe, retain, and distribute the source media.

Getting Started with Whisper STT

1

Choose Model Size

Tiny (~120MB) minimizes the download, Base (~210MB) is a middle option, and Small (~590MB) uses more storage and memory. Compare them on a representative clip; download size is approximate.

2

Upload or Record Audio

Upload an audio file (WAV, MP3, WebM, etc.) or record directly in the browser. The tool decodes audio in a background worker for smooth performance.

3

Transcribe

Choose Fast timing for a segment-level transcript or Precise subtitles mode for model-aligned word timing. Streaming output shows draft segments while transcription progresses.

4

Export Results

Play the audio against the synchronized transcript, seek from the waveform or text, then download plain text, readable SRT subtitles, or WebVTT captions.

Tips for Accurate Transcription

1

Use clear audio. Low background noise and clear speech produce the best results. If possible, use a good microphone and record in a quiet environment.

2

Compare model sizes. Tiny is useful for quick drafts; larger models can improve some recordings while using more memory and time. Test the same clip and review the output rather than assuming one model is always most accurate.

3

Select the spoken language. Matching the language selector to the recording helps Whisper decode names, punctuation, and multilingual speech more consistently.

4

Try WebGPU where supported. A compatible GPU can reduce processing time, but results depend on browser, hardware, drivers, model, and clip length. The tool selects an available backend; include that backend in a bug report.