Skip to content
Browser STT workspace

Whisper Speech to Text

Upload audio or video, record from your microphone, or load a direct media URL. Transcribe privately in your browser and export TXT, SRT, or VTT.

Private transcription Subtitle exports Whisper models
120-590
tiny/base/small
MB model
99
languages
supported
TXT/SRT
+ WebVTT
exports
WebGPU
+WASM fallback
GPU+CPU

STT works best on desktop

Speech recognition uses WebGPU/WASM. Desktop Chrome or Edge gives the most reliable result.

About Whisper STT

Whisper is OpenAI's speech recognition model running directly in your browser. Its multilingual model supports 99 languages, and this tool provides a selector for commonly used languages. Use Fast timing for quick transcripts, or select Precise subtitles mode for word-level timestamps and synchronized word highlighting during playback.

Choose from three model sizes: Tiny (~120MB), Base (~210MB), or Small (~590MB). Larger downloads can improve recognition on some recordings but require more memory and time. The integration can use WebGPU where supported or WebAssembly as an alternative. Audio decoding, waveform analysis, and transcription run away from the main interface so it stays responsive.

The selected recording is processed in the browser rather than uploaded to OfflineTTS for recognition. Review the result with the interactive waveform player, seek by clicking the transcript, and export plain text, SRT subtitles, or WebVTT captions. Model delivery and ordinary website requests are covered by the Privacy Policy.

Try our TTS tool: Kokoro TTS (54 voices Β· Best quality) Β· Kitten TTS (8 voices Β· Lightest) Β· Piper TTS (25 voices Β· Fastest CPU) Β· Supertonic TTS (31 languages Β· Local) Β· Pocket TTS (8 voices Β· Voice cloning)

What Whisper Does Not Decide

A transcript is a machine draft

Whisper does not verify names, numbers, quotations, legal meaning, medical terminology, or facts in the recording. This interface also does not perform speaker identification. Overlapping speakers, music, accents, compression, and low recording levels can change both words and punctuation. Compare the result with the waveform and source before using it as captions, minutes, research evidence, or an accessibility deliverable.

Timing modes are estimates

Fast mode provides segment timing. Precise subtitles mode downloads a separate alignment model and can provide word-level timing for synchronized highlighting, but β€œprecise” does not mean frame-perfect. Review cue starts, ends, reading speed, line breaks, and meaningful sounds in a subtitle editor. You are responsible for consent to record, transcribe, retain, and distribute the source media.

Getting Started with Whisper STT

1

Choose Model Size

Tiny (~120MB) minimizes the download, Base (~210MB) is a middle option, and Small (~590MB) uses more storage and memory. Compare them on a representative clip; download size is approximate.

2

Upload or Record Audio

Upload an audio file (WAV, MP3, WebM, etc.) or record directly in the browser. The tool decodes audio in a background worker for smooth performance.

3

Transcribe

Choose Fast timing for a segment-level transcript or Precise subtitles mode for model-aligned word timing. Streaming output shows draft segments while transcription progresses.

4

Export Results

Play the audio against the synchronized transcript, seek from the waveform or text, then download plain text, readable SRT subtitles, or WebVTT captions.

Tips for Accurate Transcription

1

Use clear audio. Low background noise and clear speech produce the best results. If possible, use a good microphone and record in a quiet environment.

2

Compare model sizes. Tiny is useful for quick drafts; larger models can improve some recordings while using more memory and time. Test the same clip and review the output rather than assuming one model is always most accurate.

3

Select the spoken language. Matching the language selector to the recording helps Whisper decode names, punctuation, and multilingual speech more consistently.

4

Try WebGPU where supported. A compatible GPU can reduce processing time, but results depend on browser, hardware, drivers, model, and clip length. The tool selects an available backend; include that backend in a bug report.

Whisper STT Architecture & Subtitling Guide

OfflineTTS Whisper STT brings OpenAI's foundational speech recognition model directly into client-side JavaScript. By processing sensitive audio files entirely inside your device's memory, medical, legal, corporate, and private personal recordings remain confidential.

Local Audio Decoding

Audio files are converted to single-channel 16,000Hz floating-point arrays using Web Audio API decoders in a detached Web Worker thread. This ensures the main user interface never stutters, even when processing 30-minute podcast recordings.

SubRip & WebVTT Precision

Word-aligned timestamps produce clean subtitle cues. You can click any word in the generated transcript to immediately seek the waveform player to that exact audio position for instantaneous manual proofreading and correction.

Frequently Asked Questions about Whisper Speech to Text

How does Whisper STT transcribe audio locally in the browser?

OfflineTTS Whisper Speech-to-Text uses OpenAI Whisper models ported to ONNX Runtime Web. When you upload or record audio, a background Web Worker decodes the audio into standardized 16kHz PCM arrays and feeds chunks into the neural model directly on your device via WebGPU or WebAssembly.

Is any audio or transcript data sent to your servers?

No. Unlike cloud transcription services (such as Google Speech-to-Text or Rev), your audio file and transcribed text remain 100% inside your browser session. No media files or generated transcripts are uploaded to any server.

What audio and video file formats can I upload for transcription?

You can upload common audio formats including MP3, WAV, M4A, AAC, FLAC, OGG, and WebM, as well as MP4 and WebM video files. The browser's native AudioContext and Web Audio APIs extract and decode the audio track automatically.

What is the difference between Fast mode and Precise subtitles mode?

Fast mode transcribes audio in standard sentence-level segments for rapid reading. Precise subtitles mode activates an additional alignment model that calculates exact word-level start and end timestamps, making it ideal for synchronized subtitles and karaoke-style captioning.

Can I export formatted subtitles for YouTube and video editors?

Yes. You can export your transcribed speech as plain text (TXT), SubRip Subtitles (SRT with sequential numbering and millisecond timecodes), or WebVTT (VTT with cue timings). These subtitle files can be directly imported into YouTube Studio, Adobe Premiere, DaVinci Resolve, or Final Cut Pro.

Which Whisper model size should I choose (Tiny, Base, or Small)?

Whisper Tiny (~120MB) downloads rapidly and is great for clear English speech on low-power devices. Whisper Base (~210MB) balances download speed with higher multilingual vocabulary accuracy. Whisper Small (~590MB) delivers state-of-the-art accuracy with superior resilience against background noise and accents.