💬Audio to Text
Transcribe speech to text + SRT captions. Free, private, no signup — files never leave your browser.
First run downloads the AI speech model once (~40MB), then it’s fast. 100% private — on-device.
Spoken words to text, with timestamps.
On-device AI listens to your audio or video and writes out every word — with times you can turn into captions. Nothing is uploaded.
How to use Audio to Text
Turn spoken words into written text. Drop an audio or video file and on-device AI writes out what is said, with timestamps. Download the transcript as text or as SRT captions for your videos. Nothing is uploaded.
- 1Drop an audio or video file (MP3, WAV, M4A, MP4).
- 2Click Transcribe and let the AI listen (a minute or two).
- 3Copy the text or download it as .txt / .srt captions.
Use cases
What audio to text does
- ✓Word-accurate transcription with sentence timestamps, ready for captions.
- ✓One-click .txt and .srt downloads, plus copy-to-clipboard.
- ✓Language is auto-detected across 99 languages, including Hindi and English.
- ✓Model downloads once (~40MB) and is cached — transcription runs fully offline after that.
- ✓Private by construction: your recordings never leave the device.
Specs
- Speed
- Roughly real-time on short clips, on-device
- Input
- MP3, WAV, M4A, MP4
- Output
- Plain text, .txt, timestamped .srt
- Privacy
- On-device AI, zero uploads
- Cost
- Free, unlimited, no signup
Frequently asked questions
Is my audio uploaded anywhere?
No. The Whisper speech model runs fully in your browser. Your recording never leaves your device.
Which languages are supported?
Dozens, including English, Hindi, Spanish, French and German — the model auto-detects the language.
What do I get at the end?
Plain transcript text plus timestamped segments you can download as .txt or .srt caption files.
Why does the first run take a while?
The AI model (~40MB) downloads once and is cached. Transcription itself runs locally and takes roughly the length of short clips.