Drop a recording or a video, or dictate live. AI running in your browser writes out every word with timestamps, ready to copy or save as subtitles. Nothing is uploaded.
This is local AI. The same approach runs private assistants, search, and document processing on your own hardware, with no data sent to an AI provider.
Local AI for your businessHow it works
A file, a video, or your microphone. The language is detected for you.
The transcript appears as it's made. The first run downloads the model once; after that it starts instantly.
Copy the text, or save it as TXT, SRT, or VTT subtitles. Click any timestamp to hear that moment.
Why use it
Confidential interviews, medical notes, legal calls: the audio is processed on your device and never uploaded.
Whisper recognizes the language by itself, from English and Spanish to Swahili, Yoruba, and Hausa.
Every line is timestamped. Download SRT or VTT and add captions to a video in any editor or on YouTube.
Click a timestamp to play that exact moment, so fixing a name or a number takes seconds.
Dictate an email, a memo, or a draft, and watch the words appear as you speak.
Once the model is downloaded, your browser keeps it, so transcription works even without internet.
Audio to Text runs OpenAI's open Whisper speech model inside your browser. Your audio and its transcript never leave your device, so there's nothing for anyone to keep, leak, or train on. The model itself downloads once, from this site. We count anonymous usage (file name, type, and size, the language heard, browser, and IP) to keep the tool improving.
Read exactly what we recordFAQ
Drop it above. The words are written out with timestamps; copy the text or download it as TXT, SRT, or VTT.
No. The speech model runs in your browser, so the recording never leaves your device. Only the model itself is downloaded, once, from this site.
Your browser downloads the speech model the first time (about 80 to 200 MB, depending on your device). It's saved, so every run after that starts straight away.
It uses Whisper, the open model behind many paid transcription services, sized to run well in a browser. Clear speech comes out very accurately; check names and numbers by clicking their timestamps.
Yes, from your microphone: words appear as you speak. For calls, webinars, and classes, use the free AI Note Taker, which listens to your computer's audio and writes notes instead of a wall of text.
Use the AI Note Taker. It hears the call from your computer (no bot joins), writes the key points, decisions, and to-dos as it goes, and gives you a PDF at the end.
Yes. Drop the video, then download SRT or VTT. Both work in YouTube, Premiere, DaVinci Resolve, CapCut, and VLC.
99 languages, detected automatically. Accuracy is best for widely spoken languages like English, Spanish, French, and German.
On computers and phones with a modern graphics chip it uses the GPU, which is fast. Otherwise it runs on the processor, using every core, which is slower but works everywhere.
More free tools
Passport, visa, and upload photos
Photo Resizer
Pick where the photo is going (passport, visa, or a size limit) and get a file that passes, first time.
OpenStrip hidden data from files
Metadata Remover
Remove GPS location, camera details, and author names from photos, PDFs, and documents.
OpenSee what a file reveals
Metadata Viewer
See where a photo was taken, which device made it, who wrote a document, and whether it says it's AI.
OpenWant tools like this in your own product?
Eligapris builds fast, private web apps and on-device AI. The same engineering behind Audio to Text, for your team.