How to transcribe audio in 3 steps
- Choose a file
Drop it on the box above or click to browse. You can also record your voice with the microphone. - Click “Transcribe”
Pick the audio language and quality. The AI model downloads the first time and then stays cached in your browser. - Review and download
Fix anything directly in the text, then download it as TXT or as SRT/VTT subtitles.
A free alternative to paid transcription services
Most transcription websites let you try a few minutes and then ask for a monthly subscription, or cap how many files you can upload per day. That makes sense for them: every minute they transcribe on their servers costs them money.
Herramientas Libres works differently. The speech recognition model runs inside your own browser, using your computer or phone. Since we don't pay for servers per minute, we can keep it free and unlimited. The site is supported by ads.
| Feature | Herramientas Libres | Typical paid services |
|---|---|---|
| Price | Free | $8–20 per month, or pay per minute |
| Minute or file limits | None | A few minutes or files per day on free plans |
| Sign-up | Not needed | Almost always required |
| Where is your audio processed? | On your device | Uploaded to their servers |
| SRT/VTT subtitles | Included | Often paid plans only |
| Speaker identification | No (not yet) | Yes, in some |
What you can transcribe
- Voice notesVoice messages from chat apps (.opus, .ogg, .m4a) work like any other file.
- Interviews and podcastsGet the text with timestamps so you can find every quote.
- Lectures and talksTurn a class recording into searchable notes.
- MeetingsKeep written minutes without sending confidential conversations to third parties.
- VideosUpload MP4, MOV or WEBM directly; the audio is extracted in your browser.
- DictationClick “Record with microphone”, speak and get the written text.
Supported formats
Audio: MP3, WAV, M4A, AAC, OGG, OPUS, FLAC and WEBM. Video: MP4, MOV, WEBM and, in most browsers, MKV. The tool uses your browser's own decoder, so exact support depends on it: if a file won't open, try another browser or convert it to MP3.
Which quality should I choose?
- Fast (≈40 MB): great for phones and clear audio with one speaker.
- Balanced (≈80 MB): recommended on computers. Good accuracy and speed.
- Most accurate (≈250 MB): for difficult audio (noise, accents, technical terms). Slower and needs more memory.
Each model downloads only once. If your browser supports WebGPU (recent Chrome and Edge do), transcription runs on your graphics card and is much faster.
Real privacy
Many sites promise privacy but still upload your audio. Here there is no server to receive it: the page downloads the public speech recognition model and your audio is processed in your browser's memory. You can check it yourself: once the model is loaded, disconnect from the internet and transcription keeps working. See our privacy policy.
Questions? Read the FAQ.