Transcribe video to text
Get the words out of a video: a plain transcript to paste somewhere, or ready-made SRT and VTT subtitles with timings. Recognition uses Whisper and handles 99 languages, including accented and mixed speech.
Converted on our servers, deleted right after
How to use
- 1
Choose the output — plain text, or SRT/VTT subtitles with timecodes.
- 2
Drop in a video or audio file. Leave the language on Detect unless the recording is short or noisy.
- 3
Read the transcript on the page, copy it, or download the file.
FAQ
Which languages are recognised?
Whisper covers 99 languages and detects the spoken one automatically. English, Russian, Spanish, Portuguese, German and French are the strongest; rarer languages are noticeably weaker.
How accurate is it?
On clear speech, close to a careful human typist. Accuracy drops with background noise, several people talking over each other, heavy accents and low-bitrate audio. Always read it through before publishing.
How long does it take?
Roughly as long as the recording: a 10-minute video takes about 10 minutes. You can close the tab — the result waits in your account, or on this page if you keep it open.
Why does this one upload my file?
Speech recognition needs a model far too large to run in a browser tab. The media is processed on our servers and deleted the moment the transcript is ready.
Is there a limit?
One transcript a day without an account, up to 10 minutes of media and 500 MB. Premium raises it to 45 minutes per file with no daily limit.
What is the difference between SRT and VTT?
Both are subtitle files with timings. SRT is understood by nearly every player and by YouTube; VTT is the format HTML5 video uses on the web. The text is identical.