Translate Video to Text and Subtitles

Upload a video in any language. Get the transcript translated, and an SRT that still matches the original timing.

5 minutes freeNo accountTimestamps preservedSubtitles, not dubbing
Translate into

Drop your file here

or click to browse

Audio: MP3 · WAV · M4A · FLAC · OGG · AAC · AIFF · WMA · AMR · OPUS

Video: MP4 · MOV · AVI · MKV · WebM · FLV · WMV

Up to 200 MB · 5 min free preview

Up to 200 MB · first 5 minutes translated free

Free preview — the first five minutes of your video, transcribed and translated, no account. Output is text and subtitles, never a dubbed voice track. URL paste works for YouTube, TikTok, Instagram, Google Drive and direct links to a media file; for Vimeo or Loom, download the file and upload it.

VexaScribe translates what is said in a video into another language, and gives it back as text and subtitles. Upload an MP4, MOV, MKV, WebM, AVI or M4V — or paste a YouTube, TikTok, Instagram or Google Drive link — and you get two transcripts, the original and the translation, exportable as SRT, VTT, TXT, DOCX or JSON.

It runs in two stages. Speech is transcribed in the source language with Whisper Large-v3 Turbo, which covers 99 languages, and that transcript is then translated cue by cue. The part that matters for video is what happens to the timing: every cue keeps its original start and end times, so the translated SRT drops straight onto the untouched video and lands in sync in YouTube Studio, Premiere, DaVinci, CapCut or an HTML5 player. Target languages are whatever the translation engine supports — around 133 on the free engine, and no fixed list on premium.

The output is text and subtitles, never a dubbed voice track — the speaker's own audio is left alone. Translation is included rather than billed per word, so you pay for minutes of audio and not for how much was said in them. The preview above translates the first five minutes of one video without an account; signing up raises that to 30 minutes free with no card, and paid plans start at $2 a month.

Subtitles or dubbing — which one do you need?

These are different products, and “video translator” is used for both. Subtitling translates what was said into readable text laid over the original audio. Dubbing replaces the audio with a synthetic voice in the new language. This page does subtitling and transcripts. If you need a dubbed voice track, the note below tells you where to go instead.

If you need…Subtitles (this page)Dubbing (elsewhere)
The speaker's real voice preservedYes — the original audio is untouched. Usually required for interviews, research, documentary and anything evidential.No — the voice is replaced.
Viewers who will not read subtitlesPoor fit — young or low-literacy audiences, or anyone watching while doing something else.Better fit.
A searchable, quotable text recordYes — the transcript is the deliverable, not a by-product.Not usually part of the output.
To review it before publishingStraightforward — an SRT is a text file you can edit anywhere.Harder — changing a line means regenerating audio.
To satisfy a captions requirementSubtitles are the format most accessibility rules ask for.Does not meet a captions requirement on its own.

If dubbing is what you actually want

We do not do it, and there is no setting here that will. AI dubbing with voice cloning is a separate product category — HeyGen and Synthesia both build it as their core product and are the obvious places to start. The two pair up in practice: a translated transcript from here is a usable input script for a dubbing tool, which is a reason to run the free preview even if you end up dubbing elsewhere.

Where this fits in your workflow

This is a step in a pipeline rather than a destination. Depending on what you are making, it sits in one of four places.

Before the edit — so you can cut on meaning

Run the raw footage through first and edit against a transcript you can actually read. If you are cutting an interview in a language you do not speak, this is the difference between editing on waveforms and editing on what was said. Speaker labels come through on both the source and the translation, so you can see who said what before a single cut is made.

After the edit — subtitles for the locked cut

The more common case. Export the finished video, get the translated SRT back, attach it in YouTube Studio or drop it on the timeline in Premiere, DaVinci or CapCut. Because the cue timings are the original ones, nothing needs re-syncing. Do this after picture lock rather than before — if the cut changes, the timings move with it.

As the script for a dubbing tool

If you have decided you want a dubbed voice track, a dubbing tool still needs to know what to say. A translated transcript with timings is exactly that input, and you can fix a mistranslated name here — where it is a text edit — rather than after it has been spoken in a synthetic voice.

Releasing in several languages at once

Transcribe the video once, then translate that same transcript into each language you need. You are charged for the transcription a single time no matter how many languages follow. Each language is its own translation rather than one bulk operation, so add them as you need them instead of committing up front.

What you get back

Two transcripts — source language and translation — each exportable as plain text, DOCX, SRT, VTT or JSON. The part that matters for video is that the translated SRT keeps the original cue numbers and timestamps, so it drops onto the untouched video and lands in sync.

Source — Spanish
2
00:00:04,120 --> 00:00:07,480
Entonces empezamos el proyecto
en marzo del año pasado.
Translated — English
2
00:00:04,120 --> 00:00:07,480
So we started the project
in March of last year.

Same cue number, same in and out times to the millisecond. Only the text changes. One consequence worth knowing: translated text is rarely the same length as the source, so a cue that sat comfortably in Spanish can be tight in English or leave a gap. For anything you are publishing, read the SRT once before you upload it.

How it works

  1. Upload the video or paste a link. MP4, MOV, MKV, WebM, AVI and M4V up to 5 GB. Links work for YouTube, TikTok, Instagram and Google Drive, or any direct link to a media file.
  2. Pick the target language. The source is auto-detected — override it if the audio opens in one language and continues in another.
  3. Export. Translated SRT or VTT for the video, TXT or DOCX for reading, JSON if you are building something on top of it.

What happens to a large file

A multi-gigabyte upload is not one fragile request. The file is split into 5 MB pieces and sent ten at a time, each retried on its own if the connection drops. A wobbly connection costs you one piece rather than the whole upload.

Who said what

Speakers are separated automatically and labelled through the transcript. If you already know how many people are in the recording, saying so up front helps. Those labels carry into the translated subtitles too, not just the source transcript.

Fixing it before you export

Brand names and unusual surnames are where machine transcription most often slips. Edit the text and export again — the file is not re-transcribed and nothing is charged a second time.

What you can take away

Both languages, in five formats: SRT and VTT for subtitles, TXT and DOCX for reading, and JSON if you are feeding something downstream. You can also download the extracted audio on its own.

How long it takes

Median 3–5 minutes per hour of video on the standard tier and 1–2 minutes on premium, measured across 7,432 production jobs. Translation adds roughly a minute per 5,000 characters of transcript. This is the one performance figure we publish, because it is the one we actually measure — and it describes speed, not correctness.

Formats, links and limits

Video files

MP4, MOV, MKV, WebM, AVI and M4V. Audio is extracted from the container automatically, so there is nothing to convert first. Up to 5 GB per file on an account; the no-account preview caps at 200 MB.

Pasting a link

YouTube, TikTok, Instagram and Google Drive work, and so does a direct link to a media file — a .mp4 on your own server, or a signed storage URL. Vimeo and Loom page links do not, because those pages are not direct media links; download the file and upload it instead. Private and age-restricted videos cannot be fetched from a link either.

The free preview, stated plainly

Without an account you get the first five minutes of one video, once every seven days. For most videos that is a sample rather than a finished job — it exists so you can see the output on your own footage before deciding anything. An account raises it to 30 minutes free, and paid plans start at $2.

What one video actually costs

Transcription and translation draw on the same pool of minutes. A 60-minute video runs to roughly 9,000 words: 60 minutes to transcribe, plus 11 more if you translate on premium — call it 71 minutes for the finished bilingual set. Free translation adds nothing.

PlanMinutesHour-long videos a monthEach
Starter — $22002$1.00
Basic — $51,00014$0.36
Pro — $102,50035$0.29
Studio — $206,00084$0.24

Translating a video into a second language does not re-charge you for the transcription, and a language you have already generated on the premium engine stays generated — coming back for that version later costs nothing.

A more realistic month

Eight ten-minute videos, each released in two languages. That is 80 minutes of transcription plus roughly 15 minutes of premium translation — about 95 minutes, comfortably inside the $5 plan with room left over. The second language costs only its own translation, because the transcription is already done and paid for.

What the alternative costs

A professional translator working from the same transcript, at ten cents a word — the low end of the range — would charge somewhere near $900 for a 60-minute video. That is not an argument that the two are interchangeable; a human catches things neither engine will. It is an argument about where the review budget goes: machine-translate first, then pay a person to check the parts that matter.

Why translate here rather than somewhere else

Most tools that will translate a video are video editors with translation attached. We are a transcription product with translation attached. That difference shows up in four practical ways — and in one where it goes against us, which is at the bottom.

Translation is bundled, not metered

You are charged for minutes of audio, not for words of output. A dense, fast-talking forty-minute video costs the same to translate as a sparse one of the same length. Tools that bill per character get more expensive precisely when your video has more in it, which is the wrong way round.

Two engines, and the choice is yours per file

The free engine is fine for understanding what was said. The premium one is for when the wording will be published — it keeps proper names as names instead of translating them into words, and turns idioms into natural equivalents. You pick per file, so a rough cut and a final deliverable do not have to cost the same.

Premium translations are stored, not re-billed

Translate a file into Spanish on the premium engine and that result is kept against the file. Coming back for it next month deducts nothing. It means iterating on a video — checking the Spanish, then adding German — gets cheaper rather than repeatedly charging you. The free engine re-runs each time, so this particular benefit is a premium one.

You can fix the text before you export

A brand name or an unusual surname will occasionally come out wrong. Correct it in the transcript and re-export — the file is not re-transcribed and you are not charged again. On a tool that renders subtitles into the video, the equivalent fix means regenerating the video.

Where a video editor is the better choice

We hand back a subtitle file. We cannot burn subtitles into the picture, style them, or give you a finished video with the captions baked in — there is no rendering step in our pipeline at all. If what you want is a ready-to-post video with styled captions on screen, an in-browser editor like VEED or Kapwing is built for that and we are not. Use us when you want the file; use them when you want the render.

How good is it, honestly

We do not publish an accuracy percentage for video translation, and the reason is worth stating rather than hiding. The benchmark usually quoted is OpenAI's Whisper paper (arXiv:2212.04356, Table 13), which reports word error rates on the FLEURS test set — 4.2% English, 3.0% Spanish, 4.5% German, 8.3% French, 14.7% Mandarin. Those figures are real, and they do not describe what you will get here.

  • •Different model. That table benchmarks Whisper Large-v2. Our standard tier runs Large-v3 Turbo, and premium runs a different engine again.
  • •Different audio. FLEURS is clean read speech. Video is room tone, music beds, overlapping speakers and phone microphones.
  • •Transcription is half the job. WER measures the source transcript. Translation is a second stage with its own failure modes, and the two do not simply add up.

The limit worth knowing before you publish

Both translation engines work one cue at a time. Neither builds a glossary and carries it through the file, so a term translated one way at minute two is not guaranteed to be translated identically at minute forty. Premium is more likely to get each line right — it keeps proper names as names rather than translating them into words, and turns idioms into natural equivalents — but it is not reading the surrounding conversation while it does. For anything going out publicly, check recurring names and technical terms across the whole file.

What this is not good at

Five things worth knowing before you upload, because each of them would be annoying to discover afterwards.

  • ●No burned-in subtitles. You get a subtitle file, not a finished video with captions on screen. There is no rendering step in our pipeline at all. For attaching the file, see how to add subtitles to a video.
  • ●Not live. Finished files only — there is no streaming or real-time captioning mode.
  • ●Music beds and crosstalk hurt twice. Anything that degrades the transcript degrades the translation built on top of it, so a loud score or three people talking over each other costs you more than it would on a transcription-only job.
  • ●Unmeasured beyond the common languages. The premium engine has no language list and will attempt whatever you name. For languages Google Translate does not list, it will still produce output — but we have not measured accuracy there and will not claim it.
  • ●Premium transcription covers fewer languages, not more. Standard handles 99, premium 55. Premium is the better engine on the languages it supports, but if yours is outside that set the standard tier is the one that can do the job.

Common language pairs

Any of the 99 languages the standard tier covers can be the source. The target list is whatever the translation engine supports — around 133 on the free engine, and no fixed list on premium, which will attempt languages the free one does not cover at all. Four pairs have their own pages with detail specific to that language:

Those pages are written around audio files, but the workflow and the engines are identical — a video in that language behaves the same way once the audio is extracted.

Putting translated subtitles on a YouTube video

Paste the video URL, pick the target language, export the SRT. In YouTube Studio open the video, go to Subtitles, add the language and upload the file. It appears as a selectable subtitle track alongside the original audio.

This is clearly worth doing when a video has no captions at all, or when the auto-captions are poor enough that YouTube's own translate-captions option is translating errors — transcribing the audio directly avoids inheriting those mistakes. Where the auto-captions are already clean, the gain is smaller and mostly about controlling the file yourself.

Working with someone else's video, or just want the words? The YouTube transcript tool pulls a transcript without the translation step. For the mechanics of attaching subtitles in Premiere, DaVinci, CapCut or on a phone, see how to add subtitles to a video.

Frequently asked questions

Does this dub the video or add subtitles?

Subtitles and text. You get the transcript translated into your target language plus a timestamped SRT or VTT file that plays over the original audio — the speaker's own voice is untouched. It does not generate a synthetic voice track in the new language. If you need dubbing, that is a different product category; HeyGen and Synthesia build it as their core product. The two pair up in practice, because a translated transcript is a usable input script for a dubbing tool.

Does the translated SRT still match the original video timing?

Yes. Each cue keeps its original number and its in and out timestamps to the millisecond; only the text inside changes. So the file drops onto the untouched video and lands in sync. One thing to watch: translated text is rarely the same length as the source, so a cue that sat comfortably in one language can be tight in another or leave a gap. For anything you are publishing, read the SRT once before uploading it.

Which video formats and links work?

MP4, MOV, MKV, WebM, AVI and M4V, up to 5 GB per file on an account. Audio is extracted from the container automatically, so there is nothing to convert first. For links, YouTube, TikTok, Instagram and Google Drive work, and so does a direct link to a media file such as an .mp4 on your own server. Vimeo and Loom page links do not, because those pages are not direct media links — download the file and upload it instead. Private and age-restricted videos cannot be fetched from a link either.

How much of my video can I translate for free?

Without an account, the first five minutes of one video, once every seven days. For most videos that is a sample rather than a finished job — it exists so you can judge the output on your own footage before deciding anything. Creating an account raises it to 30 minutes free with no card. Paid plans start at $2 a month for 200 minutes.

What does it cost to translate a full video?

Transcription and translation draw on the same pool of minutes. A 60-minute video is roughly 9,000 words: 60 minutes to transcribe, plus about 11 more if you translate on the premium engine — around 71 minutes for the finished bilingual set. On the $5 plan that works out near $0.36 per hour-long video. Free translation adds nothing on top. Translating the same file into a second language does not re-charge you for the transcription, and a language you have already generated stays generated.

How long does it take?

Median 3 to 5 minutes per hour of video on the standard tier, and 1 to 2 minutes on premium, measured across 7,432 production jobs. Translation adds roughly a minute per 5,000 characters of transcript. That is the one performance figure we publish because it is the one we actually measure, and it describes speed rather than correctness.

How accurate is it?

We do not publish an accuracy percentage for video translation. The benchmark usually quoted is OpenAI's Whisper paper (arXiv:2212.04356, Table 13), which reports word error rates on the FLEURS test set — 4.2% English, 3.0% Spanish, 4.5% German, 8.3% French, 14.7% Mandarin. Those figures are real but they do not describe this pipeline: they benchmark Whisper Large-v2 on clean read speech, our standard tier runs Large-v3 Turbo, and translation is a second stage with its own failure modes. Also worth knowing before you publish: both translation engines work one cue at a time and neither carries a glossary across the file, so recurring names and technical terms are worth checking across the whole transcript.

Can I put the translated subtitles back on YouTube?

Yes. Export the SRT, then in YouTube Studio open the video, go to Subtitles, add the language and upload the file. It appears as a selectable subtitle track alongside the original. This is most clearly worth doing when a video has no captions at all, or when the auto-captions are poor enough that YouTube's own translate-captions option is translating errors — transcribing the audio directly avoids inheriting those mistakes.

Methodology & disclosure

Where the numbers come from. The word error rates in the accuracy section are from OpenAI's Whisper paper (arXiv:2212.04356, Table 13, FLEURS test set) and benchmark Whisper Large-v2 — a different model from the one running here, which is why they are presented as context rather than as our accuracy. The processing-speed figures are our own: median times across 7,432 production jobs. We do not maintain independent translation-quality benchmarks and do not publish an accuracy percentage for translation.

Engines. Transcription and translation are two separate choices, and their language coverage is not the same. For transcription, the standard tier runs Whisper Large-v3 Turbo across 99 languages; the premium tier is opt-in, costs two minutes of quota per minute of audio, and covers 55. Speaker labels use pyannote.audio. For translation, the free engine is Google Translate and the premium engine sends each cue through a language model with no fixed language list. Engine choice may change as newer models release. Not HIPAA-eligible — do not upload PHI.

Disclosure. VexaScribe is a hosted transcription and translation service, and this page recommends it for translated-subtitle work. It is not the right tool for AI voice dubbing, which we do not offer; HeyGen and Synthesia are named in that section because dubbing is their core product, and we have no commercial relationship with either. Product limits and pricing on this page were verified against our own configuration on 24 September 2026.

Related tools

Translate your first video free

The preview at the top of this page translates the first five minutes of any video without an account. An account raises that to 30 minutes free, no card.