Translate the transcript to 130+ languages
Publish your video internationally in one click, timestamps preserved
Upload MP4, MOV, WEBM, MKV, or AVI. Get a broadcast-quality transcript with speaker labels and timestamps in 99 languages. First 5 minutes free, no signup.
Video-to-text turns the spoken audio in a video into a written transcript using AI speech recognition. VexaScribe accepts MP4, MOV, WEBM, MKV, and AVI directly — no ffmpeg step, no MP3 conversion. Get a transcript with speaker labels, timestamps, and 99-language auto-detection. Export as TXT, DOCX, SRT, VTT, or JSON.
Drop your file here
or click to browse
Audio: MP3 · WAV · M4A · FLAC · OGG · AAC · AIFF · WMA · AMR · OPUS
Video: MP4 · MOV · AVI · MKV · WebM · FLV · WMV
Up to 200 MB · 5 min free preview
Publish your video internationally in one click, timestamps preserved
Get the 5 key moments — skip the full replay
Ask "when did they announce the release date?" and get the timestamp
TXT · DOCX · SRT · VTT · JSON — one processing
A video transcript generator is a tool that takes a video file (or URL) and produces a written transcript automatically, without a human transcriber typing it out. Modern AI-based generators like VexaScribe use OpenAI Whisper Large-v3 or Speechmatics Melia-1 under the hood, hit 92-97% accuracy on clean English audio, and process a 60-minute video in 5-10 minutes. The tool at the top of this page is a video transcript generator with a 5-minute free preview.
| Aspect | AI generator | Manual transcription |
|---|---|---|
| Time per hour of video | 5-10 minutes processing | 4-6 hours of typing |
| Cost per hour of video | $0.24-$3.00 (or $2/mo for 200 minutes) | $90-$150 (professional service) |
| Accuracy on clean audio | 92-97% | 99%+ (with review) |
| Best for | Publishing, subtitles, research, show notes | Legal filings, ADA-compliant broadcast captions |
For most publishing and content workflows the AI generator is the correct pick — the accuracy gap is small enough that a 5-10 minute editing pass closes it. Human transcription only wins when a stamp of certification is legally required.
Every format below works with the tool at the top of this page. Video resolution has no effect on transcript accuracy — the audio track is what the engine reads. Rule of thumb: if the file plays in VLC or QuickTime, it transcribes.
| Format | Common source |
|---|---|
| .mp4 | The universal default — H.264 or H.265 codec, used by TikTok, Instagram, most cameras, screen recorders, and iPhone. MP4-specific guide → |
| .mov | Apple QuickTime container — older iPhones, Final Cut Pro / DaVinci Resolve ProRes exports. |
| .webm | Web-native (VP9/VP8 + Opus codec) — Chrome and OBS screen recorders, browser captures, Loom exports. |
| .mkv | Matroska container — open format, common for desktop screen recorders and downloaded video archives. |
| .avi | Legacy Microsoft container — older screen recorders, archived footage. |
| .wmv | Windows Media Video — Windows Movie Maker output and older corporate recordings. |
| .flv | Legacy Flash video — older lecture libraries and corporate training archives. |
| .mpg / mpeg | Older camcorders and broadcast archives. |
MP4 is the universal video container — H.264 or H.265 video codec with AAC or Opus audio. TikTok, Instagram, YouTube exports, iPhone recordings (iOS 11+), Android recordings, and most screen recorders all output MP4. Video resolution doesn't affect transcription accuracy — a 4K MP4 with poor mic audio transcribes worse than a 720p MP4 with a lavalier close to the speaker. For MP4-specific workflows including iPhone captures and screen recordings, see our MP4 to text guide.
MOV is Apple's QuickTime container — used by older iPhones (iOS 10 and below), Mac screen captures (Cmd+Shift+5), Final Cut Pro exports, and DaVinci Resolve ProRes output. The underlying codec is often the same H.264 as MP4, just wrapped in a QuickTime header. MOV files transcribe identically to MP4 — no special handling needed on this page's tool. If your MOV file plays in QuickTime, it transcribes.
MKV (Matroska) is the open-source container of choice for OBS Studio screen recordings, high-quality video archives, and downloaded video files with multiple audio tracks. When an MKV has multiple audio tracks (some Zoom exports do), our tool uses the first track by default — re-process on a different track from the editor if the first track is silent or wrong-language. MKV files can hold FLAC or Opus audio, both of which transcribe with no accuracy penalty.
Audio-only files also work: MP3, M4A, WAV, AAC, FLAC, OGG, OPUS. For audio-specific guides see MP3 to text, M4A to text, WAV to text, and OGG to text.
The workflow is the same for a 15-second TikTok and a 3-hour lecture recording.
Drag your MP4, MOV, WEBM, MKV, or AVI file directly into the tool zone at the top of this page — no pre-conversion, no ffmpeg, no audio extraction. Anonymous previews accept files up to 200 MB.
99 languages, auto-detected from the first speech in the audio track. Override manually only if the video opens with music or long silence. Non-Latin scripts (Chinese, Japanese, Arabic, Hindi) and right-to-left languages (Arabic, Hebrew) render cleanly in every export format.
Automatic speaker diarization labels up to 8 distinct voices. Spot-check names, companies, and product SKUs — proper nouns are what AI gets wrong most (20-30% error rate even on clean audio). The editor's search-and-replace fixes all instances of a name in one pass.
TXT or DOCX for a transcript document. SRT or VTT for subtitle files (YouTube, Premiere Pro, Final Cut, DaVinci Resolve). JSON with word-level timestamps for custom video editors and developer pipelines. The same processed file gives you every format — no re-upload needed.
Every transcript includes segment-level timestamps and automatic speaker labels. Word-level timestamps are available on paid plans. Below is an anonymized excerpt of what you'll see in the editor and in the exported TXT / DOCX.
[00:00:03] Speaker 1: The thing about video transcription [00:00:06] Speaker 1: is that most people don't realize it's just [00:00:09] Speaker 1: the audio track underneath. [00:00:12] Speaker 2: Right — the container format doesn't matter. [00:00:15] Speaker 2: MP4, MOV, MKV are all just wrappers. [00:00:19] Speaker 1: Exactly. The AI only sees the audio. [00:00:22] Speaker 1: A 4K video with bad mic audio [00:00:25] Speaker 1: transcribes worse than a 480p video with a good mic.
Prefer paragraphs without timestamps or speaker labels? Toggle them off before export — the DOCX and TXT outputs can be paragraph-only for blog posts and articles. SRT and VTT always preserve timing (subtitle formats require it). JSON keeps the raw word-level timing array for developers building custom video editors or search interfaces.
Modern AI video transcription achieves 92-97% word accuracy on clean English audio — 3-8% Word Error Rate. Accuracy is governed by audio quality, not video resolution — a 4K video with bad mic audio transcribes worse than a 480p video with good mic audio.
Real-world accuracy varies dramatically by content type. The ranges below are approximate and align with published third-party evaluations on the Hugging Face Open ASR Leaderboard.
| Content type | Accuracy | Editing time |
|---|---|---|
| Studio explainer / single-speaker tutorial (treated room, good mic) | 95-97% | 5-10 min/hr |
| Podcast-style video (two speakers, dedicated mics) | 93-96% | 10-15 min/hr |
| Zoom / Teams / Meet recording (built-in laptop mics) | 91-95% | 10-15 min/hr |
| Lecture / classroom recording (ceiling or wireless mic) | 89-94% | 10-20 min/hr |
| TikTok / Instagram Reel (in-app recording with music underneath) | 82-90% | 15-25 min/hr |
| Vlog (outdoor, handheld, ambient noise) | 80-88% | 20-30 min/hr |
| Phone-recorded interview (compressed audio, near-field) | 85-92% | 15-20 min/hr |
| Multi-language / heavy accents | 82-90% | 15-25 min/hr |
Proper nouns — brand names, product SKUs, technical jargon, foreign names — miss more often than regular vocabulary (20-30% error rate even on clean audio). Full accuracy methodology and per-model comparison on our Whisper accuracy analysis.
The language dropdown pre-selects based on your browser locale. A Turkish speaker on a Turkish browser sees Turkish selected by default; a Portuguese speaker sees Portuguese; a Chinese speaker sees Mandarin. This is deliberate, not cosmetic — it fixes a bug that affects most third-party transcript tools built on Supadata's caption API.
The bug: when a video has multiple auto-generated caption tracks (TikTok and Instagram videos often have three or four — Vietnamese, Indonesian, English, plus the original language), a tool that asks the caption API for “auto” can get a random language back. That's where the notorious “why is my English video's transcript in Spanish” complaint comes from. Pre-selecting the user's actual browser language cuts the miss rate significantly.
All 99 Whisper Large-v3 languages are supported. Non-Latin scripts render cleanly in the editor and every export format. Right-to-left scripts preserve reading direction in DOCX and TXT. Auto-detection doesn't work perfectly on very short clips (under 3 seconds) — for a 5-second Reel, override manually if needed.
For translation, VexaScribe supports 130+ target languages — transcribe in the source language, then translate the transcript in one click from the editor. For dedicated video translation workflows see video translator.
Yes — genuinely free, with two tiers. The anonymous preview (200 MB max, 5-minute transcript, one per week per IP) requires no signup and shows the full transcript with speaker labels and timestamps. The free account (30 minutes/month, email only, no card) unlocks the editor, full-length transcripts, and every export format (TXT, DOCX, SRT, VTT, JSON). Paid plans start at $2/month for 200 minutes.
| Tier | Price | Limits | What's included |
|---|---|---|---|
| Anonymous preview | Free · no signup | 200 MB max file · 5-min preview · 1 per week per IP | Full 5-min transcript with speaker labels + timestamps · TXT/SRT download |
| Free account | Free · email only | 30 minutes/month · no card required | Full transcript · editor · TXT/DOCX/SRT/VTT/JSON export · 130+ language translation |
| Starter | $2/month | 200 minutes/month | Everything above · AI summary · AI Chat · priority processing |
Free alternatives without any signup include YouTube's built-in transcript (English-primary, ~65% accuracy, YouTube-only) and self-hosted OpenAI Whisper (Python + GPU required, free forever). See free transcription options for the honest comparison.
This page's tool is one of five ways to get from video to text. Honest comparison of when each makes sense:
Drop the video into VexaScribe, Otter, HappyScribe, or similar. 5-10 minutes per hour of video, 92-97% accuracy on clean English audio. $0.004-$0.05 per minute at list rates. Best for: publishing, research, subtitles, show notes, meeting notes, and anything that isn't legally binding.
If the video is already on YouTube, click ... under the video → Show transcript. Free, instant, but Google's older ASR engine returns roughly 60-70% accuracy on accented English and less on non-English content. Not a real substitute for AI transcription for publishing use.
Rev.com, Scribie, or GMR — human transcribers deliver 99%+ accuracy at $1.50-$2.50 per audio minute with 12-48 hour turnaround. Only worth the cost when you need certified verbatim for legal filings, ADA-compliant broadcast captions, or other stamp-of-certification use cases.
Install openai-whisper or faster-whisper on a machine with a GPU. Free MIT license, offline, no per-minute cost, complete privacy. Requires Python skills and ~10 GB VRAM for Large-v3. Best for: high volume, privacy-sensitive content, or anyone comfortable with ML tooling.
Play the video and type as you listen. 4-6 hours of typing per hour of video for a skilled transcriber. Only worth it for very short clips, audio too poor for AI to handle, or heavily technical content where you know the domain better than any model.
The tool at the top of this page fits most video-transcription jobs. Four scenarios where a different tool is a better fit:
Best for teams needing shared transcription workflows, per-project libraries, and human-review passes on top of the AI output. Higher per-minute cost than VexaScribe but bundles the collaboration UX. Fits 2+ person editorial teams.
Anonymous 5-minute preview with no signup, then 30 minutes free on signup, then $2/month. Speechmatics Melia-1 primary engine with Whisper Large-v3 fallback. Best for one-off video transcription, creators, and researchers.
Purpose-built for live Zoom, Teams, and Google Meet meetings with a bot participant that joins the call and transcribes in real time. For finished video files that you already downloaded, generic tools beat Otter on price and format flexibility.
Human transcription at $1.50-$2.50 per audio-minute with 12-48 hour turnaround. Only worth the cost when you need certified verbatim for legal filings, ADA-compliant broadcast captions, or other stamp-of-certification use cases. AI-first tools cover everything else at 100-500× lower cost.
For SRT files specifically, our SRT generator is purpose-built with subtitle-specific export. For batch upload of many videos at once, see bulk transcription. For a full 14-model API comparison see our transcription API comparison.
No, ChatGPT does not directly transcribe video files. ChatGPT can accept short audio inputs (under about a minute) via voice mode but does not accept MP4, MOV, or WEBM video uploads. Use a dedicated tool like VexaScribe to extract audio from the video and transcribe in one step — drop the video file, get a full transcript with timestamps and speaker labels. VexaScribe runs OpenAI Whisper Large-v3 (the same model family behind many AI products) and handles the video-to-audio extraction server-side.
Four steps: (1) drop the video file into a transcription tool that accepts video directly (VexaScribe, Otter, HappyScribe, Rev — modern tools take MP4, MOV, WEBM, MKV, AVI without pre-extracting audio); (2) let it auto-detect the language (99 languages supported on Whisper-based tools); (3) spot-check speaker labels and proper nouns; (4) export as TXT or DOCX for a transcript document, SRT or VTT for subtitles. Processing takes 5-10 minutes per hour of video. No conversion to MP3 or WAV needed — that advice is from 2018-era tools.
Depends on use case. For one-off videos: VexaScribe gives a 5-minute anonymous preview (no signup) and 30 minutes free on signup. For unlimited free with privacy: OpenAI Whisper installed locally on your computer — free forever, offline, requires Python setup. For a YouTube video you own: YouTube auto-captions are free but ~60-70% accuracy vs 92-97% on modern Whisper-based tools. Avoid unknown "free video to text" sites — many use pre-2022 speech engines with significantly worse output.
Yes — and you should, in one upload. The same transcription pass produces the timestamped text; you then export it once as TXT/DOCX (transcript for blog posts, research, show notes) and once as SRT/VTT (subtitle file for YouTube, TikTok, Reels, Premiere Pro, Final Cut, DaVinci Resolve). The timestamps are identical — they're just serialized differently. Don't upload the video twice.
No. All modern AI transcription tools accept video files directly and extract the audio track internally. The "extract audio first" advice comes from older online converters (2018-era Google Cloud Speech v1, IBM Watson) that only accepted audio inputs. With current tools including VexaScribe, drag your MP4 / MOV / WEBM / MKV / AVI in directly — the server handles the audio-track extraction.
Universal formats that always work: MP4 (H.264 or H.265 codec), MOV (Apple QuickTime — iPhone recordings, Mac screen captures), WEBM (browser recordings, Loom exports), MKV (high-quality archives, OBS output), AVI (older Windows recordings), and WMV. Also works: FLV and MPEG/MPG. Rule of thumb: if the file plays in VLC or QuickTime, it transcribes. Video resolution has no effect on accuracy — the audio track is what the engine reads. For MP4-specific tips (iPhone recordings, screen captures), see our MP4 to text guide.
For anonymous previews: 200 MB per file (roughly 15-25 minutes of standard 720p video, or hours of audio-only). Anonymous previews truncate the transcript at 5 minutes — you see enough to know exactly what the full output will look like. On a free or paid account: up to 5 GB per file with no hard duration cap. Files longer than ~90 minutes are auto-chunked at silence boundaries during processing.
Modern Whisper-based transcribers achieve 92-97% word accuracy on clean English audio per the Open ASR Leaderboard and OpenAI's published Whisper benchmarks. Accuracy is governed by AUDIO quality, not video quality — a 4K video with bad mic audio transcribes worse than a 480p video with good mic audio. Accents (85-92%), background noise or music (80-90%), technical jargon, and overlapping speakers all reduce accuracy.
Yes, if the source is non-English. Whisper's translation mode takes a video in Spanish, French, German, Mandarin, and 90+ other languages and outputs an English transcript. The output is idiomatic English (not word-for-word) and works best when source audio is clear. For cross-language translation in the other direction (e.g. English video to Spanish transcript), first transcribe in the source language, then translate the text — VexaScribe supports 130+ target translation languages on any paid plan.
Two very different things. YouTube's auto-captions use Google's older ASR engine — accuracy is roughly 60-70% on accented English and drops fast for non-English content. Fine for low-stakes viewing. Modern AI video transcription (VexaScribe, Whisper, Speechmatics, AssemblyAI) uses transformer-based models trained on hundreds of thousands of hours of audio — 92-97% accuracy on clean English, with speaker labels, professional formatting, and export to TXT/DOCX/SRT/VTT. For publishing, business, or research use, the accuracy gap is dramatic enough that YouTube auto-captions aren't a real substitute.
These three words get used interchangeably, but they solve different problems. If you're producing a video and someone hands you a “transcript,” make sure it's the thing you actually need — the file format, timing, and viewer expectations differ.
| Attribute | Transcript | Subtitles | Captions |
|---|---|---|---|
| Primary format | TXT, DOCX (paragraph document) | SRT, VTT (timed cues) | SRT, VTT + speaker + non-speech cues |
| Use case | Reading, publishing, quoting, search | Overlay on video for translation or preference | Accessibility for deaf/hard-of-hearing viewers |
| Timestamp granularity | Paragraph or sentence (optional) | Per-cue (~2s chunks) | Per-cue with sound cues [music], [laughter] |
| Viewer needs | None — text is standalone | Same-audio speakers can skip; wrong-language speakers need them | Required for deaf/HoH viewers by ADA/WCAG |
| VexaScribe export | TXT · DOCX · JSON | SRT · VTT | SRT · VTT (add non-speech cues manually) |
VexaScribe produces all three formats from a single upload: a plain transcript (TXT/DOCX), timed subtitles (SRT/VTT), and captions if you add non-speech cues in the editor. One processing, every format — see the export options after your preview finishes.
Video to text is the process of converting the spoken audio in a video file into a written transcript using AI speech recognition. It's the same job people describe as “extract text from video” or “extract audio from video to text” — the tool pulls the audio track out of the container, runs it through a speech model, and returns a written document. Modern tools accept MP4, MOV, WEBM, MKV, and AVI directly — the audio track is extracted server-side and passed to a model like OpenAI Whisper Large-v3 or Speechmatics Melia-1. The output is a timestamped, speaker-labeled document exportable as TXT, DOCX, SRT, VTT, or JSON.
The tool itself hasn't really changed in the past two years — Whisper Large-v3, Speechmatics Melia-1, and a handful of commercial APIs deliver similar accuracy on similar workloads. What DID change is that people found half a dozen different names for the same job.
Content-type accuracy ranges on this page are approximate. Whisper Large-v3 baseline accuracy is consistent with published results on the Hugging Face Open ASR Leaderboard. See our Whisper accuracy analysis for the full third-party picture. First-party benchmark figures previously cited here were withdrawn on and have been removed.
VexaScribe runs Speechmatics Melia-1 as the primary transcription engine. Whisper Large-v3 serves as fallback for languages Melia-1 doesn't cover. Language auto-detection uses the user's browser locale to avoid the Supadata multi-caption-track ambiguity.
Anonymous 200 MB file cap and 5-minute preview truncation verified against production configuration (2026-08-19). Free tier of 30 minutes/month verified against user creation defaults. Translation support for 130+ target languages via Google Translate backing. Paid plans start at $2/month for 200 minutes — see /pricing for the current tier list.
We don't claim “100% accuracy” or “free forever.” The anonymous preview is free with a 5-minute cap, the free account has a monthly cap, paid plans have volume tiers. Competitor claims in the “when to reach for another tool” section are as of 2026-08-19 vendor documentation. Full editorial policy at /about/editorial-standards.