Video to Text — Transcribe Any Video Free (MP4, MOV, WEBM)

Upload MP4, MOV, WEBM, MKV, or AVI. Get a broadcast-quality transcript with speaker labels and timestamps in 99 languages. First 5 minutes free, no signup.

What is video-to-text transcription?

Video-to-text turns the spoken audio in a video into a written transcript using AI speech recognition. VexaScribe accepts MP4, MOV, WEBM, MKV, and AVI directly — no ffmpeg step, no MP3 conversion. Get a transcript with speaker labels, timestamps, and 99-language auto-detection. Export as TXT, DOCX, SRT, VTT, or JSON.

Drop your file here

or click to browse

Audio: MP3 · WAV · M4A · FLAC · OGG · AAC · AIFF · WMA · AMR · OPUS

Video: MP4 · MOV · AVI · MKV · WebM · FLV · WMV

Up to 200 MB · 5 min free preview

After you get your transcript — do more with it

Translate the transcript to 130+ languages

Publish your video internationally in one click, timestamps preserved

AI Summary of your video

Get the 5 key moments — skip the full replay

AI Chat over the video transcript

Ask "when did they announce the release date?" and get the timestamp

Export in every format

TXT · DOCX · SRT · VTT · JSON — one processing

Whisper Large-v3 + Speechmatics Melia-199 languagesGDPR EU hosting

Video Transcript Generator — How It Works

A video transcript generator is a tool that takes a video file (or URL) and produces a written transcript automatically, without a human transcriber typing it out. Modern AI-based generators like VexaScribe use OpenAI Whisper Large-v3 or Speechmatics Melia-1 under the hood, hit 92-97% accuracy on clean English audio, and process a 60-minute video in 5-10 minutes. The tool at the top of this page is a video transcript generator with a 5-minute free preview.

Video Transcript Generator vs Manual Transcription

AspectAI generatorManual transcription
Time per hour of video5-10 minutes processing4-6 hours of typing
Cost per hour of video$0.24-$3.00 (or $2/mo for 200 minutes)$90-$150 (professional service)
Accuracy on clean audio92-97%99%+ (with review)
Best forPublishing, subtitles, research, show notesLegal filings, ADA-compliant broadcast captions

For most publishing and content workflows the AI generator is the correct pick — the accuracy gap is small enough that a 5-10 minute editing pass closes it. Human transcription only wins when a stamp of certification is legally required.

Supported Video Formats

Every format below works with the tool at the top of this page. Video resolution has no effect on transcript accuracy — the audio track is what the engine reads. Rule of thumb: if the file plays in VLC or QuickTime, it transcribes.

FormatCommon source
.mp4The universal default — H.264 or H.265 codec, used by TikTok, Instagram, most cameras, screen recorders, and iPhone. MP4-specific guide
.movApple QuickTime container — older iPhones, Final Cut Pro / DaVinci Resolve ProRes exports.
.webmWeb-native (VP9/VP8 + Opus codec) — Chrome and OBS screen recorders, browser captures, Loom exports.
.mkvMatroska container — open format, common for desktop screen recorders and downloaded video archives.
.aviLegacy Microsoft container — older screen recorders, archived footage.
.wmvWindows Media Video — Windows Movie Maker output and older corporate recordings.
.flvLegacy Flash video — older lecture libraries and corporate training archives.
.mpg / mpegOlder camcorders and broadcast archives.

MP4 Transcript Notes

MP4 is the universal video container — H.264 or H.265 video codec with AAC or Opus audio. TikTok, Instagram, YouTube exports, iPhone recordings (iOS 11+), Android recordings, and most screen recorders all output MP4. Video resolution doesn't affect transcription accuracy — a 4K MP4 with poor mic audio transcribes worse than a 720p MP4 with a lavalier close to the speaker. For MP4-specific workflows including iPhone captures and screen recordings, see our MP4 to text guide.

MOV Transcript Notes

MOV is Apple's QuickTime container — used by older iPhones (iOS 10 and below), Mac screen captures (Cmd+Shift+5), Final Cut Pro exports, and DaVinci Resolve ProRes output. The underlying codec is often the same H.264 as MP4, just wrapped in a QuickTime header. MOV files transcribe identically to MP4 — no special handling needed on this page's tool. If your MOV file plays in QuickTime, it transcribes.

MKV Transcript Notes

MKV (Matroska) is the open-source container of choice for OBS Studio screen recordings, high-quality video archives, and downloaded video files with multiple audio tracks. When an MKV has multiple audio tracks (some Zoom exports do), our tool uses the first track by default — re-process on a different track from the editor if the first track is silent or wrong-language. MKV files can hold FLAC or Opus audio, both of which transcribe with no accuracy penalty.

Audio-only files also work: MP3, M4A, WAV, AAC, FLAC, OGG, OPUS. For audio-specific guides see MP3 to text, M4A to text, WAV to text, and OGG to text.

How to Transcribe a Video in 4 Steps

The workflow is the same for a 15-second TikTok and a 3-hour lecture recording.

  1. 1

    Drop the video into the uploader

    Drag your MP4, MOV, WEBM, MKV, or AVI file directly into the tool zone at the top of this page — no pre-conversion, no ffmpeg, no audio extraction. Anonymous previews accept files up to 200 MB.

  2. 2

    Let language auto-detection do its job

    99 languages, auto-detected from the first speech in the audio track. Override manually only if the video opens with music or long silence. Non-Latin scripts (Chinese, Japanese, Arabic, Hindi) and right-to-left languages (Arabic, Hebrew) render cleanly in every export format.

  3. 3

    Review speakers and fix proper nouns

    Automatic speaker diarization labels up to 8 distinct voices. Spot-check names, companies, and product SKUs — proper nouns are what AI gets wrong most (20-30% error rate even on clean audio). The editor's search-and-replace fixes all instances of a name in one pass.

  4. 4

    Export in the format you actually need

    TXT or DOCX for a transcript document. SRT or VTT for subtitle files (YouTube, Premiere Pro, Final Cut, DaVinci Resolve). JSON with word-level timestamps for custom video editors and developer pipelines. The same processed file gives you every format — no re-upload needed.

Sample Transcript Output

Every transcript includes segment-level timestamps and automatic speaker labels. Word-level timestamps are available on paid plans. Below is an anonymized excerpt of what you'll see in the editor and in the exported TXT / DOCX.

[00:00:03] Speaker 1: The thing about video transcription
[00:00:06] Speaker 1: is that most people don't realize it's just
[00:00:09] Speaker 1: the audio track underneath.
[00:00:12] Speaker 2: Right — the container format doesn't matter.
[00:00:15] Speaker 2: MP4, MOV, MKV are all just wrappers.
[00:00:19] Speaker 1: Exactly. The AI only sees the audio.
[00:00:22] Speaker 1: A 4K video with bad mic audio
[00:00:25] Speaker 1: transcribes worse than a 480p video with a good mic.

Prefer paragraphs without timestamps or speaker labels? Toggle them off before export — the DOCX and TXT outputs can be paragraph-only for blog posts and articles. SRT and VTT always preserve timing (subtitle formats require it). JSON keeps the raw word-level timing array for developers building custom video editors or search interfaces.

How Accurate Is AI Video Transcription?

Modern AI video transcription achieves 92-97% word accuracy on clean English audio — 3-8% Word Error Rate. Accuracy is governed by audio quality, not video resolution — a 4K video with bad mic audio transcribes worse than a 480p video with good mic audio.

Real-world accuracy varies dramatically by content type. The ranges below are approximate and align with published third-party evaluations on the Hugging Face Open ASR Leaderboard.

Content typeAccuracyEditing time
Studio explainer / single-speaker tutorial (treated room, good mic)95-97%5-10 min/hr
Podcast-style video (two speakers, dedicated mics)93-96%10-15 min/hr
Zoom / Teams / Meet recording (built-in laptop mics)91-95%10-15 min/hr
Lecture / classroom recording (ceiling or wireless mic)89-94%10-20 min/hr
TikTok / Instagram Reel (in-app recording with music underneath)82-90%15-25 min/hr
Vlog (outdoor, handheld, ambient noise)80-88%20-30 min/hr
Phone-recorded interview (compressed audio, near-field)85-92%15-20 min/hr
Multi-language / heavy accents82-90%15-25 min/hr

Proper nouns — brand names, product SKUs, technical jargon, foreign names — miss more often than regular vocabulary (20-30% error rate even on clean audio). Full accuracy methodology and per-model comparison on our Whisper accuracy analysis.

99 Languages + Auto-Detection: The Honest Version

The language dropdown pre-selects based on your browser locale. A Turkish speaker on a Turkish browser sees Turkish selected by default; a Portuguese speaker sees Portuguese; a Chinese speaker sees Mandarin. This is deliberate, not cosmetic — it fixes a bug that affects most third-party transcript tools built on Supadata's caption API.

The bug: when a video has multiple auto-generated caption tracks (TikTok and Instagram videos often have three or four — Vietnamese, Indonesian, English, plus the original language), a tool that asks the caption API for “auto” can get a random language back. That's where the notorious “why is my English video's transcript in Spanish” complaint comes from. Pre-selecting the user's actual browser language cuts the miss rate significantly.

All 99 Whisper Large-v3 languages are supported. Non-Latin scripts render cleanly in the editor and every export format. Right-to-left scripts preserve reading direction in DOCX and TXT. Auto-detection doesn't work perfectly on very short clips (under 3 seconds) — for a 5-second Reel, override manually if needed.

For translation, VexaScribe supports 130+ target languages — transcribe in the source language, then translate the transcript in one click from the editor. For dedicated video translation workflows see video translator.

Is It Really Free? What You Get Before Signup

Yes — genuinely free, with two tiers. The anonymous preview (200 MB max, 5-minute transcript, one per week per IP) requires no signup and shows the full transcript with speaker labels and timestamps. The free account (30 minutes/month, email only, no card) unlocks the editor, full-length transcripts, and every export format (TXT, DOCX, SRT, VTT, JSON). Paid plans start at $2/month for 200 minutes.

TierPriceLimitsWhat's included
Anonymous previewFree · no signup200 MB max file · 5-min preview · 1 per week per IPFull 5-min transcript with speaker labels + timestamps · TXT/SRT download
Free accountFree · email only30 minutes/month · no card requiredFull transcript · editor · TXT/DOCX/SRT/VTT/JSON export · 130+ language translation
Starter$2/month200 minutes/monthEverything above · AI summary · AI Chat · priority processing

Free alternatives without any signup include YouTube's built-in transcript (English-primary, ~65% accuracy, YouTube-only) and self-hosted OpenAI Whisper (Python + GPU required, free forever). See free transcription options for the honest comparison.

5 Methods to Create a Transcript from a Video

This page's tool is one of five ways to get from video to text. Honest comparison of when each makes sense:

1.AI transcription tool (fastest, cheapest, good enough for most use cases)

Drop the video into VexaScribe, Otter, HappyScribe, or similar. 5-10 minutes per hour of video, 92-97% accuracy on clean English audio. $0.004-$0.05 per minute at list rates. Best for: publishing, research, subtitles, show notes, meeting notes, and anything that isn't legally binding.

2.YouTube auto-captions (free but low-accuracy, YouTube-only)

If the video is already on YouTube, click ... under the video → Show transcript. Free, instant, but Google's older ASR engine returns roughly 60-70% accuracy on accented English and less on non-English content. Not a real substitute for AI transcription for publishing use.

3.Human transcription service (slow, expensive, certified accuracy)

Rev.com, Scribie, or GMR — human transcribers deliver 99%+ accuracy at $1.50-$2.50 per audio minute with 12-48 hour turnaround. Only worth the cost when you need certified verbatim for legal filings, ADA-compliant broadcast captions, or other stamp-of-certification use cases.

4.Self-hosted OpenAI Whisper (free forever, requires setup)

Install openai-whisper or faster-whisper on a machine with a GPU. Free MIT license, offline, no per-minute cost, complete privacy. Requires Python skills and ~10 GB VRAM for Large-v3. Best for: high volume, privacy-sensitive content, or anyone comfortable with ML tooling.

5.Manual transcription (slowest, most tedious, sometimes necessary)

Play the video and type as you listen. 4-6 hours of typing per hour of video for a skilled transcriber. Only worth it for very short clips, audio too poor for AI to handle, or heavily technical content where you know the domain better than any model.

When to Reach for Another Tool (Honest)

The tool at the top of this page fits most video-transcription jobs. Four scenarios where a different tool is a better fit:

  1. 1.HappyScribe — team-collaboration for podcasters and video producers

    Best for teams needing shared transcription workflows, per-project libraries, and human-review passes on top of the AI output. Higher per-minute cost than VexaScribe but bundles the collaboration UX. Fits 2+ person editorial teams.

  2. 2.VexaScribe — best free preview + 30-min free tier

    Anonymous 5-minute preview with no signup, then 30 minutes free on signup, then $2/month. Speechmatics Melia-1 primary engine with Whisper Large-v3 fallback. Best for one-off video transcription, creators, and researchers.

  3. 3.Otter.ai — best for live meeting recording

    Purpose-built for live Zoom, Teams, and Google Meet meetings with a bot participant that joins the call and transcribes in real time. For finished video files that you already downloaded, generic tools beat Otter on price and format flexibility.

  4. 4.Rev.com — best for certified verbatim (human review)

    Human transcription at $1.50-$2.50 per audio-minute with 12-48 hour turnaround. Only worth the cost when you need certified verbatim for legal filings, ADA-compliant broadcast captions, or other stamp-of-certification use cases. AI-first tools cover everything else at 100-500× lower cost.

For SRT files specifically, our SRT generator is purpose-built with subtitle-specific export. For batch upload of many videos at once, see bulk transcription. For a full 14-model API comparison see our transcription API comparison.

Frequently Asked Questions

Can ChatGPT transcribe a video?

No, ChatGPT does not directly transcribe video files. ChatGPT can accept short audio inputs (under about a minute) via voice mode but does not accept MP4, MOV, or WEBM video uploads. Use a dedicated tool like VexaScribe to extract audio from the video and transcribe in one step — drop the video file, get a full transcript with timestamps and speaker labels. VexaScribe runs OpenAI Whisper Large-v3 (the same model family behind many AI products) and handles the video-to-audio extraction server-side.

How do I transcribe a video to text?

Four steps: (1) drop the video file into a transcription tool that accepts video directly (VexaScribe, Otter, HappyScribe, Rev — modern tools take MP4, MOV, WEBM, MKV, AVI without pre-extracting audio); (2) let it auto-detect the language (99 languages supported on Whisper-based tools); (3) spot-check speaker labels and proper nouns; (4) export as TXT or DOCX for a transcript document, SRT or VTT for subtitles. Processing takes 5-10 minutes per hour of video. No conversion to MP3 or WAV needed — that advice is from 2018-era tools.

What's the best free video transcription tool in 2026?

Depends on use case. For one-off videos: VexaScribe gives a 5-minute anonymous preview (no signup) and 30 minutes free on signup. For unlimited free with privacy: OpenAI Whisper installed locally on your computer — free forever, offline, requires Python setup. For a YouTube video you own: YouTube auto-captions are free but ~60-70% accuracy vs 92-97% on modern Whisper-based tools. Avoid unknown "free video to text" sites — many use pre-2022 speech engines with significantly worse output.

Can I get a transcript AND subtitles from the same video?

Yes — and you should, in one upload. The same transcription pass produces the timestamped text; you then export it once as TXT/DOCX (transcript for blog posts, research, show notes) and once as SRT/VTT (subtitle file for YouTube, TikTok, Reels, Premiere Pro, Final Cut, DaVinci Resolve). The timestamps are identical — they're just serialized differently. Don't upload the video twice.

Do I need to extract audio from the video first?

No. All modern AI transcription tools accept video files directly and extract the audio track internally. The "extract audio first" advice comes from older online converters (2018-era Google Cloud Speech v1, IBM Watson) that only accepted audio inputs. With current tools including VexaScribe, drag your MP4 / MOV / WEBM / MKV / AVI in directly — the server handles the audio-track extraction.

What video formats can I transcribe?

Universal formats that always work: MP4 (H.264 or H.265 codec), MOV (Apple QuickTime — iPhone recordings, Mac screen captures), WEBM (browser recordings, Loom exports), MKV (high-quality archives, OBS output), AVI (older Windows recordings), and WMV. Also works: FLV and MPEG/MPG. Rule of thumb: if the file plays in VLC or QuickTime, it transcribes. Video resolution has no effect on accuracy — the audio track is what the engine reads. For MP4-specific tips (iPhone recordings, screen captures), see our MP4 to text guide.

How long can the video be? What's the file size limit?

For anonymous previews: 200 MB per file (roughly 15-25 minutes of standard 720p video, or hours of audio-only). Anonymous previews truncate the transcript at 5 minutes — you see enough to know exactly what the full output will look like. On a free or paid account: up to 5 GB per file with no hard duration cap. Files longer than ~90 minutes are auto-chunked at silence boundaries during processing.

How accurate is an AI video transcriber?

Modern Whisper-based transcribers achieve 92-97% word accuracy on clean English audio per the Open ASR Leaderboard and OpenAI's published Whisper benchmarks. Accuracy is governed by AUDIO quality, not video quality — a 4K video with bad mic audio transcribes worse than a 480p video with good mic audio. Accents (85-92%), background noise or music (80-90%), technical jargon, and overlapping speakers all reduce accuracy.

Can I translate a video to text in English?

Yes, if the source is non-English. Whisper's translation mode takes a video in Spanish, French, German, Mandarin, and 90+ other languages and outputs an English transcript. The output is idiomatic English (not word-for-word) and works best when source audio is clear. For cross-language translation in the other direction (e.g. English video to Spanish transcript), first transcribe in the source language, then translate the text — VexaScribe supports 130+ target translation languages on any paid plan.

How is video transcription different from YouTube auto-captions?

Two very different things. YouTube's auto-captions use Google's older ASR engine — accuracy is roughly 60-70% on accented English and drops fast for non-English content. Fine for low-stakes viewing. Modern AI video transcription (VexaScribe, Whisper, Speechmatics, AssemblyAI) uses transformer-based models trained on hundreds of thousands of hours of audio — 92-97% accuracy on clean English, with speaker labels, professional formatting, and export to TXT/DOCX/SRT/VTT. For publishing, business, or research use, the accuracy gap is dramatic enough that YouTube auto-captions aren't a real substitute.

Transcript vs Subtitles vs Captions: What's the Difference?

These three words get used interchangeably, but they solve different problems. If you're producing a video and someone hands you a “transcript,” make sure it's the thing you actually need — the file format, timing, and viewer expectations differ.

AttributeTranscriptSubtitlesCaptions
Primary formatTXT, DOCX (paragraph document)SRT, VTT (timed cues)SRT, VTT + speaker + non-speech cues
Use caseReading, publishing, quoting, searchOverlay on video for translation or preferenceAccessibility for deaf/hard-of-hearing viewers
Timestamp granularityParagraph or sentence (optional)Per-cue (~2s chunks)Per-cue with sound cues [music], [laughter]
Viewer needsNone — text is standaloneSame-audio speakers can skip; wrong-language speakers need themRequired for deaf/HoH viewers by ADA/WCAG
VexaScribe exportTXT · DOCX · JSONSRT · VTTSRT · VTT (add non-speech cues manually)

VexaScribe produces all three formats from a single upload: a plain transcript (TXT/DOCX), timed subtitles (SRT/VTT), and captions if you add non-speech cues in the editor. One processing, every format — see the export options after your preview finishes.

What Is Video to Text?

Video to text is the process of converting the spoken audio in a video file into a written transcript using AI speech recognition. It's the same job people describe as “extract text from video” or “extract audio from video to text” — the tool pulls the audio track out of the container, runs it through a speech model, and returns a written document. Modern tools accept MP4, MOV, WEBM, MKV, and AVI directly — the audio track is extracted server-side and passed to a model like OpenAI Whisper Large-v3 or Speechmatics Melia-1. The output is a timestamped, speaker-labeled document exportable as TXT, DOCX, SRT, VTT, or JSON.

The tool itself hasn't really changed in the past two years — Whisper Large-v3, Speechmatics Melia-1, and a handful of commercial APIs deliver similar accuracy on similar workloads. What DID change is that people found half a dozen different names for the same job.

Methodology & Sources

Accuracy claims

Content-type accuracy ranges on this page are approximate. Whisper Large-v3 baseline accuracy is consistent with published results on the Hugging Face Open ASR Leaderboard. See our Whisper accuracy analysis for the full third-party picture. First-party benchmark figures previously cited here were withdrawn on and have been removed.

Engine used on this page

VexaScribe runs Speechmatics Melia-1 as the primary transcription engine. Whisper Large-v3 serves as fallback for languages Melia-1 doesn't cover. Language auto-detection uses the user's browser locale to avoid the Supadata multi-caption-track ambiguity.

Pricing and product limits

Anonymous 200 MB file cap and 5-minute preview truncation verified against production configuration (2026-08-19). Free tier of 30 minutes/month verified against user creation defaults. Translation support for 130+ target languages via Google Translate backing. Paid plans start at $2/month for 200 minutes — see /pricing for the current tier list.

Editorial standards

We don't claim “100% accuracy” or “free forever.” The anonymous preview is free with a 5-minute cap, the free account has a monthly cap, paid plans have volume tiers. Competitor claims in the “when to reach for another tool” section are as of 2026-08-19 vendor documentation. Full editorial policy at /about/editorial-standards.

Ready to transcribe your video?

Scroll back up and drop your file. First 5 minutes free, no signup. Sign up when you want the full transcript and every export format.