Captions — the short answer
Captions are on-screen text that reproduces the dialogue plus important non-speech audio (sound effects, music cues, speaker labels) of a video. They exist primarily for viewers who cannot hear the audio — deaf or hard-of-hearing audiences, viewers watching muted, and viewers in loud environments. Modern captions come in five types: closed captions (CC, toggleable), open captions (burned-in), SDH (subtitle-delivered captions on streaming), live captions (real-time), and AI-generated captions (produced by ASR models).
Expanded definition — what captions actually contain
A caption track for a video is more than a transcript on screen. It reproduces the full audio experience in text so a viewer who can't hear the audio understands what's happening. Concretely, a well-produced caption track includes:
- Spoken dialogue — verbatim or lightly edited for readability, timed to appear as the words are spoken.
- Speaker identifications — LISA:, MARK (V.O.):, NARRATOR: — especially when multiple speakers appear off-screen or unseen.
- Sound effects — [door slams], [phone ringing], [thunder], [laughter]. Anything that adds to plot, mood, or continuity.
- Music cues — [upbeat rock music], [♪ You Are My Sunshine ♪], [ominous instrumental]. Music that sets mood or contains recognizable lyrics.
- Non-speech vocalizations — [sighs], [laughs nervously], [gasps], [groans] — when they convey emotion or meaning.
- Foreign-language dialogue — either translated in the caption itself or bracketed as [speaking Spanish] when the audience isn't meant to understand it.
- Off-screen sound — [distant siren], [car engine approaching] — sounds that affect the scene but aren't on camera.
What captions do NOT typically include: purely ambient audio that doesn't affect the scene, background chatter without narrative purpose, or minor sounds that would clutter the caption pane without serving comprehension. Style guides (Netflix Timed Text, BBC Subtitle Guidelines) explicitly recommend omitting these.
The five types of captions in modern use
Captions exist in five practical forms today, distinguished by how they're delivered and whether the viewer controls their visibility. The right type depends on the platform, the audience, and the delivery constraints. Extraction table:
| Type | Toggleable? | Non-speech content? | Typical delivery |
|---|---|---|---|
| Closed Captions (CC) | Yes — viewer turns on/off | Yes — [door slams], [music] | Separate track (CEA-608/708 on broadcast; .srt/.vtt/TTML on streaming) |
| Open Captions (Burned-In) | No — always visible | Yes, when the creator adds them | Part of the video pixels (encoded permanently into the frame) |
| SDH — Subtitles for the Deaf and Hard of Hearing | Yes — viewer turns on/off | Yes — [door slams], [music], speaker IDs | Subtitle track (text or bitmap), typically IMSC1.1 or WebVTT on streaming |
| Live Captions | Yes on video platforms; open on some live displays | Sometimes — depends on the captioner | Real-time text stream (stenography for broadcast; ASR for consumer platforms) |
| AI-Generated Captions | Yes — delivered as standard closed captions | Only if the tool explicitly detects and marks them | SRT, VTT, TXT, or embedded as CC on video platforms |
Closed Captions (CC)
The default caption format for regulated broadcast and mainstream streaming. Viewers can turn them on or off in the player. On US broadcast TV, closed captions are encoded as CEA-608 (line 21 in analog NTSC) or CEA-708 (digital ATSC) data. On streaming and web platforms, they're delivered as separate text-based subtitle files (SRT, WebVTT, TTML) selected via the language menu.
Typically used for: Broadcast TV, streaming platforms, YouTube, Blu-ray/DVD, Zoom, corporate video
Open Captions (Burned-In)
Captions burned into the video pixels themselves — always visible, cannot be turned off. There's no separate data track; the captions are part of the video frame. Standard for social media where autoplay is silent and viewers can't (or won't) toggle CC on. Also used at live events where a display shows captions regardless of viewer preference, and in some cinema accessibility screenings.
Typically used for: Social media (TikTok, Reels, LinkedIn, Instagram), live event displays, accessibility cinema, marketing video
SDH — Subtitles for the Deaf and Hard of Hearing
A hybrid: uses subtitle infrastructure (bottom-center text, styled like translation subtitles) but includes caption-grade content (speaker IDs, sound effects, music cues). Streaming platforms use SDH because their delivery pipelines don't carry broadcast caption data formats (CEA-608/708) — they deliver captions through subtitle tracks instead. When you pick 'English [SDH]' or 'English [CC]' on Netflix, you're getting an SDH track.
Typically used for: Netflix, Disney+, HBO Max, Amazon Prime Video, Apple TV+, Blu-ray, DVD
Live Captions
Generated in real time as speech happens. Two main production methods: professional stenographers using stenotype machines (broadcast news, courts, congressional hearings) achieving ~95%+ accuracy with 2-4 second latency, and AI-powered automatic speech recognition (Zoom, Teams, YouTube Live) achieving 85-95% accuracy with 3-5 second latency. Live captions are increasingly available at conferences, city council meetings, and public events via ASR-powered caption displays.
Typically used for: Live TV news, congressional hearings, sports broadcasts, live streams (YouTube, Twitch), Zoom / Teams / Google Meet, live events with caption displays
AI-Generated Captions
Produced by automatic speech recognition (ASR) models like OpenAI Whisper Large-v3, Deepgram Nova, or Google Cloud Speech. Modern models achieve 90-95% accuracy on clean English audio at a fraction of human-transcription cost (~$0.005-0.10 per audio minute vs $1.25-3.50 per minute for human transcription). Standard for creator captioning, podcast episode captioning, meeting transcription, and volume production. Typically not sufficient alone for broadcast or legal compliance without a human review pass.
Typically used for: Creator workflows (YouTube, podcasts, corporate video), volume production, budget-constrained captioning
For the closed-caption-specific deep dive — CEA-608/708 technical detail, how to add or remove CC on TV/YouTube/streaming, the trademark history of the “CC” symbol — see what is closed captioning?
Captions vs subtitles — the practical difference
The single most common confusion around captions is how they differ from subtitles. Both are on-screen text; both accompany video. The difference is the viewer they're designed for:
- Captions assume the viewer cannot hear the audio. They include dialogue plus non-speech information (sound effects, speaker IDs, music cues).
- Subtitles assume the viewer can hear the audio but doesn't understand the spoken language. They include dialogue only.
| Aspect | Captions | Subtitles |
|---|---|---|
| Primary purpose | Accessibility — viewer cannot hear audio | Translation — viewer doesn't speak the language |
| Includes dialogue | Yes | Yes |
| Includes speaker IDs | Yes (LISA:, MARK:) | Usually no |
| Includes sound effects | Yes ([door slams], [thunder]) | No |
| Includes music cues | Yes ([upbeat music]) | Sometimes (song lyrics only) |
| Assumes viewer can hear? | No | Yes |
| Same language as audio? | Usually yes | Usually different language |
| Legally required (US)? | Yes — FCC, CVAA, ADA case law | No |
| Originated | 1972 PBS open captions; 1980 mainstream CC | 1930s foreign-language film |
Regional terminology matters. In the United States, “captions” and “subtitles” are two distinct words for two distinct things. In the UK, Ireland, Australia, and much of continental Europe, “subtitles” is used for both accessibility and translation — BBC subtitles, for instance, include speaker IDs and sound effects (they're accessibility captions in US terminology). Context and delivery platform disambiguate.
For the full comparison — including SDH, closed vs open, US/UK/EU legal frameworks, and how to decide which your video needs — see captions vs subtitles.
Where you see captions — platform by platform
Captions are now a default expectation across nearly every video platform in 2026. The specific implementation varies by platform:
YouTube
Auto-captions (Google ASR, 85-92% English), creator-uploaded SRT/VTT, community-contributed on some channels
Look for the CC button on the player
Netflix / Disney+ / HBO Max / Amazon Prime
SDH tracks in 10-40 languages per title (IMSC1.1 delivery), forced-narrative subtitles for foreign-language dialogue
Language menu → look for [CC] or [SDH] tag
TikTok / Instagram Reels / Shorts
Built-in auto-captions plus burned-in creator captions (95%+ of creators add captions in 2026)
Autoplay is muted by default — captions are essential for engagement
LinkedIn Video
Native auto-captions plus SRT upload; enterprise-standard for professional accessibility
70%+ of LinkedIn video is watched with sound off
Zoom / Microsoft Teams / Google Meet
Live AI captions and real-time transcript; recording captions in some tiers
Standard feature since 2020-2021 rollout
Broadcast TV (US, UK, EU)
CEA-608/708 on US ATSC; DVB-Subtitle in EU; Teletext (legacy)
Regulated by FCC (US), Ofcom (UK), AVMSD (EU)
Blu-ray / DVD
SDH tracks in multiple languages; sometimes traditional CC on legacy discs
Selectable from disc menu
Live events / conferences
Increasingly available via ASR-powered caption displays (Otter, Wordly, Ai-Media Scribblr)
Legally expected for regulated public events under ADA/EAA
How captions are created — five production methods
The captioning pipeline you pick trades accuracy for cost, speed, and latency. The five common methods in 2026:
| Method | Accuracy | Turnaround / latency | Cost per audio minute |
|---|---|---|---|
| Human transcription | 99%+ | 24-72 hours turnaround typical | $1.25 - $3.50 |
| AI + human review | 97-99% | 1-24 hours typical | $0.20 - $1.00 |
| AI-only transcription | 90-95% on clean English (Whisper Large-v3) | 5-15 min per audio hour | $0.005 - $0.10 |
| Live stenography (real-time human) | ~95%+ | 2-4 seconds | $2.00 - $6.00 |
| Live ASR (real-time AI) | 85-95% | 3-5 seconds | Included in platform / $0.05-0.50 |
Human transcription
Typical use: Broadcast, legal, medical, high-stakes brand video, compliance-critical accessibility
Providers: Rev, 3PlayMedia, VITAC, Verbit (with human tier)
AI + human review
Typical use: Streaming platform SDH (Netflix, Disney+), professional creator content, corporate video needing 98%+
Providers: 3PlayMedia (AI+review tier), Rev AI+human, Verbit AI+human, HappyScribe hybrid
AI-only transcription
Typical use: Creator content, podcasts, meeting transcripts, volume production, YouTube video captioning
Providers: VexaScribe (Whisper Large-v3), Otter.ai, Descript, Rev AI, Deepgram, AssemblyAI
Live stenography (real-time human)
Typical use: Live TV news, congressional hearings, courts, live sports, high-stakes live events
Providers: VITAC, CaptionMax, National Court Reporters Association members
Live ASR (real-time AI)
Typical use: Zoom / Teams / Meet meetings, YouTube live streams, Twitch, conference caption displays
Providers: Zoom (built-in), Google Meet (built-in), YouTube Live (built-in), Otter Live, Wordly
For most creator workflows in 2026, AI-only transcription is the pragmatic choice — 90-95% accuracy on clean English audio at pennies per minute, exported as SRT or VTT and imported into any video platform. VexaScribe generates captions using OpenAI's Whisper Large-v3 in 99 languages; see caption generator for the tool page and how accurate is Whisper for the underlying accuracy math. For broadcast, legal, or medical compliance, budget for AI + human review or full human transcription instead.
Who uses captions — well beyond accessibility
Captions were designed for deaf and hard-of-hearing viewers. In 2026 they're used by a much broader audience. A widely-cited Ofcom study (2006) found roughly 80% of UK caption users were not deaf or hard-of-hearing. Modern caption usage rates on Netflix, YouTube, and Amazon Prime run 50-80% among general audiences.
Who actually turns captions on today:
- Deaf and hard-of-hearing viewers — the original and primary audience; captions are essential accessibility.
- Silent-autoplay viewers — Facebook, Instagram, TikTok, and LinkedIn autoplay videos muted. 85%+ of Facebook video is watched with sound off (Digiday, Facebook internal data); captions are required for engagement, not optional.
- Viewers in loud environments — gyms, public transport, open-plan offices, coffee shops, bars, factories.
- Viewers in quiet environments — bedtime, hospitals, libraries, sleeping partner nearby, sleeping baby.
- Non-native English speakers — captions bridge the comprehension gap for accented speech, technical vocabulary, or rapid dialogue.
- Language learners — captions in the target language reinforce vocabulary and pronunciation.
- Viewers with attention or processing differences — captions provide a redundant channel that supports comprehension for many neurodivergent viewers.
- Viewers of accented or technical content — even native speakers catch details more reliably by reading than by listening in rapid or heavily-accented speech.
- Content re-purposers — captions get pulled into transcripts for SEO, blog posts, podcast show notes, and searchable video archives.
For creators and businesses, captions have crossed the threshold from accessibility feature to core distribution feature — video without captions loses reach on social, muted-autoplay platforms, and reduced-accessibility audiences.
Legal requirements at a glance — US, UK, EU
Captioning law varies by jurisdiction, platform, and content type. This is a summary — for a deeper legal breakdown, see captions vs subtitles which covers US, UK, and EU frameworks in detail. Consult a lawyer for specific compliance questions.
| Jurisdiction / context | Requirement | Quality standard |
|---|---|---|
| US — Broadcast | Required by FCC (47 CFR §79.1) on all broadcast, cable, and satellite TV | Four quality standards: accuracy, synchronicity, completeness, placement |
| US — Online (previously aired) | Required by CVAA (2010) — online video previously shown on US TV with captions must retain captions | Same quality as original TV airing |
| US — Online (general) | Increasingly required via ADA case law — NAD v. Netflix (2012), NAD v. Harvard (2019-2020) | WCAG 2.1 Level A / Level AA as baseline |
| US — Federal agencies | Required by Section 508 of the Rehabilitation Act | Aligned with WCAG 2.0 Level AA (2017 refresh) |
| UK — Broadcast | Communications Act 2003 + Ofcom Code — 80% subtitling for qualifying channels, higher for PSBs | Ofcom quality guidelines |
| EU — Media services | AVMSD (Directive 2018/1808) — Member States ensure media services accessible continuously and progressively | Set by national implementations |
| EU — Products / e-commerce | European Accessibility Act (Directive 2019/882) — effective 28 June 2025 | Covers audiovisual media access services in scope |
| International — Web | WCAG 2.1 SC 1.2.2 Level A (prerecorded) and 1.2.4 Level AA (live) | Accepted global accessibility baseline |
The trend line: captioning requirements are expanding, not contracting. The EAA (June 2025) brought EU e-commerce and audiovisual media services under a common accessibility framework. ADA case law in the US continues to extend captioning obligations to online video. WCAG 2.2 (2023) reaffirmed 1.2.2 (captions on prerecorded video, Level A) and 1.2.4 (captions on live video, Level AA) as accepted global baselines.
File formats used for captions
Captions are delivered in different formats depending on the platform. For most creator workflows in 2026, SRT is the right default — supported by YouTube, Vimeo, Premiere, DaVinci, CapCut, VLC, and virtually every platform. Broadcast and OTT delivery use specialized formats.
| Format | Description | Where used |
|---|---|---|
| SRT (SubRip) | Plain text; sequence numbers + timestamps. Most universal caption format. | YouTube, Vimeo, most creator platforms, VLC, Premiere, DaVinci, CapCut |
| WebVTT (.vtt) | W3C standard for HTML5 <track> element; supports styling, positioning, regions. | HTML5 web video, HLS streaming, video.js and other web players |
| SCC (Scenarist) | Binary CEA-608 byte pairs; used to deliver legacy US broadcast captions. | US broadcast (CEA-608 delivery), legacy TV workflows |
| TTML / IMSC1.1 | XML-based; W3C standard with rich styling and multi-language support. | Netflix (IMSC1.1), professional OTT delivery, broadcast |
| CEA-608 / CEA-708 | On-air broadcast caption data formats decoded by TV sets. | US analog NTSC (608, 'line 21 captions'); digital ATSC (708) |
| DVB-Subtitle | Bitmap-based subtitle format used in European digital broadcast. | European digital TV (DVB-T, DVB-S) |
| ASS / SSA | Advanced SubStation Alpha — rich positioning, fonts, karaoke effects. | Anime, fan-subtitle communities, karaoke, styled fansubs |
Practical recommendation for creators: generate an SRT first — it works everywhere. Convert to WebVTT with a tool like our SRT ↔ VTT converter if you're embedding via HTML5 <track>. Use SCC or TTML only when delivering to US broadcast or professional OTT pipelines. For creating an SRT from scratch, see how to create an SRT file.
Captions vs transcripts — related but different
A common conflation. Captions are time-synced text that appears on screen at the moment the audio plays. Transcripts are plain-text records of what was said, without timing or synchronization to video. Both come from the same source material — often produced together — but they serve different jobs.
- Captions serve accessibility during video playback. They're embedded or attached to the video and appear synchronized to speech.
- Transcripts serve searchability, SEO, and use cases where the reader isn't watching the video — podcast show notes, meeting summaries, legal records, LLM inputs.
- Most modern captioning tools produce both from a single transcription pass: word-level timestamps generate SRT captions on export and TXT/DOCX transcripts on separate download.
For the transcription-first perspective (what the audio-to-text step actually does), see what is audio transcription? For plain-text output workflows, see transcribe audio to text.
Frequently asked questions about captions
What are captions in simple terms?
Captions are on-screen text that represent both the dialogue and important non-speech audio of a video — sound effects, music cues, speaker identifications. They exist primarily so viewers who cannot hear the audio (deaf, hard-of-hearing, watching muted) can follow the content in full. Captions differ from subtitles: subtitles cover only spoken dialogue and assume the viewer can hear. Captions can be closed (toggleable by the viewer), open (permanently burned into the video), or delivered as SDH (Subtitles for the Deaf and Hard of Hearing) on streaming platforms.
What are the different types of captions?
Five main types in 2026. (1) Closed captions (CC) — viewer toggles them on or off; delivered as a separate data track (CEA-608/708 on broadcast, .srt/.vtt/TTML on streaming). (2) Open captions — burned into the video pixels; always visible; can't be turned off; standard for social media where autoplay is muted. (3) SDH — Subtitles for the Deaf and Hard of Hearing; caption-grade content delivered as a subtitle track; the standard on Netflix, Disney+, and other streaming platforms. (4) Live captions — generated in real time by human stenographers (broadcast news, congressional hearings) or by AI-powered ASR (Zoom, Teams, Google Meet, YouTube live streams). (5) AI-generated captions — produced by ASR models like OpenAI Whisper Large-v3, achieving 90-95% accuracy on clean English audio; the fastest and most affordable production method for online video.
What's the difference between captions and subtitles?
Captions include dialogue plus non-speech audio (sound effects like [door slams], speaker IDs like SARAH:, music cues like [upbeat music]) and exist for viewers who cannot hear the audio. Subtitles include dialogue only and exist for viewers who can hear but don't understand the spoken language. Regional terminology: in the US, 'captions' typically means accessibility tracks and 'subtitles' means translation. In the UK, Ireland, Australia, and much of continental Europe, 'subtitles' is used for both purposes. For the full comparison, see captions vs subtitles. For the closed-caption-specific deep dive, see what is closed captioning.
Who uses captions besides deaf and hard-of-hearing viewers?
Most caption users are not deaf or hard-of-hearing. A widely-cited Ofcom 2006 study found roughly 80% of UK caption users were hearing viewers. Modern caption users include: viewers watching muted on social media (autoplay), viewers in loud environments (gyms, public transport, open offices), non-native speakers using captions for comprehension, learners using captions for language study, viewers with attention or processing differences, viewers of accented or technical content who catch details more reliably by reading, and viewers watching in situations where audio would disturb others (bedtime, hospitals, libraries). Netflix, YouTube, and Amazon Prime report caption usage rates of 50-80% among general audiences.
Where do you see captions?
Nearly every modern video platform. YouTube (auto-captions plus creator-uploaded SRT/VTT). Netflix, Disney+, HBO Max, Amazon Prime Video (SDH tracks in multiple languages). TikTok, Instagram Reels, LinkedIn Video (in-app auto-captions plus increasingly burned-in creator captions). Zoom, Microsoft Teams, Google Meet (live AI captions for meetings). Broadcast TV in the US, UK, and EU (regulated by FCC / Ofcom / AVMSD, delivered as CEA-608/708 or DVB-Subtitle). Blu-ray and DVD (SDH tracks). Live captioning is now standard on major news broadcasts, sports, congressional hearings, and increasingly at live events and conferences via ASR displays.
How are captions created?
Four production pipelines in 2026, in rough order of accuracy vs cost. (1) Human transcription (Rev, 3PlayMedia, VITAC): 99%+ accuracy, ~$1.25-3.50 per minute of video, standard for broadcast and legal compliance. (2) AI + human review: 97-99% accuracy, ~$0.20-1.00 per minute; most streaming platform (Netflix, Disney+) SDH tracks use this workflow. (3) AI transcription (OpenAI Whisper, VexaScribe, Otter, Descript, Deepgram): 90-95% accuracy on clean English audio, ~$0.005-0.10 per minute, dominant for creator and podcast workflows. (4) Live captioning by stenographers or AI ASR (Zoom, YouTube live, broadcast news): real-time delivery with 3-5 second latency and 85-95% accuracy depending on speaker clarity and vocabulary difficulty.
Are captions legally required?
Depends on jurisdiction and context. In the US: FCC rules (47 CFR §79.1) require captions on all broadcast, cable, and satellite TV. The CVAA (2010) extends captioning to online video previously aired on TV. ADA case law — NAD v. Netflix (2012) and NAD v. Harvard/MIT (2019-2020) — requires captioning of streaming and public educational content. Section 508 applies to federal agencies and federally-funded institutions. In the UK: the Communications Act 2003 and Ofcom Code set subtitling targets for broadcasters. In the EU: the Audiovisual Media Services Directive (AVMSD, revised 2018) and the European Accessibility Act (Directive 2019/882, effective 28 June 2025) require captioning for in-scope audiovisual services. WCAG 2.1 Level A requires captions on prerecorded video; Level AA requires captions on live video. Even where not strictly required by law, ADA case law is expanding and captioning is increasingly a business expectation.
What file formats are used for captions?
Several, each for different delivery contexts. SRT (SubRip) is the most universal — plain text, sequence numbers, HH:MM:SS,mmm timestamps; supported everywhere. WebVTT (.vtt) is the W3C HTML5 standard, required for the HTML5 <track> element, supports positioning and styling. SCC (Scenarist Closed Caption) is the binary CEA-608 format used in US broadcast delivery. TTML and IMSC1.1 are XML-based, used in professional OTT delivery (Netflix uses IMSC1.1). ASS/SSA supports advanced positioning and styling, popular in anime and fan-subtitle communities. CEA-608/708 are the actual on-air broadcast caption data formats decoded by TV sets. For creator workflows in 2026, SRT is the right default — convert to VTT for HTML5 web embedding, use SCC or TTML only for broadcast/OTT delivery.
What's the difference between captions and transcripts?
Captions are time-synced text that appears on screen at the moment the audio plays. Transcripts are plain-text records of what was said, without timing or synchronization to video. Transcripts include all content but display separately from the video (typically in a text panel, chapter markdown, or standalone document). Captions serve accessibility during video playback; transcripts serve searchability, SEO, and use cases where the reader is not watching the video (podcast show notes, meeting summaries, legal records). Most modern workflows produce both from the same source: a transcript with word-level timestamps generates SRT captions on export and TXT/DOCX transcripts on separate download. See what is audio transcription for the transcription-first perspective.
How accurate are AI-generated captions?
OpenAI Whisper Large-v3 (the current state-of-the-art open ASR model, used by VexaScribe and most modern captioning tools) achieves roughly 90-95% accuracy on clean English audio, measured as WER (word error rate) of 5-10% on standard benchmarks like LibriSpeech. Non-English major languages (Spanish, French, German, Japanese) range 85-93% depending on the language and audio quality. YouTube auto-captions run 85-92% on English, lower for accents and non-English. Human captioning services advertise 99%+ accuracy — the FCC does not specify a numeric accuracy requirement for broadcast, only that captions 'match the spoken words to the fullest extent possible.' For broadcast, legal, or medical use, AI + human review is the professional standard. For creator and podcast workflows, AI-only is typically acceptable after a quick review pass for proper nouns and technical terms.
Generate captions for your video — free 30 minutes
Upload your audio or video, get captions in 99 languages with word-level timestamps. Export as SRT, VTT, or burned-in MP4. Whisper Large-v3 accuracy. No credit card.