Home/Audio to SRT

Audio to SRT — Convert MP3, WAV, M4A to SRT Subtitle Files (Free)

Drop in MP3, WAV, M4A, AAC, OGG, FLAC, WMA, or Opus and get a standards-compliant SRT back — HH:MM:SS,mmm timecodes, UTF-8 encoding, ready for YouTube, Premiere, or any player. The first 5 minutes are free without an account.

Drop your file here

or click to browse

Audio: MP3 · WAV · M4A · FLAC · OGG · AAC · AIFF · WMA · AMR · OPUS

Video: MP4 · MOV · AVI · MKV · WebM · FLV · WMV

Up to 200 MB · 5 min free preview

The preview transcribes the first 5 minutes and the .srt downloads without an account · files up to 200 MB free, 5 GB with a free account

Speaker labels99 languagesNo account needed

By VexaScribe Editorial · Verified

TL;DR

VexaScribe converts MP3, WAV, M4A, AAC, OGG, FLAC, WMA, and Opus audio to SRT subtitle files. Three steps: upload, transcribe, export. A 1-hour file typically transcribes in about 2 minutes on the premium model, 3–5 on standard. Output is standards-compliant SRT with HH:MM:SS,mmm timecodes, UTF-8 encoding, ready for YouTube, Premiere Pro, DaVinci Resolve, or any player.

How to Convert Audio to SRT

End-to-end, a 1-hour audio file becomes a downloadable SRT in roughly 5 minutes of wall-clock time — most of which is upload, not transcription.

1

Upload the audio file

Drag-and-drop MP3, WAV, M4A, AAC, OGG, FLAC, WMA, or Opus into VexaScribe. Files up to 5GB and 4+ hours long are supported. If your audio lives at a public URL (podcast RSS enclosure, direct MP3 link, cloud storage share), paste the URL for the paste-URL flow. iPhone Voice Memos M4A, Zoom cloud recordings, Google Meet recordings — all upload directly with no conversion.

2

Run Whisper Large-v3 transcription

Pick the spoken language (auto-detect handles 100+ languages, but specifying explicitly is more reliable). Enable speaker diarization for multi-speaker audio. A 1-hour file typically completes in about 2 minutes on the premium model, 3–5 on standard. The dashboard shows progress — you don't need to keep the tab open.

3

Export the SRT

Click Export → SRT. The download is a plain-text .srt file with numbered cues, HH:MM:SS,mmm timecodes, UTF-8 encoding, and speaker prefixes if diarization was enabled. VTT, DOCX, and PDF exports available in the same dialog. Ready for YouTube Studio, Vimeo, LinkedIn Video, or import into Premiere Pro, DaVinci Resolve, or Final Cut Pro.

Supported Audio Formats

Whisper Large-v3 was trained on 680,000 hours of multilingual audio drawn from the open web, so essentially every consumer audio codec works. Upload the format you have — don't pre-convert.

FormatExtensionNotes
MP3.mp3Universal podcast format. Accepted at any bitrate.
WAV.wavUncompressed PCM. Highest fidelity, largest file size.
M4A.m4aAAC in an MP4 container. The default export for iPhone Voice Memos, Zoom, and Google Meet — uploads directly, no conversion.
AAC.aacRaw AAC audio stream. Common in podcast production and mobile recording apps.
OGG.oggVorbis in an Ogg container. Open-source and patent-free.
FLAC.flacFree Lossless Audio Codec. Compressed but bit-exact to the source PCM, so nothing is discarded.
WMA.wmaWindows Media Audio. Legacy Microsoft codec — accepted but rare in modern workflows.
Opus.opusModern low-latency codec used by Discord, WebRTC, and WhatsApp voice notes. Efficient at low bitrates.

Two things that do affect the result

Bitrate. Anything at 128 kbps or higher is fine, and the format you picked barely matters at that point. Below roughly 64 kbps, compression starts discarding detail that speech models rely on — usually surfacing as errors on quiet speech and accented voices. If you have a choice at export time, pick the higher bitrate.

Overlapping speech. Cross-talk is the hardest case for any diarization system, ours included. If you are recording a group, individual mics per speaker produce dramatically better speaker labels than one shared mic in the middle of a table.

Common Use Cases

Podcast episode → captioned social clips

Record the episode, export as MP3 or M4A, upload to VexaScribe, export SRT. The SRT doubles as source for Apple Podcasts / Spotify transcripts, your website transcript page, and burned-in captions for social video clips.

Interview / research audio → timed transcript

Journalists, academic researchers, and podcast interviewers use SRT to pull citable quotes with exact timing. Jump to 00:37:22 in the source to verify a quote — that's the point.

Zoom / Meet / Teams recording → captioned share-out

Zoom exports M4A audio (or MP4 video); Google Meet exports M4A; Teams exports MP4. Upload the audio with speaker diarization on, export the SRT, attach to the recording share link so remote colleagues watch with captions.

Voice memo / dictation → searchable notes

iPhone Voice Memos M4A files upload directly. Export SRT to keep the timestamped structure, or DOCX/TXT for a flat text version. Useful for meeting recap, fieldwork, and note-taking on the go.

Radio show / archived broadcast

Convert broadcast MP3 or WAV archives into SRT for accessibility, search indexing, or podcast repurposing. Diarization separates hosts, guests, and callers.

Oral history / academic research

Historians and qualitative researchers use SRT to anchor quotes to source audio for citation. UTF-8 output handles non-English interviewee names and place names cleanly.

Attaching Audio-Derived SRT to Video

An SRT is a timed subtitle file — useful only when it's paired with playback. Three common patterns:

1. Static-image podcast video. Wrap the podcast audio in a video container with a still image (episode artwork, waveform, or looping brand animation). Any video editor works — Premiere, DaVinci Resolve, CapCut, even FFmpeg. Import the SRT as a caption track or burn it into the video for social platforms that strip caption files (Instagram Reels, some TikTok flows).

2. Separately recorded video (talking-head or interview). If the audio was recorded on a dedicated device (Zoom H6, lav mic into recorder) alongside a video track, the SRT generated from the clean audio track will need slight alignment nudging against the video track — cameras and audio recorders drift by a few frames over long takes. Use your NLE's sync-by-waveform feature to lock the audio track to the video first, then import the SRT.

3. YouTube / Vimeo upload. No editing required — upload the video and SRT separately in the platform's subtitle section. YouTube Studio → Content → Subtitles → Add Language → Upload File. Vimeo Advanced settings during publish. This is the cleanest path for platforms that render SRT natively.

For the reverse case — starting with video and generating SRT directly from the video track — use video to SRT.

SRT vs Plain Transcript — Which Do You Actually Want?

Audio produces two kinds of output. Pick the one that matches your use case.

SRT (this page)

Numbered cues with start/end timecodes. Designed to overlay on video during playback. Use for: video captions, podcast video repurposing, jumping to exact quote timing, accessibility compliance for prerecorded video.

Plain transcript (DOCX / TXT / PDF)

Flat text with optional speaker labels but no per-word timing. Use for: reading, editing, LLM input, sharing as a document, blog posts derived from podcasts. See MP3 to text, WAV to text, M4A to text.

Format spec, working code sample, and encoding notes for SRT are covered at what is an SRT file.

When to Hire a Human

Legal depositions, courtroom transcripts, medical records, and regulated accessibility deliverables need human review or human authorship — mishearings on names, dosages, jurisdictions, and citations carry real consequences there, and a machine pass is a starting point rather than a deliverable. For a captioned podcast episode, a lecture recording, a marketing interview, or a YouTube upload, auto-generated SRT plus a quick review pass is the right tool. See AI vs human transcription for the full comparison.

Audio to SRT FAQ

How do I convert MP3 to SRT?

Upload the MP3 to VexaScribe (30 minutes free at signup), let it transcribe the audio (about 2 minutes for a 1-hour file on the premium model, 3–5 on standard), then click Export → SRT. The output is a plain-text .srt with numbered cues, HH:MM:SS,mmm timecodes, and UTF-8 encoding — ready for YouTube, Premiere Pro, DaVinci Resolve, or any player that accepts SRT.

How do I convert WAV to SRT?

Same three-step workflow as MP3. WAV is lossless PCM, so it's the highest-fidelity input format — but re-encoding an MP3 to WAV before uploading cannot restore detail the MP3 already discarded, so there is no reason to do it. Upload the file you have. VexaScribe accepts WAV up to 5GB per file.

How do I convert M4A to SRT?

Drag the M4A into VexaScribe. M4A is the default export for iPhone Voice Memos, Zoom cloud recordings, and Google Meet — all of them upload directly with no conversion. Transcription takes about 2 minutes per hour of audio on the premium model, 3–5 on standard. Export as SRT when done.

Can I convert audio to SRT for free?

Yes, up to 30 minutes on signup with no card required. Whisper Large-v3 (which powers VexaScribe) is MIT-licensed and free to self-host if you have Python and a GPU, but the DIY route lacks the cue-splitting, editor, and speaker labels. TurboScribe has a limited free tier with ads.

How accurate is audio-to-SRT?

It depends far more on the recording than on the file format. Clean single-speaker audio captured close to a decent microphone is the best case; accented speech, specialist vocabulary, background noise, and overlapping speakers each make it harder. We publish accuracy figures only where we have measured them — rather than quote a range for your recording sight-unseen, run the free preview on your own file and judge the output directly. The first 5 minutes need no account.

How do I attach an SRT file to a video after generating from audio?

Three options depending on your target. (1) Video already exists: import the SRT into Premiere Pro / DaVinci Resolve / Final Cut as a caption track and export the video with subtitles burned in or as a soft-sub track. (2) Static-image video (podcast episode): create a video with a still image using any editor, then attach the SRT. (3) YouTube / Vimeo: upload the video and SRT separately in the platform's subtitle section — no burning required.

What's the difference between audio-to-SRT and audio-to-text?

Audio-to-text produces a flat transcript with no timing — one long text file suitable for reading, editing, or feeding to an LLM. Audio-to-SRT produces a timed subtitle file — numbered cues with start/end timecodes designed to overlay on video during playback. Use SRT when the output needs to sync to video; use plain text when timing isn't required.

Can I convert a voice memo to SRT?

Yes. iPhone Voice Memos exports M4A by default, which uploads directly to VexaScribe. Same three-step workflow. Voice memos are usually single-speaker close-mic recordings — the easiest case for any speech model. Useful for turning field notes, dictation, or thinking-out-loud sessions into timed transcripts you can search later.

How long does it take?

Far less than real time. A 1-hour file typically processes in about 2 minutes on the premium model and 3–5 on standard. Upload usually takes longer than transcription does, depending on your connection; SRT export is instant. You don't need to keep the tab open — the dashboard shows progress and files stay in your workspace.

Convert your audio to SRT — free, no account ↑