Speech to Text — dictate live, or upload a recording
Talk and watch your words appear, or upload a recording and get a transcript with speaker labels and timestamps. Free either way, no sign-up to start.
Start dictating — free, no sign-up
Click the microphone and talk. Your words appear as you speak. Nothing is uploaded to VexaScribe and nothing is stored — your browser does the recognition and the text stays on this page until you copy or clear it.
Works in Chrome, Edge, and Safari.
Have a recording instead? Upload an audio file
Drop your audio file here
or click to browse
Audio: MP3 · WAV · M4A · FLAC · OGG · AAC · AIFF · WMA · AMR · OPUS
Video: MP4 · MOV · AVI · MKV · WebM · FLV · WMV
Up to 200 MB · first 5 minutes free, no account
Free 5-minute preview · no account · files up to 200 MB · sign up free for 30 minutes and 5 GB files
Which one do you need?
If you are about to talk, dictate: your words appear as you speak, free, with no account, and nothing to upload. If you already have a recording, upload it instead: you get speaker labels, timestamps and a downloadable transcript. Dictation handles one voice live; uploads handle several, after the fact.
Dictate
- Instant — text appears as you speak
- No upload, no account, no wait
- One speaker
- 18 languages
- Needs Chrome, Edge or Safari
Upload a recording
- Speaker labels and timestamps
- Downloadable transcript file
- Several speakers, separated
- 99 languages, detected automatically
- Works in any browser
The trade is speed against structure. Dictation is the fastest route from a thought to written words, but it produces one undifferentiated block of text from one voice. An upload takes a minute or two and gives you something with shape — who said what, and when they said it.
How dictation works here
- 1
Click the microphone
Your browser asks for microphone permission once. Nothing installs, and nothing is sent to VexaScribe.
- 2
Talk normally
Words appear as you speak. Say "period", "comma" or "new line" for punctuation, or add it afterwards in the box.
- 3
Copy or download
Take the text as a copy-paste or a .txt file. Closing the tab discards it — nothing is stored.
Dictation runs for as long as you keep talking. Long silences can end a session in some browsers — if the text stops appearing, click the microphone again and carry on. Nothing already dictated is lost.
What you get from an uploaded recording
Dictation gives you a block of text. An upload gives you a structured transcript: every line carries a timestamp, and each speaker is separated and labelled.
- 00:04Speaker 1Can you walk me through how you first got into the field?
- 00:09Speaker 2Almost by accident, honestly. I was doing something else entirely.
- 00:16Speaker 1What changed?
Speakers are separated automatically, and renaming one in the editor applies to every line that speaker has. Exports include TXT, DOCX, SRT, VTT, CSV and JSON — the last carrying a start time, end time and confidence score for every individual word.
The free preview covers the first five minutes of any file up to 200 MB with no account. Transcribing audio files covers the upload path in full, including measured accuracy by language.
Where your voice actually goes
Most dictation tools say “private, in your browser” and leave it there. The honest answer is more specific, and it depends on which browser you are using.
The dictation tab
Nothing reaches VexaScribe and nothing is stored — closing or clearing the tab discards the text. Recognition is handled by your browser, and where that happens varies: Chrome and Edge send audio to Google's servers, while Safari processes many languages on-device. If you need a guarantee that audio never leaves your machine, use the dictation built into macOS or Windows, or run Whisper locally.
The upload tab
Files go to us over TLS 1.2+ and are stored encrypted at rest. We do not train models on your audio and we do not sell user data. Files can be deleted from the dashboard at any time, and account deletion is self-serve. A preview run without an account is not attached to any profile.
If a recording is confidential enough that this matters, the safest option is the one that never touches a network: your operating system's own dictation for live speech, or a locally-run model for files. We would rather say that than pretend otherwise.
Which browsers support voice typing
Live dictation uses the Web Speech API, which is not implemented everywhere. The upload tab has no such requirement and works in any browser.
| Browser | Dictation | Notes |
|---|---|---|
| Chrome (desktop) | Works | Audio is sent to Google's servers for recognition |
| Edge | Works | Chromium-based, same engine and same behaviour |
| Safari (macOS & iOS) | Works | Apple processes many languages on-device |
| Chrome (Android) | Works | Gboard's own mic is often smoother on mobile |
| Firefox | Not supported | No Web Speech API — use the upload tab instead |
Firefox has never shipped the Web Speech API for recognition. If that is your browser, record on your phone and use the upload tab — or use Firefox's operating-system dictation, which works in any text field.
Accuracy: talking versus recording
Dictating usually beats transcribing, and the reason is not the software. When you dictate you are speaking clearly, close to a microphone, in a room you chose. A recording is whatever the room gave you.
What breaks dictation
- Speaking faster than you would to a person — the engine needs phrase boundaries to work with.
- Unusual proper nouns, which come back as the nearest common word.
- Background noise the browser cannot gate out, especially other voices.
- Long pauses, which end the session in some browsers mid-thought.
What breaks a recording
- Distance from the microphone — a phone across the table picks up as much room as speech.
- Crosstalk, where two people overlap and attribution lands on the wrong turn.
- Strong accents against a language with thinner training coverage.
- Technical vocabulary and acronyms, which sit outside everyday language.
- Music underneath speech, which shares the same frequency range as the voice.
For uploads we publish measured word error rates by language and audio condition rather than a single marketing number — clear audio, noisy interviews and heavily accented speech land in different places, and one figure would not survive contact with your recording. See how accurate transcription is for the full breakdown.
Languages
The two tabs use different engines, so they support different numbers of languages. Claiming one number for both would be wrong.
18
for live dictation
English (US and UK), Spanish (Spain and Mexico), French, German, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Japanese, Korean, Mandarin, Hindi, Arabic and Indonesian. Pick yours before you start talking.
99
for uploaded files
Detected automatically from the audio — you do not choose. The wider coverage is why a recording in a less common language is better uploaded than dictated.
The dictation list is shorter because the browser's recognition engine is a different system with narrower coverage than the model behind the upload path.
Find your file type
Every format below uploads here directly — no converting first. If you are not sure what you have, the descriptions should place it.
- M4AWhat the iPhone Voice Memos app records, and what most Apple devices hand you.
- MP3The default almost everywhere else — podcast downloads, dictaphones, most exports.
- WAVWhat a field recorder or a DAW produces. Uncompressed, large, best quality.
- OGGCommon from open-source tools and some messaging apps.
- MP4 and videoAny video file — the audio track is extracted for you.
- iPhone voice memosGetting the file off the phone is the awkward part; that page covers it.
- VoicemailCarrier-specific export steps, which differ more than you would expect.
Find your situation
The tool is the same; what differs is the workflow around it and what to watch for.
- A meeting recordingZoom, Meet and Teams all export a file you can upload. Speaker labels matter most here.
- An interviewTwo speakers, quotes you may publish — the case where checking attribution pays off.
- A lecture or classLong, single-speaker, often technical vocabulary. Usually the easiest kind of audio.
- A podcast episodeShow notes, chapter markers, and turning one episode into a written post.
- Research interviewsCoding in NVivo, ATLAS.ti or MAXQDA, where the export format decides your workflow.
- Writing in Google DocsDocs has its own built-in dictation. That page compares it with this one.
Dictation for dyslexia and writing difficulty
Speaking and spelling are separate skills. For writers with dyslexia, dysgraphia, or anyone who finds a blank page harder than a conversation, dictation removes the step that causes the trouble — you say the sentence, and the spelling is handled.
It also changes what a first draft costs. Talking through an idea is faster than typing it for most people, and the resulting text tends to be more direct, because spoken sentences are shorter than written ones.
Two practical notes: dictate in whole sentences rather than word by word, since the engine uses surrounding words to choose between homophones. And expect to edit punctuation afterwards — say “period” and “comma” if you prefer, but many people find it faster to talk freely and punctuate at the end.
Dictating on your own device
Your operating system already has dictation, and for short notes it is often the better choice — it works in any text field, not just a web page.
Windows
Press Windows key + H in any text field. The first use downloads a speech pack; after that it runs on-device.
macOS
Press the microphone key, or Control twice, with Dictation enabled in System Settings. On Apple Silicon it runs on-device for supported languages.
iPhone and Android
Tap the microphone on the keyboard. Both handle automatic punctuation, and both run on-device on recent models. On a phone this is usually smoother than a browser tab.
What this page adds on top: a wider language list than some OS defaults, no session time-outs, text you can copy or download as a file — and the upload tab, which system dictation cannot do at all.
Speech to text, or text to speech?
They are opposite directions, and the names are close enough that people land on the wrong one constantly.
Speech to text
Audio in, writing out. Talking or a recording becomes text you can read, edit and search. That is this page.
Text to speech
Writing in, audio out. You paste text and press play to hear it read aloud — what tools like NaturalReader and ElevenLabs do. Not this page.
What is actually free
Free means different things on the two tabs, so here are both, in numbers.
Dictation
Free with no account and no time limit. Your browser does the recognition, so it costs us nothing to run — which is what makes that sustainable rather than a trial in disguise.
Uploads
The first five minutes of any file up to 200 MB, with no account. A free account adds 30 minutes and files up to 5 GB, still with no card. Paid plans start at $2/month for 200 minutes — transcription costs real money to run, and someone pays for it either way.
Common questions
Is there a free speech to text?
Yes, two of them here. Browser dictation is free with no account and no time limit — your browser does the recognition, so it costs nothing to run. Uploading a recording is free for the first five minutes with no account, or 30 minutes with a free account and no card.
How do I turn on voice typing?
Click the microphone on the Dictate tab and allow microphone access when the browser asks. Words appear as you speak. Nothing installs, and you can copy the text or download it as a .txt file when you are done.
Why doesn't dictation work in Firefox?
Firefox has never implemented the Web Speech API that browser dictation depends on. Chrome, Edge and Safari all support it. In Firefox, either use the upload tab — which works everywhere — or your operating system's own dictation, which works in any text field.
Is speech to text accurate?
Dictating into a close microphone in a quiet room is the most accurate case, usually well above 95% on clear speech. Recorded audio varies far more: clear podcast-quality audio runs 3-6% word error rate, noisy interviews 8-15%, and heavily accented or jargon-dense speech 10-20%.
Does this record or store my voice?
The dictation box does not send anything to VexaScribe and does not store your text — closing or clearing the tab discards it. Your browser handles recognition, and where that happens depends on the browser: Chrome and Edge send audio to Google's servers, while Safari processes many languages on-device. If you need a guarantee that audio never leaves your machine, use macOS or Windows dictation, or run Whisper locally.
Can it tell two speakers apart?
On an uploaded recording, yes — speakers are separated automatically and labelled, and renaming one applies to every line that speaker has. Live dictation cannot: it treats everything the microphone hears as one voice, so a two-person conversation comes back as one block of text.
What's the difference between speech to text and transcription?
Mostly which end you are at. Speech to text is the general term for turning spoken words into writing, including live dictation. Transcription usually means doing it to a recording that already exists, and typically implies a structured result — speaker labels, timestamps, an exportable file.
What's the difference between speech to text and text to speech?
Opposite directions. Speech to text takes audio and gives you writing — that is this page. Text to speech takes writing and gives you audio, which is what tools like NaturalReader and ElevenLabs do. If you are about to talk, you want this page; if you are about to paste text in and press play, you want a text-to-speech tool.
How big a file can I upload?
Up to 200 MB with no account, and the first five minutes are transcribed. With a free account, up to 5 GB and 10 hours per file. For reference, an hour of typical MP3 audio is roughly 30-60 MB, so length is rarely the limit — uncompressed WAV files are what usually exceed it.
Does it add punctuation automatically?
On uploaded recordings, yes — sentences and paragraph breaks are added for you. In live dictation, punctuation is partly manual: say "period", "comma" or "new line" as you go, or talk freely and add it in the text box afterwards, which most people find faster.
Can I dictate from an audio file?
Not on the Dictate tab — that listens to your microphone live. Use the Upload tab instead: it takes 17 audio and video formats and returns a transcript with speaker labels and timestamps.