Speech to Text — dictate live, or upload a recording

Talk and watch your words appear, or upload a recording and get a transcript with speaker labels and timestamps. Free either way, no sign-up to start.

Start dictating — free, no sign-up

Click the microphone and talk. Your words appear as you speak. Nothing is uploaded to VexaScribe and nothing is stored — your browser does the recognition and the text stays on this page until you copy or clear it.

Your text will appear here as you speak.

Works in Chrome, Edge, and Safari.

Have a recording instead? Upload an audio file

Which one do you need?

If you are about to talk, dictate: your words appear as you speak, free, with no account, and nothing to upload. If you already have a recording, upload it instead: you get speaker labels, timestamps and a downloadable transcript. Dictation handles one voice live; uploads handle several, after the fact.

Dictate

  • Instant — text appears as you speak
  • No upload, no account, no wait
  • One speaker
  • 18 languages
  • Needs Chrome, Edge or Safari

Upload a recording

  • Speaker labels and timestamps
  • Downloadable transcript file
  • Several speakers, separated
  • 99 languages, detected automatically
  • Works in any browser

The trade is speed against structure. Dictation is the fastest route from a thought to written words, but it produces one undifferentiated block of text from one voice. An upload takes a minute or two and gives you something with shape — who said what, and when they said it.

How dictation works here

  1. 1

    Click the microphone

    Your browser asks for microphone permission once. Nothing installs, and nothing is sent to VexaScribe.

  2. 2

    Talk normally

    Words appear as you speak. Say "period", "comma" or "new line" for punctuation, or add it afterwards in the box.

  3. 3

    Copy or download

    Take the text as a copy-paste or a .txt file. Closing the tab discards it — nothing is stored.

Dictation runs for as long as you keep talking. Long silences can end a session in some browsers — if the text stops appearing, click the microphone again and carry on. Nothing already dictated is lost.

What you get from an uploaded recording

Dictation gives you a block of text. An upload gives you a structured transcript: every line carries a timestamp, and each speaker is separated and labelled.

TranscriptEN · 2 speakers
  1. 00:04Speaker 1Can you walk me through how you first got into the field?
  2. 00:09Speaker 2Almost by accident, honestly. I was doing something else entirely.
  3. 00:16Speaker 1What changed?

Speakers are separated automatically, and renaming one in the editor applies to every line that speaker has. Exports include TXT, DOCX, SRT, VTT, CSV and JSON — the last carrying a start time, end time and confidence score for every individual word.

The free preview covers the first five minutes of any file up to 200 MB with no account. Transcribing audio files covers the upload path in full, including measured accuracy by language.

Where your voice actually goes

Most dictation tools say “private, in your browser” and leave it there. The honest answer is more specific, and it depends on which browser you are using.

The dictation tab

Nothing reaches VexaScribe and nothing is stored — closing or clearing the tab discards the text. Recognition is handled by your browser, and where that happens varies: Chrome and Edge send audio to Google's servers, while Safari processes many languages on-device. If you need a guarantee that audio never leaves your machine, use the dictation built into macOS or Windows, or run Whisper locally.

The upload tab

Files go to us over TLS 1.2+ and are stored encrypted at rest. We do not train models on your audio and we do not sell user data. Files can be deleted from the dashboard at any time, and account deletion is self-serve. A preview run without an account is not attached to any profile.

If a recording is confidential enough that this matters, the safest option is the one that never touches a network: your operating system's own dictation for live speech, or a locally-run model for files. We would rather say that than pretend otherwise.

Which browsers support voice typing

Live dictation uses the Web Speech API, which is not implemented everywhere. The upload tab has no such requirement and works in any browser.

BrowserDictationNotes
Chrome (desktop)WorksAudio is sent to Google's servers for recognition
EdgeWorksChromium-based, same engine and same behaviour
Safari (macOS & iOS)WorksApple processes many languages on-device
Chrome (Android)WorksGboard's own mic is often smoother on mobile
FirefoxNot supportedNo Web Speech API — use the upload tab instead

Firefox has never shipped the Web Speech API for recognition. If that is your browser, record on your phone and use the upload tab — or use Firefox's operating-system dictation, which works in any text field.

Accuracy: talking versus recording

Dictating usually beats transcribing, and the reason is not the software. When you dictate you are speaking clearly, close to a microphone, in a room you chose. A recording is whatever the room gave you.

What breaks dictation

  • Speaking faster than you would to a person — the engine needs phrase boundaries to work with.
  • Unusual proper nouns, which come back as the nearest common word.
  • Background noise the browser cannot gate out, especially other voices.
  • Long pauses, which end the session in some browsers mid-thought.

What breaks a recording

  • Distance from the microphone — a phone across the table picks up as much room as speech.
  • Crosstalk, where two people overlap and attribution lands on the wrong turn.
  • Strong accents against a language with thinner training coverage.
  • Technical vocabulary and acronyms, which sit outside everyday language.
  • Music underneath speech, which shares the same frequency range as the voice.

For uploads we publish measured word error rates by language and audio condition rather than a single marketing number — clear audio, noisy interviews and heavily accented speech land in different places, and one figure would not survive contact with your recording. See how accurate transcription is for the full breakdown.

Languages

The two tabs use different engines, so they support different numbers of languages. Claiming one number for both would be wrong.

18

for live dictation

English (US and UK), Spanish (Spain and Mexico), French, German, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Japanese, Korean, Mandarin, Hindi, Arabic and Indonesian. Pick yours before you start talking.

99

for uploaded files

Detected automatically from the audio — you do not choose. The wider coverage is why a recording in a less common language is better uploaded than dictated.

The dictation list is shorter because the browser's recognition engine is a different system with narrower coverage than the model behind the upload path.

Find your file type

Every format below uploads here directly — no converting first. If you are not sure what you have, the descriptions should place it.

  • M4AWhat the iPhone Voice Memos app records, and what most Apple devices hand you.
  • MP3The default almost everywhere else — podcast downloads, dictaphones, most exports.
  • WAVWhat a field recorder or a DAW produces. Uncompressed, large, best quality.
  • OGGCommon from open-source tools and some messaging apps.
  • MP4 and videoAny video file — the audio track is extracted for you.
  • iPhone voice memosGetting the file off the phone is the awkward part; that page covers it.
  • VoicemailCarrier-specific export steps, which differ more than you would expect.

Find your situation

The tool is the same; what differs is the workflow around it and what to watch for.

  • A meeting recordingZoom, Meet and Teams all export a file you can upload. Speaker labels matter most here.
  • An interviewTwo speakers, quotes you may publish — the case where checking attribution pays off.
  • A lecture or classLong, single-speaker, often technical vocabulary. Usually the easiest kind of audio.
  • A podcast episodeShow notes, chapter markers, and turning one episode into a written post.
  • Research interviewsCoding in NVivo, ATLAS.ti or MAXQDA, where the export format decides your workflow.
  • Writing in Google DocsDocs has its own built-in dictation. That page compares it with this one.

Dictation for dyslexia and writing difficulty

Speaking and spelling are separate skills. For writers with dyslexia, dysgraphia, or anyone who finds a blank page harder than a conversation, dictation removes the step that causes the trouble — you say the sentence, and the spelling is handled.

It also changes what a first draft costs. Talking through an idea is faster than typing it for most people, and the resulting text tends to be more direct, because spoken sentences are shorter than written ones.

Two practical notes: dictate in whole sentences rather than word by word, since the engine uses surrounding words to choose between homophones. And expect to edit punctuation afterwards — say “period” and “comma” if you prefer, but many people find it faster to talk freely and punctuate at the end.

Dictating on your own device

Your operating system already has dictation, and for short notes it is often the better choice — it works in any text field, not just a web page.

Windows

Press Windows key + H in any text field. The first use downloads a speech pack; after that it runs on-device.

macOS

Press the microphone key, or Control twice, with Dictation enabled in System Settings. On Apple Silicon it runs on-device for supported languages.

iPhone and Android

Tap the microphone on the keyboard. Both handle automatic punctuation, and both run on-device on recent models. On a phone this is usually smoother than a browser tab.

What this page adds on top: a wider language list than some OS defaults, no session time-outs, text you can copy or download as a file — and the upload tab, which system dictation cannot do at all.

Speech to text, or text to speech?

They are opposite directions, and the names are close enough that people land on the wrong one constantly.

Speech to text

Audio in, writing out. Talking or a recording becomes text you can read, edit and search. That is this page.

Text to speech

Writing in, audio out. You paste text and press play to hear it read aloud — what tools like NaturalReader and ElevenLabs do. Not this page.

What is actually free

Free means different things on the two tabs, so here are both, in numbers.

Dictation

Free with no account and no time limit. Your browser does the recognition, so it costs us nothing to run — which is what makes that sustainable rather than a trial in disguise.

Uploads

The first five minutes of any file up to 200 MB, with no account. A free account adds 30 minutes and files up to 5 GB, still with no card. Paid plans start at $2/month for 200 minutes — transcription costs real money to run, and someone pays for it either way.

See all plans.

Common questions

Is there a free speech to text?

Yes, two of them here. Browser dictation is free with no account and no time limit — your browser does the recognition, so it costs nothing to run. Uploading a recording is free for the first five minutes with no account, or 30 minutes with a free account and no card.

How do I turn on voice typing?

Click the microphone on the Dictate tab and allow microphone access when the browser asks. Words appear as you speak. Nothing installs, and you can copy the text or download it as a .txt file when you are done.

Why doesn't dictation work in Firefox?

Firefox has never implemented the Web Speech API that browser dictation depends on. Chrome, Edge and Safari all support it. In Firefox, either use the upload tab — which works everywhere — or your operating system's own dictation, which works in any text field.

Is speech to text accurate?

Dictating into a close microphone in a quiet room is the most accurate case, usually well above 95% on clear speech. Recorded audio varies far more: clear podcast-quality audio runs 3-6% word error rate, noisy interviews 8-15%, and heavily accented or jargon-dense speech 10-20%.

Does this record or store my voice?

The dictation box does not send anything to VexaScribe and does not store your text — closing or clearing the tab discards it. Your browser handles recognition, and where that happens depends on the browser: Chrome and Edge send audio to Google's servers, while Safari processes many languages on-device. If you need a guarantee that audio never leaves your machine, use macOS or Windows dictation, or run Whisper locally.

Can it tell two speakers apart?

On an uploaded recording, yes — speakers are separated automatically and labelled, and renaming one applies to every line that speaker has. Live dictation cannot: it treats everything the microphone hears as one voice, so a two-person conversation comes back as one block of text.

What's the difference between speech to text and transcription?

Mostly which end you are at. Speech to text is the general term for turning spoken words into writing, including live dictation. Transcription usually means doing it to a recording that already exists, and typically implies a structured result — speaker labels, timestamps, an exportable file.

What's the difference between speech to text and text to speech?

Opposite directions. Speech to text takes audio and gives you writing — that is this page. Text to speech takes writing and gives you audio, which is what tools like NaturalReader and ElevenLabs do. If you are about to talk, you want this page; if you are about to paste text in and press play, you want a text-to-speech tool.

How big a file can I upload?

Up to 200 MB with no account, and the first five minutes are transcribed. With a free account, up to 5 GB and 10 hours per file. For reference, an hour of typical MP3 audio is roughly 30-60 MB, so length is rarely the limit — uncompressed WAV files are what usually exceed it.

Does it add punctuation automatically?

On uploaded recordings, yes — sentences and paragraph breaks are added for you. In live dictation, punctuation is partly manual: say "period", "comma" or "new line" as you go, or talk freely and add it in the text box afterwards, which most people find faster.

Can I dictate from an audio file?

Not on the Dictate tab — that listens to your microphone live. Use the Upload tab instead: it takes 17 audio and video formats and returns a transcript with speaker labels and timestamps.