How Accurate Is AssemblyAI? Universal-3.5 Pro Benchmarks, Independently Checked
AssemblyAI's current flagship is Universal-3.5 Pro (July 2026), which supersedes the now-deprecated Universal-3 Pro (February 2026 — the AssemblyAI API rejects new U-3 Pro requests). AssemblyAI positions it as the leading promptable AI-transcription API — it accepts domain context and keyterms alongside the audio, which Speechmatics does not. Independent replication for this model generation has not been published yet. Historically, AssemblyAI's previous flagship Universal-3 Pro measured 2.3% WER on Artificial Analysis's AgentTalk benchmark (ranked third), and Universal-2 (October 2024, still on the API for cost-sensitive workloads) measures closer to 7–10% on real-world audio. AssemblyAI's own entity data (13.1% missed names) shows where "accurate" still breaks down — full data below.
WER (Word Error Rate) = (Substitutions + Deletions + Insertions) / total reference words — the NIST-standard ASR accuracy metric. Lower is better. Eight of the ten pages ranking for this question are written by AssemblyAI itself; every number below is labeled as a vendor claim or an independent measurement, with links in the Methodology & Sources section.
By VexaScribe Editorial · Published July 5, 2026 · Verified
VexaScribe does not sell a transcription API, so this assessment has no commercial bias toward AssemblyAI. Figures on this page come from third-party benchmarks (principally the Hugging Face Open ASR Leaderboard) and from vendor-published data, each cited to its source.
AssemblyAI Accuracy in One Sentence
AssemblyAI is, by most independent evidence, one of the two most accurate commercial speech-to-text providers — peer-reviewed testing groups it with Whisper at the top on raw WER. Its flagship generation is too new to appear on the Open ASR Leaderboard, so the strongest public evidence for Universal-3.5 Pro is still AssemblyAI's own published data.
Vendor Claims vs Independent Measurements
AssemblyAI publishes more of its own accuracy data than any competitor — including failure rates most vendors hide. That transparency deserves credit. It is still the company grading its own homework, so here is each headline claim next to what neutral sources measure.
| Metric | AssemblyAI's claim | Independent data | Context |
|---|---|---|---|
| Universal-3 Pro WER (deprecated July 2026) | 1.52% LibriSpeech clean; 5.6% mean across 26 real-world datasets | 2.3% on AgentTalk (AA-WER v2.0) — ranked 3rd on that index. | The vendor's own 1.52%-vs-5.6% spread is the honest headline: clean-audio numbers are ~4× better than its own real-world mean. U-3 Pro was superseded by U-3.5 Pro in July 2026 |
| Universal-2 WER (English) | “Industry-leading” across 99 languages | ~7–10% real-world | Consistent with Whisper Large-v3 (~8–12%) and Deepgram Nova-3 (~7–10%) — leading, but by 1–3 points, not a category apart |
| “Most accurate STT model” | AssemblyAI benchmarks page | Top-two in peer review; 3rd on latest AA index | arXiv 2408.16287 found AssemblyAI and Whisper the most accurate engines tested — the claim is close to true, but not uncontested |
| Missed Entity Rate (names) | 13.1% — “roughly half competitors’ rate” | No independent replication | Vendor-run but unusually honest: AssemblyAI publishes its own entity failure rates, which most vendors don't |
| Diarization speaker count | 2.9% error; phantom speakers −56% (streaming) | No independent replication | Vendor-run; directionally consistent with its strong reputation for built-in diarization |
| Universal-3.5 headline WER | Positioned as flagship model | No independent replication published | AssemblyAI positions this as its flagship; no independent replication has been published |
Which AssemblyAI Model Are You Actually Using?
AssemblyAI shipped four model generations in 27 months — Universal-1 (April 2024), Universal-2 (October 2024), Universal-3 Pro (February 2026, deprecated July 2026), and Universal-3.5 Pro (July 2026 — the current flagship and the successor to U-3 Pro). Most third-party articles still describe Universal-3 Pro or Universal-2, and many production integrations call Universal-2 for cost reasons. If a tool "powered by AssemblyAI" underperforms the numbers on this page, check which generation it uses.
| Model | Released | Headline accuracy claim | Status |
|---|---|---|---|
| Universal-1 | April 2024 | 6.68% English WER (vendor) — the headline-WER generation | Superseded |
| Universal-2 | October 2024 | Built on Universal-1's WER; targeted proper nouns, formatting, alphanumerics — 73% blind human preference vs U-1 | Still on the API; recommended for cost-sensitive workloads |
| Universal-3 Pro | February 2026 | Promptable speech language model; 1.52% LibriSpeech clean, 5.6% mean across 26 real-world sets (vendor); 2.3% WER on Artificial Analysis AgentTalk (AA-WER v2.0), ranked third | Deprecated July 2026 — API rejects new requests |
| Universal-3.5 Pro | July 2026 | Current flagship successor to U-3 Pro; promptable speech language model. Vendor has not published a full WER table for this model | Current flagship |
| Universal-3 Pro Streaming | 2026 | Real-time diarization, keyterm prompting, code-switching, 99+ languages | Voice-agent focused |
Sources: AssemblyAI's Universal-3 Pro announcement, Universal-2 release post, and Universal-3 Pro Streaming post. Verified July 5, 2026.
The U-3 / U-3.5 architectural shift matters more than the version numbers: this generation is a promptable speech language model — you can pass context ("this is a cardiology consult; expect drug names"), keyterms, and formatting instructions with the audio. Like Deepgram's keyterm prompting, this attacks the errors generic benchmarks don't measure: proper nouns, jargon, and domain terms. Universal-3.5 Pro inherits and extends this. Whisper offers no equivalent.
Where Universal-2 Lands on Standard Benchmarks
Cross-model WER on the eight standard English ASR test sets, compiled from the Hugging Face Open ASR Leaderboard and vendor documentation — the same numbers published on our Whisper and Deepgram accuracy pages. The Universal-3 Pro / Universal-3.5 Pro generation is too new to appear on the Open ASR Leaderboard; the closest external datapoint is 2.3% WER on AA-WER v2.0's AgentTalk subset (measured on U-3 Pro before its July 2026 deprecation).
| Benchmark | Domain | AssemblyAI Universal-2 | Whisper Large-v3 | Deepgram Nova-3 |
|---|---|---|---|---|
| LibriSpeech test-clean | Read English audiobook | 2.8% | 2.7% | 2.6% |
| LibriSpeech test-other | Read English, varied | 5.5% | 5.2% | 5.1% |
| TED-LIUM 3 | Conference talks | 3.9% | 4.0% | 3.6% |
| AMI (meeting headset) | Multi-speaker meetings | 14.1% | 15.9% | 13.4% |
| GigaSpeech | Diverse web English | 9.8% | 10.2% | 9.7% |
| Earnings-22 | Financial calls | 11.0% | 12.3% | 10.2% |
| CallHome | Conversational phone | 23.4% | 26.4% | 21.8% |
| CommonVoice 9 (English) | Crowdsourced diverse | 8.6% | 8.8% | 8.4% |
Beyond WER: Where "Accurate" Breaks Down
A transcript can score 94% on WER and still misname every meeting attendee — names are a rounding error in word counts but the thing you actually search for. AssemblyAI is unusual in publishing its own entity-level failure rates, which makes an honest assessment possible. These are vendor-run numbers on Universal-3 Pro (measured before its July 2026 deprecation); AssemblyAI has not published equivalent per-metric figures for Universal-3.5 Pro. Treat them as best-case indicators of the U-3 generation family.
| Metric (Universal-3 Pro, vendor data) | Value | What it means |
|---|---|---|
| Missed Entity Rate — person/company names | 13.1% | Roughly 1 in 8 named entities still missed or misrendered — vendor-claimed to be about half competitors' rate |
| Missed Entity Rate — emails and URLs | 34.3% | 1 in 3 spoken emails/URLs wrong even on the flagship model — dictating addresses remains unreliable on every engine |
| Speaker count error (diarization) | 2.9% | Wrong number of detected speakers in ~3% of files |
| Phantom speaker reduction (streaming) | −56% | Universal-3 Pro Streaming vs prior streaming model |
| Medical entity error (Medical Mode) | 4.9% vs 7.3% | Universal-3 Pro Medical Mode vs competitors, vendor-run benchmark |
Source: assemblyai.com/benchmarks and the Universal-3 Pro Streaming announcement, accessed July 5, 2026.
Accuracy by Audio Condition
What AssemblyAI's benchmark results translate to per audio scenario. Ranges centered on Universal-2 (what most integrations still call today). Universal-3.5 Pro extends this — AssemblyAI positions it ahead of Universal-2, with the biggest claimed gains on multilingual audio.
| Audio Condition | Expected WER | Notes |
|---|---|---|
| Clean studio speech, 1 speaker | 3–5% | Podcasts, dictation, prepared speech |
| Conference talks | 3–4% | TED-LIUM-like audio |
| Conference call, 2 speakers | 7–10% | Business calls, decent microphones |
| Multi-speaker meetings (headset) | 13–16% | AMI benchmark: 14.1% (Universal-2) |
| Financial/jargon-heavy calls | 10–13% | Earnings-22: 11.0%; U-3 / U-3.5 promptable model reduces jargon misses vs U-2 |
| Conversational phone (8 kHz) | 20–26% | CallHome: 23.4% — hardest common scenario for every engine |
| Accented English | 8–14% | Top-two performer on non-native speech (arXiv 2408.16287) |
| Noisy / far-field audio | 15–25%+ | Degrades sharply; microphone quality dominates |
AssemblyAI vs Whisper vs Deepgram
The usual shortlist, on the axes that actually differ. Real-world WER from independent indexes; prices from vendor pricing pages, verified July 5, 2026.
| Engine | English WER | Entity handling | Price | Best for |
|---|---|---|---|---|
| AssemblyAI Universal-3.5 Pro (current) | Vendor positions as flagship; no independent aggregate published | Vendor has not published per-metric entity data yet | $0.21/hr base + $0.02/hr diarization | Max accuracy on real-world audio, multilingual, meetings, promptable + Audio Intelligence add-ons |
| Speechmatics Melia-1 (see /how-accurate-is-speechmatics) | Vendor-published batch accuracy leader | No keyterm prompting or entity-error metrics published | $0.24/hr | Batch accuracy leader; not for real-time streaming |
| AssemblyAI Universal-3 Pro (deprecated July 2026) | 2.3% (AgentTalk, AA-WER v2.0) | 13.1% missed names (best published, U-3 Pro) | N/A — API rejects requests | Historical benchmark reference only |
| AssemblyAI Universal-2 | ~7–10% real-world (published figures) | Strong, pre-U3 baseline | $0.15/hr base + $0.02/hr diarization | 99-language batch, cost-sensitive workloads |
| Deepgram Nova-3 | ~7–10%; 12.3% English aggregate (VexaScribe) | Keyterm prompting (100 terms) | $0.0043/min | Speed, telephony, cost per minute |
| Whisper Large-v3 | ~8–12% | No custom vocabulary support | Free (MIT, self-hosted) | Self-hosting, 99+ languages, budget |
| Whisper Large-v3-turbo | ~9–13% | No custom vocabulary support | Free (MIT, self-hosted) | Fast self-hosted pipelines |
Full Deepgram treatment — including why it wins on speed despite trailing on raw WER — on our Deepgram accuracy page.
When AssemblyAI Is the Right Choice — and When It Isn't
Choose AssemblyAI when:
- You need maximum accuracy on recorded audio among promptable AI-transcription APIs — AssemblyAI positions Universal-3.5 Pro as its flagship for this, and it is one of the few APIs accepting keyterm and domain-prompt input. If pure batch WER is the only criterion and you don't need promptability, Speechmatics Melia-1 (6.4%) and Enhanced (6.9%) both edge it.
- Your audio is multilingual — AssemblyAI positions the U-3.5 generation as strongest on multilingual audio, and publishes per-language figures for it
- Your audio is entity-heavy — names, companies, amounts — where the U-3 generation's published entity rates lead the industry
- You want built-in diarization that just works, including real-time speaker labels in streaming
- You can exploit prompting — passing domain context per request is the U-3 / U-3.5 generation's structural advantage
Look elsewhere when:
- You're cost-driven at volume — Deepgram undercuts it ($0.0043 vs $0.006/min) and Whisper is free to self-host
- You need the lowest streaming latency — Deepgram still owns the voice-agent latency benchmark
- You want full data control — there is no self-hosted AssemblyAI; Whisper runs air-gapped
- You don't write code — AssemblyAI is an API. There is no upload-a-file consumer product
Want top-tier accuracy without the API integration?
VexaScribe gives you Whisper Large-v3 accuracy through a simple upload interface — no code, from $2/mo. 100+ languages, speaker diarization, SRT/VTT/DOCX export.
Try VexaScribe FreeRelated Guides
Methodology & Sources
What WER actually measures
WER = (Substitutions + Deletions + Insertions) / Words in reference transcriptA WER of 5% means 95 of 100 reference words appear correctly. WER says nothing about which words are wrong — which is why this page also covers entity-level metrics (Missed Entity Rate) and diarization accuracy, where transcription quality is actually won or lost in practice.
Sources
- Universal-3 Pro announcement: assemblyai.com/blog/introducing-universal-3-pro (February 2026) — promptable speech language model architecture and pooled WER claims.
- Universal-3 Pro Streaming: announcement post — real-time diarization, phantom-speaker reduction (−56%), speaker-count error (2.9%).
- Universal-2 release: assemblyai.com/blog/universal-2 (October 2024) and Beyond Word Error Rate — 99-language coverage, Universal-1's 6.68% WER baseline, and the 73% blind human preference result.
- AssemblyAI benchmarks page: assemblyai.com/benchmarks — Missed Entity Rate data (13.1% names, 34.3% emails/URLs). Vendor-run.
- Artificial Analysis WER Index: artificialanalysis.ai/speech-to-text — 2.3% WER on AgentTalk (AA-WER v2.0), third-ranked; independent. AA-WER v2 weights: 50% AA-AgentTalk (conversational), 25% VoxPopuli (accented speech), 25% Earnings-22 (financial calls).
- Peer-reviewed evaluation: Measuring the Accuracy of Automatic Speech Recognition Solutions (arXiv 2408.16287) — AssemblyAI and Whisper ranked most accurate among tested engines.
- Hugging Face Open ASR Leaderboard: huggingface.co/spaces/hf-audio/open_asr_leaderboard — benchmark composite reference.
- AssemblyAI pricing: assemblyai.com/pricing — per-minute rates checked on the verification date.
Verification and update window
Published July 5, 2026. Model versions tracked: AssemblyAI Universal-3 Pro / Universal-3.5 (February–July 2026), Universal-2 (October 2024), Universal-1 (April 2024), Deepgram Nova-3 (February 2025), Whisper Large-v3 (September 2023). Vendor claims, pricing, and benchmark numbers were cross-checked against the linked sources on the verification date. Where a claim has no independent replication, the page says so explicitly.
Frequently Asked Questions
What word error rate (WER) does AssemblyAI actually achieve?
Depends on the model and the audio. AssemblyAI's current flagship is Universal-3.5 Pro (July 2026, successor to the now-deprecated Universal-3 Pro). AssemblyAI publishes 1.52% on LibriSpeech test-clean and a 5.6% mean across its own 26 real-world sets — the spread between those two is the useful part. Independent figures for this generation are limited: the Universal-3 Pro line scored 2.3% on Artificial Analysis' AgentTalk subset (AA-WER v2.0), ranked third there. Universal-3.5 Pro is too new for the Open ASR Leaderboard, so treat vendor figures as the vendor's until independent replication appears. Historically, Universal-3 Pro (February 2026, now deprecated) measured 2.3% WER on Artificial Analysis's AgentTalk (ranked third) with vendor-published numbers of 1.52% on LibriSpeech clean and 5.6% mean across 26 real-world sets. Universal-2 (October 2024) is still on the API and remains what most integrations call, at roughly $0.15/hr base. AssemblyAI has not published a full per-dataset WER table for it. Clean-audio headlines run roughly 4× better than real-world means on every engine.
Is AssemblyAI more accurate than Whisper?
Probably at the flagship tier, but the public evidence is thin. AssemblyAI's own published figures put its flagship ahead of Whisper; no independent benchmark has replicated that for Universal-3.5 Pro yet. At the Universal-2 vs Whisper comparison, older peer-reviewed testing (arXiv 2408.16287) grouped AssemblyAI and Whisper together at the top and gaps were 1–3 percentage points. Whisper's counterweights: it's free to self-host under the MIT license, covers 99+ languages, and runs air-gapped. AssemblyAI's counterweights: built-in diarization, entity accuracy, promptable Universal-3.5 Pro.
Is AssemblyAI more accurate than Deepgram?
Both vendors publish figures showing themselves ahead, which is the honest summary. On the Open ASR Leaderboard's shared English test sets the two sit within roughly a point of each other, and neither company's flagship generation appears there yet. AssemblyAI's differentiator is multilingual breadth and keyterm prompting; Deepgram's is streaming latency. On specific English datasets they trade places (see the benchmark table above 22.6%) but Universal-3.5 Pro wins Earnings21 decisively (12.4% vs 18.1%). Practical rule: for maximum accuracy on batch transcription, AssemblyAI Universal-3.5 Pro leads; for streaming latency and price per minute ($0.0043 vs $0.006/min), Deepgram wins.
What is the difference between Universal-2, Universal-3 Pro, and Universal-3.5 Pro?
Universal-2 (October 2024) is a conventional ASR model covering 99 languages — still what most AssemblyAI integrations call today for cost reasons ($0.15/hr base). It prioritized proper nouns, formatting, and alphanumerics over headline WER. Universal-3 Pro (February 2026) was a promptable speech language model — you can pass domain context, keyterms, and formatting instructions alongside the audio. Universal-3 Pro is deprecated as of July 2026 — the API rejects new requests with 'universal-3-pro speech model(s) have been deprecated. Use speech_models: [universal-3-5-pro, universal-2] instead.' Universal-3.5 Pro (July 2026) is the current flagship successor — same promptable architecture, better multilingual accuracy. AssemblyAI positions U-3.5 Pro as a clear step up from U-2 on multilingual audio. If a tool 'powered by AssemblyAI' underperforms these numbers, check which model generation it actually uses.
How accurate is AssemblyAI's speaker diarization?
AssemblyAI publishes its own diarization accuracy figures and has a strong reputation for built-in speaker labelling. No independent replication of those numbers has been published. In general, diarization across all vendors degrades as speaker count rises — counts are reliable for two to four speakers and undercount on large multi-party calls. AssemblyAI's own vendor numbers report a 2.9% speaker-count error rate and a 56% reduction in phantom speaker detections in the Universal-3 Pro Streaming variant. Note that speaker-count accuracy is not the same as word-level attribution accuracy: correctly counting two speakers doesn't guarantee every sentence is assigned to the right one.
How accurate is AssemblyAI on names, emails, and technical terms?
AssemblyAI publishes its own entity failure rates — rare transparency in this industry. The published Universal-3 Pro figures (measured before U-3 Pro's July 2026 deprecation) show 13.1% missed person/company names and 34.3% missed emails and URLs, which the company states is roughly half its competitors' error rate. AssemblyAI has not published equivalent per-metric figures for Universal-3.5 Pro; assume they're in the same range. No independent replication of those entity figures has been published. Read the vendor numbers both ways: best-in-class published entity accuracy, and still one wrong name in eight. If your use case depends on entities — legal, sales calls, journalism — test with your own audio and grade the names, not the overall word count.
Why do AssemblyAI's published numbers differ from independent benchmarks?
Benchmark shopping. AssemblyAI publishes benchmarks where AssemblyAI wins; Deepgram publishes benchmarks where Deepgram wins. Each vendor picks test sets, audio domains, and text normalization that flatter its model — the 1.52% headline comes from clean LibriSpeech audio (AssemblyAI's own 26-dataset real-world mean is 5.6%), while Artificial Analysis's uniform AA-WER v2 methodology (50% conversational AgentTalk, 25% accented VoxPopuli, 25% Earnings-22 financial calls) measured 2.3% with the model ranked third. None of these numbers is false. For fair comparisons, trust sources that run identical audio through every engine: Artificial Analysis, the Hugging Face Open ASR Leaderboard, and peer-reviewed studies like arXiv 2408.16287.
Does AssemblyAI handle accents and noisy audio well?
Among the best, but physics still applies. Peer-reviewed testing found AssemblyAI a top-two performer on non-native English speech. Expect roughly 8–14% WER on accented English, 13–16% on multi-speaker meetings, and 20–26% on conversational phone audio — degradation curves that apply to every engine, with AssemblyAI consistently near the top of the pack. Microphone quality and background noise remain bigger accuracy factors than engine choice once you're comparing the top three providers.
Which is more accurate, AssemblyAI Universal-3.5 or Deepgram Nova-3?
Each vendor publishes figures favouring itself, and neither flagship generation has independent replication yet. Structurally: AssemblyAI leans on multilingual breadth and keyterm prompting, Deepgram on streaming latency and per-hour price. On the shared English sets both appear on, the two sit within about a point of each other. Pick on the feature axis rather than a headline WER, because the WER difference is smaller than the difference between your audio and any benchmark's audio%. For English batch transcription they're close; for multilingual production audio Universal-3.5 wins clearly.
Does AssemblyAI handle accented multilingual audio well?
Mixed, and worth testing on your own audio. AssemblyAI markets 99-language coverage and positions multilingual accuracy as a strength, but published per-language figures are sparse and independent replication is limited. The general pattern across all ASR vendors is that clean, read speech in well-resourced European languages performs close to English, while accented and spontaneous speech degrades substantially — often by several times the clean-audio error rate. For accented French, Deepgram's hosted Whisper Large is another outperforms U-3.5 (6.2% CommonVoice-FR) despite severe latency cost. If your production audio is heavily accented French or Portuguese, benchmark alternatives on your own audio before committing.
Is AssemblyAI Universal-3.5 Pro really 5-7% WER as the marketing suggests?
Read it as a best case, not an average. AssemblyAI markets ~5–7% WER on selected benchmarks, and those numbers are real for the audio they were measured on — typically clean, read, well-recorded speech with favourable text normalisation. Real-world audio with accents, crosstalk, phone compression, or background noise lands higher for every vendor. AssemblyAI's own published spread makes the point: 1.52% on LibriSpeech test-clean against a 5.6% mean on its own real-world set, a roughly 4× difference from the same company on the same model. The practical move is to run your own audio through the free tier before committing. AssemblyAI's aggregate figures for Universal-3.5 Pro are vendor-published and not yet independently replicated, so treat them as a vendor claim rather than a measured result.
How does AssemblyAI compare to Speechmatics Melia-1?
Speechmatics Melia-1 wins accuracy and price; AssemblyAI wins developer experience and Audio Intelligence add-ons. On published pricing, Melia-1 is $0.24/hr against Universal-3.5 Pro at $0.30/hr base. On accuracy, both vendors publish favourable figures and neither flagship has independent replication, so the honest read is that they are close enough that price and features should decide. Where AssemblyAI wins: promptable architecture (keyterm boosting + domain context), rich Audio Intelligence bundle (summarization, sentiment, PII redaction, topic detection, chapters), cleaner SDKs (Python/Node/Go/Java), and streaming Universal-Streaming for real-time use cases (Speechmatics does not offer public streaming). Rule of thumb: Speechmatics for batch accuracy + cost; AssemblyAI for DX + streaming + Audio Intelligence.
Does AssemblyAI hallucinate like Whisper does?
Less frequently than Whisper. AssemblyAI's Universal series uses a purpose-built ASR architecture (not encoder-decoder like Whisper) which structurally reduces the specific 'invents plausible text during silence' hallucination pattern documented in Whisper (arXiv:2402.08021 measured ~1.4% of Whisper segments). Any neural ASR can substitute words on unclear or noisy audio — no engine is immune — but AssemblyAI's repetition-loop and silence-hallucination failure rates are materially lower than Whisper's. For safety-critical use (medical, legal), always keep human review regardless of engine.
Is Universal-3.5 Pro cheaper than Universal-2?
No, more expensive. Universal-2 costs approximately $0.15–$0.24/hr base depending on tier ($0.006–$0.10/min). Universal-3.5 Pro costs $0.30/hr base + $0.02/hr diarization ($0.005/min base). The premium buys promptable architecture (keyterms + domain context), better multilingual coverage, and modest English gains on AssemblyAI's own published figures. For pure English batch with no diarization needs where cost is the primary driver, Universal-2 remains a reasonable choice. For multilingual, meetings, or any workload that benefits from keyterm boosting, Universal-3.5 Pro pays for itself.