Transcript Summarizer — Summarize Any Transcript with AI
2. Run it on your own transcript
Add the transcript
One free summary per week — no account needed.
0 / 5,000
Or see a meeting that has already been run — transcript in, fields out.
What was said
Alex: Okay, Q4 planning. Priya, you were leading.
Priya: Right. Three items. Mobile app rewrite. Compliance work legal flagged in September. Hiring — two engineers by year-end, reqs aren't open yet.
Alex: Real quick — is the rewrite this-quarter, or just what we said last time?
Priya: What we said. I think we can push. But the current app is losing conversion — Marcus showed 12% drop last week — so I'd rather not.
Alex: Okay, tentatively commit but flag it as movable. Compliance?
Priya: Not movable. Legal said end of November or we're not shipping in the EU. We might need to pull two people off the rewrite to hit it. That's the tradeoff. I'll draft options by Friday, we decide Monday.
Sam: If we pull two people we're not making the mobile date.
Priya: Right. That's why I said tradeoff.
Alex: Decisions: Priya drafts by Friday, we meet Monday. Reqs go up this week regardless. Anything else?
Sam: One thing — the auth vendor renewal is December 3rd, we haven't reviewed alternatives. Might not matter, but I'd feel better if someone owned it.
Alex: You own it, Sam. Report back in two weeks.
What was extracted
Reading…
Highlighting shows roughly where each field came from — every summary also keeps timestamps back into the audio.
By VexaScribe Editorial · Published · Updated
Turn an audio recording, video, or a text transcript into a structured AI summary — pick from six purpose-built types (Meeting, Sales Call, Interview, Lecture, Podcast, General). Handles multi-hour recordings with map-reduce architecture, preserves sales-call deadlines verbatim, and does not use your data to train AI models. Free 30 minutes at signup, no card.
VexaScribe generates a structured AI summary from any audio, video, or text transcript. Upload an MP3, WAV, MP4, MOV, or any of 17 supported formats up to 5 GB — or paste an existing transcript directly. We transcribe with Whisper Large-v3 (90-95% accuracy on clear audio, 99 languages), then generate a structured summary in one of 6 purpose-built types: General, Meeting, Sales Call, Interview, Lecture, or Podcast. Most 60-minute files finish in about two minutes on our premium model, three to five on standard. Under the hood: map-reduce architecture with semantic chunking — embeddings detect natural topic shifts in long recordings, each section is summarized in parallel, then consolidated. Meeting summaries get a second, focused pass over the decisions: small models reliably catch an explicit “we agreed to ship on the 14th” but under-extract the implicit kind, where a decision is reached over several turns and never stated in one sentence. Multi-hour recordings preserve their structure instead of getting flattened to a paragraph — there is no duration cap, files run to 5 GB on any account, and 10-hour recordings go through. Summaries can be generated in any of the world's major languages regardless of source language. Free tier is 30 minutes at signup — no credit card, all 6 summary types included.
Why summarize transcripts at all?
Raw transcripts are searchable but unreadable. Summarization is what turns 90 minutes of audio into a 90-second decision — and the time savings are measurable.
Numbers verify what every knowledge worker already feels: meetings and unfiltered transcripts are a productivity tax. AI summaries are the recovery mechanism. See how we verify these stats.
Starting from a recording instead? This page takes text you already have. If you have the audio or video and no transcript yet, audio to notes transcribes and summarizes in one pass — or upload it in the tool above and the transcript lands in the box for you.
How VexaScribe AI summary works
Most AI summarizers feed the entire transcript into one call and accept whatever comes back. That works on a 5-minute voice memo and falls apart on a 90-minute meeting. VexaScribe uses a map-reduce pipeline with semantic chunking — so even multi-hour recordings preserve their structure instead of getting flattened to a paragraph.
- 1
Semantic chunking
For long transcripts, we use OpenAI embeddings to find topic boundaries in the conversation — measuring sentence-to-sentence similarity, then locating 'valleys' in the similarity curve that mark topic shifts. Chunks get cut at those natural boundaries, not at arbitrary length thresholds.
- 2
Map phase
Each chunk goes to a current-generation OpenAI model in parallel, extracting chapters and type-specific structured data. Overlap segments from the previous chunk are included so context isn't lost across boundaries.
- 3
Reduce phase
A final call consolidates all chunk results into one coherent summary — deduplicating action items, merging only chapters that are genuinely the same topic, and producing the final structured output for your chosen summary type.
Why this matters in practice: a 2-hour planning meeting summarized naively flattens 4 distinct decisions into one vague paragraph. With semantic chunking, the AI sees four topic segments, summarizes each separately, then consolidates — preserving the actual structure of the conversation. Powered by current-generation OpenAI models for both map and reduce phases.
Pick your summary type (6 options)
The right summary depends on what kind of recording you uploaded. VexaScribe ships six purpose-built templates — each returns a different structured field set. Pick one before you click Generate, and switch any time without re-uploading.
Quality is consistent across types — instruction-tuned models match human-written quality on standardized news benchmarks (Zhang et al., TACL 2023). The differentiator is which structured fields surface what you actually need from the conversation. Picking a tool? See our honest comparison of 10 podcast transcription tools (Descript, Castmagic, Otter, Rev, and more).
What each type returns
Every summary is structured data with named fields, not a block of prose — which is the practical difference from pasting into a general chatbot. All types include an executive summary, and chapters with timestamps back into the audio wherever the recording has enough structure to divide.
| Type | Use it for | Fields you get back |
|---|---|---|
| General | Any recording where you just want the gist | Executive summary, topics, key quotes, and chapters with timestamps |
| Meeting | Team meetings, standups, board calls | Action items with assignee and deadline, plus decisions with rationale — on every meeting summary we sampled. Open questions and blockers when the meeting has them |
| Sales call | Discovery calls, demos, negotiations | Client needs, objections with the response given, competitor mentions, pricing discussions, next steps with owner and deadline, sentiment, deal stage, and BANT-style qualification signals |
| Interview | Job interviews, research interviews, journalism | Notable exchanges (topic, question, answer), strengths, and a written overall assessment. Concerns when the interview surfaces any |
| Lecture | Classes, talks, training sessions | Key concepts with explanations, examples given, terminology defined, takeaways, review questions, and further reading |
| Podcast | Episodes, panels, long-form conversations | Speaker profiles, discussion points, key insights, guest highlights, recommendations, and where speakers agreed or disagreed |
Transcribe once, then read the same recording through any lens — re-running a different summary type on an existing transcript costs nothing and uses no minutes. A long client call can be read as a sales call on Monday and as a meeting when you need the action items on Friday.
Types of transcripts we summarize (with routing to specialized workflows)
This page is the general-purpose transcript summarizer. If your transcript type has its own dedicated workflow — meeting minutes, deposition summary, video summary — you'll get a better result on the specialized page. Direct routing:
Meeting transcripts → structured minutes / summary
Zoom / Teams / Google Meet recordings. Board / project / sales / standup / non-profit templates. Robert's Rules formatting. Action items with assignees. Use this page for the meeting-specific vocabulary and template variety.
Legal deposition transcripts → paralegal-ready summary
Page-line, topical, or chronological deposition summary formats. Court-citation-ready references. Attorney review workflow. Not a certified court reporter replacement.
This is the work where the qualifier problem above matters most. A deposition turns on precisely the words that a summary is most likely to smooth away — a witness saying “I believe so” is not the same as “yes,” and “I don't recall” is not “no.” Treat the summary as triage that tells you which forty minutes of a seven-hour deposition to read properly, never as the record. The transcript is the record; the summary is an index into it, which is why the timestamps back to the source are the part that matters.
Video transcripts → video summary
YouTube videos, webinars, recorded video content. Video-first summarizer with chapter markers, timestamp-linked highlights, and YouTube-URL paste-in workflow.
Lecture transcripts → study notes
University lectures, recorded classes, professional training. Chapters + key concepts + terminology glossary + review questions. Lecture-first transcription with the study-note workflow attached.
Podcast transcripts → show notes
Podcast episodes with speaker diarization (host + guests), episode chapters, key insights, publishable show notes. Podcast-first workflow with post-hoc summarization.
Interview transcripts → stay here
Job interviews, research interviews, and journalism all use the Interview type, and it stays on this page because the fields are the whole answer. It returns notable exchanges — each keeping the topic, the question as asked, and a summary of the answer — plus strengths, concerns, and an overall assessment. For a hiring panel that means the disagreement about a candidate's systems-design answer survives as a discrete item instead of dissolving into a paragraph. For research interviews it means you can scan twelve transcripts for the same question and compare what each person actually said. If you are recording the interviews rather than working from text, interview transcription covers the capture side.
Anything else → stay here
Focus groups, personal voice memos, phone calls, one-off recordings, transcripts you already have from another tool — the six generic summary types (Meeting / Sales / Interview / Lecture / Podcast / General) on this page cover them.
Beyond summaries: ask follow-up questions with AI Chat
A summary gives you the gist. AI Chat lets you ask anything specific the summary missed — a natural-language Q&A interface scoped to one transcript, with citation chips that link back to the exact second in the audio where each quote came from.
How it complements summaries
- ● The summary gave you the key decisions — chat lets you ask “why did we decide that?” and get the reasoning quoted from the discussion
- ● The summary listed action items — chat lets you drill into “did anyone push back on the deadline?”
- ● The summary skipped something — chat lets you find it: “did anyone mention Q4 revenue numbers?”
- ● The summary is in English — chat lets you ask in your native language and get answers in it, with original-language quotes preserved
What's structurally different
- ● Citations are validated against the source. Quotes the AI proposes that don't actually appear in the transcript are dropped before reaching you.
- ● Timestamps are real. Click a citation chip — the audio player seeks to that exact moment for verification.
- ● Constrained to the recording. If something isn't in the audio, the AI says so rather than inventing.
- ● Multi-hour recordings handled via semantic retrieval. Tested on 6+ hour recordings with 21 speakers.
Action items extraction — what gets captured
The Meeting and Sales Call summary types extract action items as task + assignee + deadline. The prompt explicitly looks for two linguistic patterns: volunteer signals (“I'll handle that,” “I can take this on”) and assignment signals (“Can you do X?”, “John, can you handle this?”). Hedged language is filtered out by design — “maybe we should...” doesn't get flagged, to avoid false positives.
Transcript excerpt (Meeting)
Tom: So we need someone to own the redline by Friday. Priya: I can take that — I'll have it back to legal by Thursday afternoon. Rahul: Should I confirm the lawyer's availability? Tom: Yes, ping them Monday and let us know. Sarah: Let's circle back on the Slack-vs-Teams decision next week — I don't think we have time today.
Extracted action items
| Task | Assignee | Deadline | Detected as |
|---|---|---|---|
| Send redline to legal | Priya | Thursday afternoon | Volunteer pattern ('I can take that') |
| Confirm lawyer's availability | Rahul | Monday | Assignment pattern ('Yes, ping them Monday') |
How long should a transcript summary be?
Industry convention: a summary is typically 5-15% of the source transcript length. Short summary ~2-3%, executive summary ~5%, detailed summary ~10-15%. The right answer depends on who's reading it and what they need to do next.
VexaScribe doesn't force a target length — the summary is structured around fields (executive summary, chapters, action items, decisions, etc.) and the field content scales naturally with the source. The table below gives you a planning reference for the most common recording lengths.
| Source length | Short summary | Executive | Detailed | When to use which |
|---|---|---|---|---|
| 60-min meeting (~9,000 words transcribed) | 180-270 words (2-3%) | 450 words (5%) | 900-1,350 words (10-15%) | Short for Slack post-meeting; Executive for stakeholder digest; Detailed for project documentation |
| 30-min sales call (~4,500 words) | 90-135 words | 225 words | 450-675 words | Short for CRM Activity note; Executive for deal review; Detailed for handoff to AE / CSM |
| 60-min podcast (~9,500 words) | 190-285 words | 475 words | 950-1,425 words | Short for social-media teaser; Executive for show notes header; Detailed for full SEO-friendly show notes |
| 90-min lecture (~13,500 words) | 270-405 words | 675 words | 1,350-2,025 words | Short for quick review; Executive for study guide; Detailed for exam prep |
Accuracy and hallucination — how we handle it
Any generative AI summary can hallucinate — write things that weren't in the source. Our mitigations are at the prompt level and the architecture level. Below: the designed safeguards, what we've actually tested, and the honest limits.
What we've verified in internal testing
Small-sample testing (2 real transcripts, 4 summary runs). Two real data points beat a single “X% accurate” vendor claim. Verification date: 2026-06-29.
- Action item extraction recall: On a 10-minute board meeting with 7 hand-annotated ground-truth action items, the Meeting summary extracted 5 of 7 (71% recall) with 0 hallucinated items. Misses were borderline cases (one was arguably-not-an-action-item). Zero false positives matters more than total recall for trust.
- Sales-call deadline preservation: On a transcript with relative deadlines (“tonight tomorrow night,” “before tuesday january 27th”), the sales_call summary returned 0 fabricated calendar dates — deadlines stayed as “before January 27th” (verbatim). The same content summarized as Meeting DID fabricate dates (became “2026-01-22” and “2026-01-27”). This is a real, measurable difference between summary types.
- Action item assignee accuracy: On the 10-minute board meeting, 4 of 5 assignees were correctly named. One failure: an action item was assigned to “Shelley Scott” when the actual speaker was “Neil” — diarization mislabeled the speaker, and the summary inherited that label.
Anti-fabrication instructions in every prompt
All six summary types are explicitly instructed to never fabricate content not present in the transcript. The map and reduce phases both receive this guardrail. The guardrail is applied on every request.
Sales-call deadline preservation (unique to this type)
Sales Call summaries are explicitly instructed to preserve deadlines in the speaker's exact words. If a prospect says “next Thursday,” the summary says “next Thursday” — we don't convert relative dates to specific calendar dates. The instruction is repeated three times across the map prompt, reduce prompt, and system role to maximize compliance.
Honest limitation: this guardrail currently applies to sales_call summaries only. Meeting summaries may still convert relative dates to calendar dates — if you're summarizing a meeting where literal deadline wording matters, pick the Sales Call summary type even if the recording is internal. The meeting-prompt guardrail is on our backlog.
Map-reduce architecture reduces drift
Because each chunk is summarized in isolation against its own source segment (with overlap context), the model can't fabricate by “blending” far-apart parts of the transcript. The reduce phase consolidates from already-grounded chunk summaries rather than re-summarizing from scratch.
- ● No automatic entity deduplication pass. On long recordings, if a speaker is referred to by different names — “Dr. Smith” early on and “Mike” later — they may appear as two separate people in the summary's speaker fields. The model is sometimes good at unifying contextually within a single chunk, but cross-chunk unification is unreliable. We're improving this; it's not in production yet.
- ● No automatic verbatim cross-check pass against the transcript. We rely on prompt-level anti-fabrication instructions; we don't run a second AI pass to verify every quote.
- ● Assignee accuracy depends on diarization quality. Well-mic'd recordings with distinct voices produce clean assignee fields. Single-microphone conference rooms can blur speakers together, in which case assignees may inherit incorrect diarization labels (“SPEAKER_00”) or be mislabeled.
- ● Meeting-summary deadlines may be converted to calendar dates. See above — this guardrail exists only in sales_call mode today.
- ● High-stakes content (legal, medical, journalism quote-checking) still warrants human review. AI summaries are designed for everyday meetings, podcasts, lectures, and sales calls — the volume problem.
The mistake to watch for is not the one people expect
Most people worry a summary will invent something — a decision nobody made, a name nobody said. That happens, but it is the rarer problem, because summarizing is condensing a source rather than recalling facts from memory. The commoner and much subtler failure is a lost qualifier:
“If the Q3 numbers hold, we may expand the Berlin team — but that is a big if, and Dan has not signed off.”
“Decision: expand the Berlin team.”
Nothing was invented. Every word maps to something in the transcript. What vanished was the condition, the hedge, and the missing approval — and a reader skimming the action items has no way to tell. This is why summaries keep chapters with timestamps back into the audio: for anything consequential, read the line it came from before you act on it. It is also why each summary type extracts into fixed fields rather than free prose — a field called decisions that expects a rationale leaves less room to flatten a maybe into a yes than an open-ended “summarize this” does.
Why not just paste the transcript into ChatGPT?
Honest answer: ChatGPT works for one-off summaries when you already have the transcript text. For audio-in, repeatable, structured workflows it's the wrong tool — but it's a real option for occasional needs. The honest scorecard:
| Criterion | ChatGPT (paste a transcript) | VexaScribe | Advantage |
|---|---|---|---|
| Audio / video input | No — text only | Yes — 17 formats up to 5 GB | Us |
| Long-recording handling | Context-window-limited on multi-hour transcripts | Semantic chunking + map-reduce — multi-hour preserved | Us |
| Structured output | Requires custom prompt every time | 6 pre-built schemas (Meeting / Sales / Interview / Lecture / Podcast / General) | Us |
| Sales-deadline literal preservation | No guardrail | Explicit instruction — 'next Thursday' stays 'next Thursday' | Us |
| Multilingual summary (transcribe X, summarize Y) | Yes | Yes — 99 input languages × 99 output | Parity |
| Export to Markdown / DOCX / Notion / Slack | Manual | 1-click | Us |
| Training on your data | ChatGPT Plus trains on chats unless opted out | Never — explicit policy | Us |
| One-off transcript you already have | Fine — paste and prompt | Overkill — sign-up required | ChatGPT |
The honest framing: ChatGPT is a great general-purpose tool for ad-hoc transcript summaries when you already have the text and don't need a specific schema. VexaScribe is built for the audio-in, content-typed, repeatable case — and is the right tool when you want to skip the “design a prompt every time” cost.
To put a number on it: if your transcript is short enough to paste in one go, you already pay for a chatbot, and prose is the output you actually want, use the chatbot. You will get a decent summary and you will not have signed up for anything. The case for a purpose-built tool starts where that one breaks down — a recording too long to paste, the same summary needed every week in the same shape, action items that have to carry an assignee and a deadline as separate fields rather than as a sentence, or a summary you need to trace back to the minute it came from.
Best transcript summarizer: VexaScribe vs Otter, Fireflies, and ChatGPT
Different starting points call for different tools. Honest positioning per job-to-be-done — not a marketing ranking. The rows below describe how each tool takes its input and what it returns, which is structural and stable. Training and retention policies change, and we do not restate other companies' policies here — check each vendor's own page for those. Ours is set out in full further down.
| Tool | Best for | Input | Multi-hour | Structured schema | Trains on your data? |
|---|---|---|---|---|---|
| VexaScribe | Summarizing existing transcripts w/ content-typed schema | Paste transcript OR upload audio/video | ✅ semantic chunking | ✅ 6 preset types | No — see privacy below |
| ChatGPT / Claude | One-off ad-hoc summaries | Paste transcript text | ⚠️ context-window limits at ~90 min | ❌ design prompt every time | Refer to vendor policy |
| Otter.ai | Live-capture on meetings (bot joins) | Live meeting + bot participation | ✅ but geared to live | Meeting-only | Refer to vendor policy |
| Fireflies.ai | Live-capture + CRM integration for sales | Live meeting + bot participation | ✅ but geared to live | Meeting + sales-call | Refer to vendor policy |
Pick VexaScribe when
You already have the transcript (from a court reporter, another tool's export, WhatsApp voice notes, an old recording) OR you record with your platform's built-in recording and upload after. You want a repeatable content-typed schema without designing a prompt every time. Multi-hour transcripts must not get flattened.
Pick a live-capture tool (Otter / Fireflies) when
You have a stable calendar of recurring meetings and want zero manual action per meeting. You accept a bot participant on your calls (check confidentiality / wiretap-consent implications for your jurisdiction). You use the CRM integration heavily (Fireflies for sales pipeline).
AI summary vs human-written: when each wins
AI wins on cost, speed, and scale. Humans still win on legal liability, deep cultural nuance, and stakes-where-a-mistake-is-fatal. The honest scorecard:
| Criterion | AI summary | Human-written | Winner |
|---|---|---|---|
| Cost per 60-min summary | ~$0.00–$0.30 | $30–$80 (freelance) | AI |
| Turnaround time | ~15 seconds | 2–24 hours | AI |
| Accuracy on standard transcripts | On par with human-written quality on standardized news benchmarks (Zhang et al., TACL 2023). Designed safeguards: explicit anti-fabrication instructions; sales-call deadlines preserved in speaker's exact words. | ~96–98% with domain-expert reviewer | Human (narrowly) |
| Cultural / idiomatic nuance | Misses sarcasm, regional idioms | Strong when reviewer shares context | Human |
| Legal / medical liability | Not certified; output is not auditable testimony | Trained transcriptionist + signed attestation | Human |
| Scale (1,000+ transcripts/week) | Trivial | Requires a team | AI |
| Format consistency across runs | High — deterministic templates | Variable across humans | AI |
| Long-tail languages (e.g., Welsh, Swahili) | Strong on top 25, weaker on long tail | Depends on reviewer availability | Tie / depends |
AI wins for everyday meetings, podcasts, and lectures — the volume problem. Humans still win when a single misquote can land you in court or harm a patient: legal depositions, medical records, high-stakes journalism. VexaScribe's stance: ship AI summaries with explicit anti-fabrication instructions and sales-call deadline preservation, route the 5% of high-stakes cases to human review. See our editorial review process.
Summarize in any language — or translate while you summarize
VexaScribe transcribes in 99 input languages and can deliver the summary in any of the world's major languages — input and output are independent. Upload a Spanish meeting recording and get English action items, or vice versa, in a single workflow.
Common pairings:
- • Spanish meeting recording → English action items (US team distribution)
- • German lecture recording → English chapters + key concepts (international students)
- • Japanese podcast → English show notes (cross-market publishing)
- • Portuguese interview → English quotes + themes (research synthesis)
For pure audio translation without summarization, see transcribe and translate audio.
Export the summary anywhere
The summary downloads as Markdown, DOCX, or plain text — and copy-to-clipboard preserves Markdown formatting so you can paste cleanly into Notion, Obsidian, Google Docs, or Slack. The full transcript exports separately as TXT, DOCX, SRT, VTT, or JSON.
| Integration | Supported export | Setup |
|---|---|---|
| Markdown file | .md with frontmatter + headings | Direct download, no setup |
| DOCX (Word) | Word-compatible with styles | Direct download, no setup |
| Plain text | .txt for any text editor | Direct download, no setup |
| Notion (paste) | Copy to clipboard with Markdown formatting | Paste into any Notion page |
| Obsidian (paste) | Markdown with wiki-link compatible headings | Paste into vault |
| Google Drive (paste) | Copy to clipboard, paste into Doc | Manual paste |
What happens to your recording
Audio uploads are encrypted in transit (TLS 1.2+) and at rest. Your recordings are not used to train AI models, and you can delete files any time.
- TLS 1.2+ in transit, encrypted at rest in AWS eu-west-2.
- Not used for training — your recordings, transcripts, and summaries are not used to train AI models. The premium transcription and diarization providers we use are contractually barred from training on your data, and summaries and translations run on OpenAI's API, which does not train on API traffic.
- Self-serve deletion any time from your dashboard.
- Deletion is immediate — removing a file or your account deletes the records and every stored audio file, transcript, and summary straight away, not on a retention schedule.
Full details in our privacy policy.
Transcript Summary — Frequently Asked Questions
How does transcript summary work on VexaScribe?
Upload an audio or video recording (MP3, WAV, M4A, MP4, MOV, and 12 other formats up to 5 GB). VexaScribe transcribes the audio with Whisper Large-v3, then generates an AI summary tailored to your chosen type — General, Meeting, Sales Call, Interview, Lecture, or Podcast. A 60-minute file typically completes in about two to six minutes including both transcription and summary — around two minutes per audio hour on our premium model, three to five on standard, plus a few seconds for the summary itself.
What audio and video formats are supported?
MP3, WAV, M4A, FLAC, OGG, AAC, AIFF, WMA, AMR, OPUS for audio, and MP4, MOV, AVI, MKV, WebM, FLV, WMV for video. Files can be up to 5 GB and 10 hours long. For video files, audio is extracted automatically.
What summary types are available?
Six purpose-built types — General, Meeting, Sales Call, Interview, Lecture, and Podcast. Every summary includes an executive summary, chapters, and key quotes. Each type then adds specialized sections: Meeting adds action items (task, assignee, deadline), decisions (filtered to exclude action items), unresolved open questions, and blockers. Sales Call adds action items with deadlines preserved in the speaker's exact words, client needs with priority and evidence quotes, objections with resolution status, competitor mentions with positioning, pricing discussions with outcomes, deal next steps, sentiment with reasoning, deal stage (Discovery / Demo / Negotiation / Closing / Closed Won/Lost), and BANT qualification signals. Lecture adds key concepts, examples with teaching insight, terminology glossary, takeaways, review questions with answer hints, and further reading. Interview adds notable exchanges, strengths, concerns, and an overall hire assessment. Podcast adds speaker profiles, discussion points per speaker, key insights, agreements and disagreements, recommendations, and guest highlights. General adds a topics list.
How does VexaScribe handle long recordings without losing structure?
Map-reduce architecture with semantic chunking. For long transcripts, we use OpenAI embeddings (text-embedding-3-small) to detect natural topic boundaries — sentence-to-sentence similarity 'valleys' that mark topic shifts. Each chunk is summarized in parallel with overlap context from the previous chunk, then a final reduce phase consolidates everything into one coherent summary, deduplicating action items and merging only chapters that are genuinely the same topic. Multi-hour recordings preserve their structure instead of getting flattened to a single paragraph.
How accurate are AI summaries, and can they hallucinate facts?
On standardized news benchmarks, instruction-tuned LLM summaries are judged on par with human-written ones (Zhang et al., TACL 2023). In our own internal testing on a 10-minute board meeting with 7 hand-annotated action items, the Meeting summary extracted 5 of 7 (71% recall) with zero hallucinated items. On the same transcript, sales_call mode preserved relative deadlines verbatim ('before January 27th') while Meeting mode fabricated calendar dates from the same input — a real, measurable difference between summary types. Safeguards: all six types are explicitly instructed not to fabricate content not in the transcript, and sales_call mode has a triple-repeated deadline-preservation guardrail. Honest limitation: the deadline guardrail exists only in sales_call mode — meeting summaries may still convert relative dates. For legally-sensitive content, humans still win — see the AI vs human comparison section.
Is my recording private — do you train on it?
No. Audio files transit over TLS 1.2+ encryption and are stored encrypted at rest in AWS eu-west-2. We don't use customer data to train AI models. Self-serve deletion any time from your dashboard. Account deletion purges all recordings, transcripts, and summaries.
Which languages does the summary support?
VexaScribe transcribes in 99 languages with automatic language detection. Summaries can be generated in any of the world's major languages — feed a Spanish meeting recording and get English action items, transcribe Turkish and summarize in English, etc. Input language and output language are independent.
What does the free tier include?
30 minutes of transcription on the free preview, with summary generation included. Paid plans: Starter $2/mo (200 min), Basic $5/mo (1,000 min), Pro $10/mo (2,500 min), Studio $20/mo (6,000 min). All plans include all 6 summary types and all export formats.
Why not just paste the transcript into ChatGPT?
ChatGPT works for one-off summaries when you already have the transcript text. VexaScribe is the integrated workflow for repeatable use: upload audio once and get a transcript plus six pre-built summary schemas; no need to design a custom prompt every time. We chunk long recordings semantically (ChatGPT is context-window-limited on multi-hour audio), preserve sales-call deadlines in the speaker's literal words, and do not use your data to train AI models. ChatGPT Plus trains on chats by default unless you opt out; we don't, ever.
What's the best AI for summarizing transcripts — VexaScribe vs Otter, Fireflies, or ChatGPT?
Different tools for different starting points. If you already have the transcript text (from a court reporter, another tool's export, or manual transcription): ChatGPT / Claude paste-and-prompt works for one-off jobs but requires you to design the prompt every time, and hits context-window limits at ~90 minutes of transcript. VexaScribe is built for this exact input — paste the transcript, pick a content type (meeting / sales / lecture / etc), get a structured schema every time, and multi-hour transcripts are chunked semantically instead of failing. If you don't have the transcript yet and want a bot to join your meeting live: Otter or Fireflies are bot-first; they join the call, record automatically, summarize post-call. VexaScribe is upload-first — you record with your platform's built-in recording (Zoom, Teams, Meet, or a dictaphone) and upload after. Honest ranking on the specific job "summarize a transcript I already have": VexaScribe is purpose-built for it and typically outperforms ChatGPT paste-and-prompt on structure and multi-hour handling; Otter and Fireflies are optimized for live-capture instead, so they're not the natural choice for this specific job.
Does the AI hallucinate when summarizing a transcript?
It can, and the two failure modes are different. Outright invention — a decision nobody made — is rare in summarization because the model is condensing a source rather than recalling facts. The commoner and subtler failure is a dropped qualifier: "we may expand to Berlin" summarized as "we will expand to Berlin". Read anything consequential against the transcript line it came from before acting on it.
How do you prevent fabricated details in the summary?
The summary is generated only from the transcript text, never from general knowledge, and each summary type extracts into fixed fields rather than free prose — a meeting summary has to fill action items, decisions, open questions and blockers, which leaves far less room to drift than an open-ended "summarize this" prompt. Meeting summaries also get a second focused pass on decisions, because small models under-extract implicit agreements.
Is an AI summary accurate enough for legal or medical transcripts?
It is useful for triage — finding which parts of a long deposition or consultation deserve attention — and it is not a substitute for the record. The transcript is the record; the summary is an index into it. For anything that carries legal or clinical weight, verify against the source text, which is why every summary keeps timestamps back to the audio.
Do summaries use up my transcription minutes?
No. Minutes are charged for transcription only. Summaries are included, and re-running a different summary type on the same transcript costs nothing — you can transcribe once and read the same recording as a meeting, then as an interview, then as a lecture.
Can I get timestamps or chapters in the summary?
Yes. Every summary includes chapters with start and end timestamps pointing back into the audio, so a point in the summary can be traced to the moment it came from rather than read on trust.
How fast is summarization?
The summary itself takes a few seconds once a transcript exists. If you are starting from audio, transcription is the part that takes time: most files finish in about two minutes per hour of audio on our premium model, and three to five minutes per hour on the standard model.
What is the paste limit for the free summary?
The free summary on this page reads up to 5,000 characters — roughly 750 words, or about five minutes of speech — and runs once per week without an account. For anything longer, upload the audio or video instead: files run up to 5 GB with no limit on how long the recording is.
Do I need an account to summarize a transcript?
Not for the free summary on this page — paste a transcript, pick a type, and read the result. A free account removes the weekly limit and the length cap, and adds AI chat, translation, and the full set of export formats.
What does the AI summary not handle perfectly?
Two honest limits worth knowing. First, entity unification across long recordings — if a speaker is referred to as both 'Dr. Smith' and 'Mike' at different points, they may appear as two separate people in the summary's speaker list. We don't run a name-canonicalization pass yet. Second, action item assignees depend on diarization quality — well-mic'd meetings with distinct voices produce clean assignee fields, but single-microphone conference rooms can blur speakers together and assignees may fall back to whoever is explicitly named in the spoken sentence. Both are areas of active improvement.
Which summary type should I pick for a 60-minute team meeting versus a 90-minute lecture?
Team meeting: pick Meeting — surfaces action items with assignees, decisions, and blockers in a structured layout. Lecture: pick Lecture — generates chapters with timestamps plus key concepts, terminology glossary, and review questions for study. Podcast: pick Podcast for publishable show notes with speaker profiles. Sales call: pick Sales Call for objections, deal stage, BANT signals, and deadlines preserved literally.
If your starting point is different
You have a recording, not a transcript
Upload the audio and get the transcript and the summary in one pass.
The meeting has not happened yet
A note-taker that joins the call and writes the summary as it goes.
It is a video
Start from a video file or a YouTube URL instead of pasted text.
You have a folder of them
Process many files in one pass rather than pasting one at a time.
You need formal minutes
Board-style minutes with attendees, motions, and votes — a different document from a summary.
The recording is in another language
Transcribe and translate together, with timings and speakers preserved.
You want to compare tools first
An honest comparison of the dedicated meeting-summary tools, ours included.