Skip to content
Advertisement
AI Tools

Best AI Transcription Tools: What Diarization Really Costs

AI Tools Tutorial Team20 min readDocumentation and user reports

Pricing and features verified August 2026

Photograph of recording studio

Photo by Kristin Hardwick via stocksnap (CC0)

Speechmatics is the best API pick, from $0.129 per audio hour (checked August 2026), and the only one of five API vendors publishing an automatic deletion timeline: 7 days for batch, never stored for real-time. Deepgram hands out the largest starting credit at $200. Sonix is the best finished app at $10 an hour pay-as-you-go, because its security page states plainly that none of your data is used for training. No vendor accuracy percentage appears anywhere below.

Key takeaways

  • Speaker diarization is a metered add-on, not a checkbox — Deepgram adds $0.0020/min, AssemblyAI adds $0.12/hr on streaming, OpenAI makes you change model
  • Speechmatics is the only API vendor of five publishing an automatic deletion timeline for your audio: 7 days batch, real-time never stored
  • Speechmatics will take 33% off your rate if you opt in to model training — an explicit data-for-discount trade
  • AssemblyAI free users cannot opt out of the model improvement program at all, and opt-out is never retroactive
  • The newest and priciest model covers the fewest languages: 18 on AssemblyAI Universal-3.5 Pro versus 99 on the cheaper Universal-2

Why this list is built differently#

Three questions decide this purchase. What does speaker separation cost as a line item? What does the contract let the vendor do with your audio? Which unit are you actually billed in?

Advertisement

How these figures were checked, and what went unpublished#

Every price, limit and policy statement below came from a primary page opened in August 2026: the vendor's pricing page, developer docs, terms of service, privacy policy or security page. Those exact URLs are in the sources list. No source, no claim.

Two well-known tools were dropped for failing it. Trint's pricing page returned a 404 on every plausible URL; Notta's returned empty content on repeated attempts. Neither has a primary price in this article because neither published one we could open.

No per-minute rate appears for Rev anywhere below, because Rev does not publish one.

Every tool, priced in the unit it actually bills#

This table is the point of the article. Building it means reading the pricing page instead of the marketing page.

Billing units and published from-prices, verified against each vendor's own pricing page in August 2026.
ToolBilled inFrom-priceFree to startDiarization costRetention window published
DeepgramAudio minute$0.0077/min, Nova-3 monolingual PAYG$200 credit, no card+$0.0020/minNo
AssemblyAIAudio hour (async) or WebSocket second (streaming)$0.15/hr, Universal-2 async$50 credit, no card+$0.02/hr async, +$0.12/hr streamingNo
OpenAIAudio minute (estimated from token billing)$0.003/min, gpt-4o-mini-transcribeNo signup credit publishedRequires a different model30-day abuse logs; ZDR on approval
SpeechmaticsAudio second, quoted per hour$0.129/hr on Pro$100 credit, no cardNot metered separately on the pricing pageYes — 7 days batch, none real-time
ElevenLabs ScribeAudio hour, or bundled plan hours$0.22/hr for Scribe v2No Scribe allowance listedNot metered separately on the pricing pageNot on the API pricing page
DescriptMedia hours AND AI credits, per seat$16/mo billed annually on Hobbyist60 media min/mo, 100 one-time creditsSpeaker detection on paid plansNo
SonixAudio hour, or bundled plan hours$10/hr pay-as-you-go30-minute trial, no cardSpeaker labels included on all plansUser-controlled deletion
RevMinutes per seat per month$25.49/mo billed yearly on Essentials45 min/mo, English onlyNot priced separatelyNo
Happy ScribeMinutes, priced in eurosEUR 8.50/mo billed annually on Basic10-minute trialNot priced separatelyNo
Billing units and published from-prices, verified against each vendor's own pricing page in August 2026.

Nine tools, and the billing unit changes in almost every row. That is why ranking these by headline price tells you close to nothing.

Advertisement

What published accuracy figures actually measure#

The Open ASR Leaderboard is the closest thing this category has to a neutral scoreboard. Its paper evaluates 86 systems across 12 datasets, scoring word error rate and inverse real-time factor — accuracy and speed, tracked separately because they trade against each other.

Look at the domain the paper's own Table 1 assigns each of those 12 datasets. Three are audiobooks: LibriSpeech clean, LibriSpeech other and MLS. GigaSpeech is "audiobook, podcast, YouTube". FLEURS is Wikipedia, CoVoST-2 is open domain, TED-LIUM v3 is TED talks, VoxPopuli is European Parliament. Earnings21 and Earnings22 are earnings calls and SPGISpeech is financial meetings.

One entry out of twelve — the AMI Meeting Corpus — carries the domain label "Meetings".

Vendor numbers skip the dataset entirely. ElevenLabs' Scribe page states that Scribe "achieves industry-leading accuracy across 90+ languages, outperforming models like Whisper, Deepgram, and Gemini in benchmark tests" — illustrated with bar charts that name no benchmark and no dataset anywhere on the page.

Four things move your real error rate, and AssemblyAI's diarization docs name three: cross-talk, background noise and speakers who sound similar. The fourth is domain jargon — product names, internal acronyms — which is exactly what you will search the transcript for later.

Speaker diarization is a metered line item, not a checkbox#

Feature tables render speaker separation as a tick. On the two biggest APIs here, it is a second bill.

  • Deepgram meters diarization at $0.0020 per minute on top of the model rate. Against Nova-3 monolingual at $0.0077 per minute, that is roughly a 26% uplift for knowing who said what.
  • AssemblyAI stacks it additively: +$0.02 per hour for standard async, +$0.065 for the experimental async version, and +$0.12 on streaming. Against Universal-Streaming at $0.15 an hour, that is an 80% uplift.
  • OpenAI does not sell it as an add-on at all. Its docs direct you to gpt-4o-transcribe-diarize, a separate model, when you need to identify who speaks in which part of a recording.

The documented constraints are harder to find than the prices. AssemblyAI's docs say each speaker should have at least 30 seconds of continuous speech for good results, and warn that setting the maximum expected speaker count too high causes over-splitting — the API returns more speaker labels than there are people in the room.

Deepgram's v2 diarizer is batch-only: a streaming request that asks for it returns a validation error. Streaming diarization also returns a speaker label with no confidence score, where pre-recorded returns both. Deepgram's docs state that diarization works on Nova-1, Nova-2, Nova-3, Enhanced and Base — and that Whisper is not supported.

Advertisement

What the contract lets each vendor do with your audio#

Four vendors, four completely different postures. This section decides the purchase for anyone handling customer calls.

OpenAI is off by default. Its data-controls docs state that since 1 March 2023, data sent to the API is not used to train or improve OpenAI models unless you explicitly opt in. Abuse-monitoring logs are retained up to 30 days. Zero Data Retention covers the transcription endpoints but requires prior approval — it is not a self-serve toggle.

Speechmatics sells the trade openly. Its pricing page describes an opt-in program that takes 33% off speech-to-text rates in exchange for permission to use your audio and transcripts to improve its models. Decline and you pay full price. That puts a number on the thing everyone else buries in clause 8.

Deepgram reserves the right in its terms. They state Deepgram "may also use Your Content to improve our Services, to develop other products and services, and for other business purposes, including training and testing our Models". API customers opt out on a per-request basis with the mip_opt_out parameter — so the protection is only as good as the one engineer who remembers it on every call path.

AssemblyAI grants the same license, with a paywall on the exit. Its terms cover testing, evaluation, benchmarking and training. Its own opt-out FAQ states that free users do not have the ability to opt out, and that opt-out is forward-looking only. Upload first, opt out later does not work.

Then the retention asymmetry, which is the buying signal. Speechmatics publishes it: batch audio, transcripts and job configuration stored 7 days then deleted automatically, real-time audio never recorded at all. OpenAI publishes 30 days for abuse-monitoring logs and offers zero retention on the transcription endpoints, but only to customers it approves. Deepgram's privacy notice says customer data "shall be retained, stored, and deleted according to our agreement with our business customer" — a contract, not a number. AssemblyAI publishes no window on its pricing page, privacy policy, terms, security page or docs.

API-first options#

Speechmatics

Best for: teams that need a retention window in writing

4.6

Pricing
Pro from $0.129/hr; free $100 credit, no card (checked August 2026)

The only vendor here that publishes an automatic deletion timeline for the audio and transcripts themselves — 7 days for batch, never stored for real-time — and it holds ISO/IEC 27001:2022, SOC 2 Type II, GDPR and HIPAA coverage. Billing is calculated to the second, and usage above 500 hours a month per transcription type triggers an automatic 20% discount.

Its multilingual model, Melia 1, switches between languages automatically without you selecting one, and seven bilingual packs cover pairs including Mandarin-English, Spanish-English and Arabic-English. Specific limitation: language identification is supported on batch transcription only, so a real-time stream cannot detect its own language — you must know it before the WebSocket opens. Files posted directly in the request body must be under 1 GB or the job is rejected; larger files have to be passed by URL.

Deepgram

Best for: high-volume batch transcription on a tight per-minute budget

4.4

Pricing
Nova-3 monolingual from $0.0077/min pay-as-you-go; $200 free credit, no card (checked August 2026)

The largest starting credit in the category and the most granular add-on pricing: redaction at $0.0020/min, entity detection at $0.0017/min, keyterm prompting at $0.0013/min, with smart formatting included. Compliance coverage spans SOC 2, GDPR, HIPAA, CCPA, PCI and Australian Privacy Principle 8.

Specific limitation: requests whose processing exceeds 10 minutes return a 504 Gateway Timeout on Nova, Base and Enhanced models — a failure you will hit on long files, not on your test clip. Maximum file size is 2 GB. The discounted Growth tier starts at $4,000 a year in pre-paid credits, redeemed against actual usage.

AssemblyAI

Best for: broad language coverage on the cheaper model tier

4.1

Pricing
Universal-2 async $0.15/hr; Universal-3.5 Pro $0.21/hr; $50 free credit, no card (checked August 2026)

Certifications are unusually specific: SOC 2 Type 1 and Type 2, PCI-DSS 4.0 Level 1 as of 31 March 2025, a GDPR third-party assessment, and processing in Dublin for customers who need EU data residency. Its diarization docs name what degrades speaker labels: overlapping speech, background noise and similar-sounding voices.

Specific limitation: streaming is billed on WebSocket session duration — the time the connection is open, not the audio you send. An idle socket left open during a coffee break bills like speech. Free-tier concurrency is capped at 5 new streams per minute against 100 on pay-as-you-go.

OpenAI

Best for: cheap batch transcription where training-off-by-default matters

4.0

Pricing
gpt-4o-mini-transcribe estimated at $0.003/min; Whisper and gpt-4o-transcribe at $0.006/min (checked August 2026)

The only API vendor here where model training is off by default rather than opt-out, and Zero Data Retention is available on the transcription and translation endpoints if you get approval. Whisper covers 98 languages, though the docs note accuracy varies by language.

Specific limitation: files can be up to 25 MB. That hard cap forces you to build chunking logic before you can transcribe a single hour-long meeting. The per-minute figures are labelled as estimates because the gpt-4o transcription models bill at token rates, so a dense, fast-talking recording costs more than a quiet one of the same length.

ElevenLabs Scribe

Best for: teams already paying ElevenLabs for voice work

3.8

Pricing
Scribe v2 at $0.22/hr; Realtime at $0.39/hr; included hours from Starter at $6/month (checked August 2026)

Covers 90+ languages and bundles speaker labelling, entity timestamps and redaction into the Scribe v2 product description. Add-ons are priced per hour: entity detection $0.070/hr, keyterm prompting $0.050/hr.

Specific limitation: the API pricing page lists included Scribe hours only from the Starter plan upward, with no free transcription allowance stated at all. The bundles are also poor value against the raw rate: Creator at $22 a month for 27 hours works out near $0.81 an hour, nearly four times the $0.22 pay-as-you-go price.

Advertisement

Finished apps#

Sonix

Best for: buyers who want a no-training commitment in plain English

4.3

Pricing
$10/hr pay-as-you-go; Core $25/month for 5 hours; 30-minute free trial (checked August 2026)

The clearest data commitment of any vendor here: none of your data processed through Sonix is used for training, deletion wipes both audio and transcripts permanently, and files stay accessible after your subscription ends. Every plan includes the editor, speaker labels, subtitles and 2FA.

Specific limitation: overage beyond your plan allowance is billed at $10 an hour — exactly the pay-as-you-go rate — so subscribing saves nothing on the marginal hour. Additional seats are $25 a month or $275 a year each, which gets expensive for a five-person team before you transcribe anything.

Descript

Best for: transcription that feeds straight into video and podcast editing

4.2

Pricing
Free 60 min/month; Hobbyist $16/month billed annually; Creator $24/month billed annually (checked August 2026)

Speaker detection covering 8+ speakers is included on Hobbyist, Creator and Business, and paid tiers support 25 languages for multi-language transcription. Annual billing saves up to 35%, and Creator's 30 media hours a month work out near $0.80 an hour of media time — a full editing suite included at that rate.

Specific limitation: Descript meters two currencies that run out independently. Media hours cover transcription, AI credits cover generative features, and exhausting either stops that half of your workflow — and the free plan's 100 AI credits are one-time, not monthly. Its privacy policy also states it "may also use your Projects to analyze and improve the Descript Service" unless you disable the Share Data with Descript setting.

Rev

Best for: high monthly minute volume per seat

3.9

Pricing
Free 45 min/month; Essentials $25.49/month billed yearly; Pro $47.99/month billed yearly (checked August 2026)

The most generous minute allowances here by a wide margin: 5,000 AI transcription and caption minutes per seat per month on Essentials, 10,000 verbatim minutes on Pro. Human transcription sits alongside the AI service, and Rev's security page states that "every transcriptionist signs a strict NDA".

Specific limitation: language coverage is gated behind the top paid tier. The free plan is English only, Essentials adds only Spanish, and you need Pro at $47.99 a month billed yearly before you reach 37+ languages. Read Rev's training statement carefully too — it promises no training of external LLMs, wording that does not address Rev's own models.

Happy Scribe

Best for: European teams billing in euros who want AI and human transcription in one place

3.5

Pricing
Basic EUR 8.50/month billed annually for 120 minutes; pay-as-you-go top-up EUR 0.20/min (checked August 2026)

The widest language claim in this roundup at 150+ languages for AI transcription, with human services covering 65+. Human transcription starts at EUR 1.75 a minute, dropping to EUR 1.66 on the Business plan.

Specific limitation: its security page publishes nothing at all about data retention, deletion timelines, or whether customer data trains AI models. For a tool that also routes files to human transcribers, that silence is the loudest thing on the page. Pricing is in euros, which adds exchange-rate noise to a US budget.

The language-coverage trap#

Shortlist by language first, because the newest and most expensive model supports the fewest of them.

AssemblyAI's newest async model, Universal-3.5 Pro at $0.21 an hour, supports 18 languages. The older, cheaper Universal-2 at $0.15 an hour supports 99. Paying more buys you 81 fewer languages.

Deepgram shows the same shape internally: Nova-3 covers 50+ languages in monolingual modes but only 10 in multilingual mode — English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. If your calls switch languages mid-sentence, that is the list you are choosing from.

The practical rule: settle your language mix first, then pick the model. Doing it the other way round means discovering the gap in production.

Advertisement

Working out which is actually cheaper for your volume#

The method is simple arithmetic, and it is the only honest way to compare a seat plan against a per-minute API. Divide each plan's published price by its published included hours, then compare that effective hourly rate against the API rates directly.

Effective hourly rates derived by dividing each vendor's published plan price by its published included hours or minutes, August 2026.
PlanPublished priceIncluded volumeEffective rate per hour
Descript Creator (annual)$24/mo30 media hoursabout $0.80
Descript Hobbyist (annual)$16/mo10 media hoursabout $1.60
Sonix Core$25/mo5 hoursabout $5.00
Sonix Pro$80/mo40 hoursabout $2.00
Rev Essentials (yearly)$25.49/mo5,000 minutes per seatabout $0.31
Rev Pro (yearly)$47.99/mo10,000 minutes per seatabout $0.29
ElevenLabs Creator$22/mo27 Scribe hoursabout $0.81
ElevenLabs Scale$299/mo450 Scribe hoursabout $0.66
Happy Scribe Business (annual)EUR 59/mo6,000 minutesabout EUR 0.59
Effective hourly rates derived by dividing each vendor's published plan price by its published included hours or minutes, August 2026.

Four traps break the naive version of this. AssemblyAI's streaming bills connection time, not audio time, so an app that holds sockets open inflates the bill invisibly. Multichannel audio is billed per channel there too — a one-hour stereo file is two billable hours. Diarization stacks additively on Deepgram and AssemblyAI. And Descript's dual meter means a plan with hours left can still stall because the credits ran dry.

Run your own hours through it. For the same exercise applied to token-billed models, see the guide to estimating LLM API costs.

Limits that break workflows before pricing does#

Scan this before you commit engineering time.

  • OpenAI: 25 MB per file, hard. Formats: mp3, mp4, mpeg, mpga, m4a, wav, webm.
  • Deepgram: 2 GB per file, plus a 504 Gateway Timeout when processing exceeds 10 minutes on Nova, Base and Enhanced, or 20 minutes on Whisper.
  • Speechmatics: 1 GB on files posted in the request body; 20,000 concurrent jobs before HTTP 429; 2 concurrent real-time sessions on the free tier against 50 on Pro.
  • AssemblyAI: free-tier streaming capped at 5 new streams per minute, versus 100 on pay-as-you-go.

Which limits shout and which whisper is the distinction that matters. A 504 or a 429 tells you something broke. A silently dropped feature does not.

Who this is not for#

If you need certified verbatim transcripts carrying a signed accuracy guarantee, none of the automated options here is the right purchase on its own. Rev and Happy Scribe both sell human transcription alongside their AI service; that is the product to price.

If your audio is mostly meetings and what you want is summaries, action items and a searchable record rather than a verbatim transcript, you are shopping in the wrong category. Read the note-taking apps roundup and the AI meeting assistants comparison instead — those tools join your calls and structure the output, which these do not.

And if the audio cannot leave your network at all, none of the hosted rates priced above apply to you. Self-hosted deployment is a different build, a different contract and a different price list.

What was left out and why#

Trint and Notta are absent because their pricing pages could not be retrieved, and a third-party blog does not count as a primary source for a price.

Otter, Fireflies, Fathom and Granola are absent for a different reason: they are meeting-notes assistants, not general transcription tools. Different job, covered in the note-taking apps roundup.

Every accuracy percentage is also left out, including the ones published by the vendors recommended above.

How to run your own two-week bake-off#

  1. Start with your worst recording, not your best

    Pick the four-person call with the bad room, the interruptions and the accent range. Clean recordings are what the benchmarks are made of, so testing on one tells you nothing new.

  2. Turn the training opt-out on before you upload anything real

    On AssemblyAI it is a paid-plan setting only an owner or admin can change, and it never applies retroactively. On Deepgram it is a per-request parameter. On Descript it is the Share Data with Descript toggle.

  3. Spend the free credits, not your budget

    At published rates, Deepgram's $200 buys roughly 430 hours at $0.0077 a minute, Speechmatics' $100 roughly 775 hours at $0.129, and AssemblyAI's $50 roughly 330 hours at $0.15. Sonix gives 30 minutes and Descript 60 minutes a month — one real meeting each.

  4. Count errors on proper nouns and jargon, not on total words

    Overall word accuracy is dominated by common words the model was always going to get right. Count the misses on product names, customer names and internal acronyms instead.

  5. Test diarization at your actual speaker count

    Run the file with your real number of participants. Check whether short contributions get attributed correctly and whether the tool invents extra speakers. Do not test with two people if your calls have six.

  6. Read the retention answer before you sign

    If the vendor cannot say how long your audio is stored and when it is deleted, that is your answer. The checklist for evaluating any AI tool before buying covers the rest.

The verdict#

Buy Speechmatics if you are building on an API and anyone in your organization will ever ask where the audio goes. From $0.129 an hour (checked August 2026), it is the only vendor here that publishes an automatic deletion timeline for your content, and its 33%-off-for-training-data offer is the most honest deal in the category.

Buy Deepgram if volume is the binding constraint and compliance is settled elsewhere. At $0.0077 a minute with $200 of starting credit, nothing here is cheaper to prototype on — budget the extra $0.0020 a minute for speaker labels from day one.

Buy Sonix if you want a finished app with a plain-English no-training commitment, and Descript if the transcript is the first step of an editing job rather than the deliverable — with its data-sharing toggle switched off on day one.

Skip anything whose pricing page you cannot open, and skip any accuracy percentage that does not name a benchmark. Those two rules alone will improve your shortlist more than any ranking will.

Frequently asked questions

How much does AI transcription cost per hour?

Published from-prices in August 2026 span Speechmatics Pro at $0.129 an audio hour up to Sonix pay-as-you-go at $10 an hour. API vendors bill per audio minute or hour; finished apps bundle hours into a seat price. On Deepgram and AssemblyAI, speaker labels are billed on top of that.

Can AI transcription tell the difference between speakers?

Yes, but it is usually a paid add-on rather than a default. Deepgram meters diarization at $0.0020 a minute on pay-as-you-go, AssemblyAI at $0.02 an hour for async and $0.12 an hour for streaming, and OpenAI makes you switch to a separate diarizing model entirely.

Do transcription services keep your recordings and use them to train AI?

It varies sharply. OpenAI states API data has not trained its models by default since March 2023. Deepgram and AssemblyAI both reserve training rights in their terms with an opt-out. Speechmatics takes 33% off if you opt in. Sonix states none of your data is used for training.

Is there a free AI transcription tool with no time limit?

Not among these. Descript's free plan gives 60 media minutes a month, Rev gives 45 English-only minutes a month, and Sonix offers a one-off 30-minute trial. The API vendors give starting credit instead: $200 at Deepgram, $100 at Speechmatics and $50 at AssemblyAI.

What is the most accurate AI transcription tool?

No published figure answers this honestly. Vendor accuracy percentages rarely name a benchmark or a dataset, and the Open ASR Leaderboard averages 86 systems across 12 datasets that are mostly audiobooks, read text and prepared talks. Test your own worst recording on two or three candidates.

Sources

  1. Deepgram — Pricing
  2. Deepgram — Terms of Service
  3. Deepgram — Privacy Policy
  4. Deepgram Docs — Diarization
  5. Deepgram Docs — Models and languages overview
  6. Deepgram Docs — Pre-recorded audio
  7. Deepgram Docs — Data privacy and compliance
  8. AssemblyAI — Pricing
  9. AssemblyAI Docs — Supported languages
  10. AssemblyAI Docs — Speaker diarization
  11. AssemblyAI Docs — How to opt out of the Model Improvement Program
  12. AssemblyAI — Terms of Service
  13. AssemblyAI — Security
  14. AssemblyAI — Privacy Policy
  15. OpenAI — API pricing
  16. OpenAI Docs — Speech to text
  17. OpenAI Docs — Data controls in the OpenAI platform
  18. Speechmatics — Pricing
  19. Speechmatics Docs — Security and compliance
  20. Speechmatics Docs — Supported languages
  21. Speechmatics Docs — Batch limits
  22. Speechmatics — Privacy Policy
  23. Speechmatics — Security
  24. ElevenLabs — API pricing
  25. ElevenLabs — Scribe speech to text
  26. Descript — Pricing
  27. Descript — Privacy Policy
  28. Sonix — Pricing
  29. Sonix — Security
  30. Rev — Pricing
  31. Rev — Security
  32. Happy Scribe — Pricing
  33. Happy Scribe — Security
  34. Open ASR Leaderboard (arXiv 2510.06961)
  35. Open ASR Leaderboard — full paper text
  36. Open ASR Leaderboard — code repository
Advertisement

AI Tools Tutorial Team

Editorial

The editorial team behind aitoolstutorial.com. Every tool is checked against its vendor's own pricing and docs before anything is published, every source is linked at the foot of the article, and every recommendation names at least one thing the tool gets wrong.