For medical transcription, Deepgram and AssemblyAI land close on capability and split on terms. Both ship a clinical speech-to-text mode, both do speaker diarization and custom medical vocabulary, and both run under a HIPAA business associate agreement. The decision isn’t which one transcribes better in the abstract. It’s which deployment and BAA terms fit your build. Deepgram signs BAAs for Enterprise customers handling ePHI and offers on-premises and VPC deployment; AssemblyAI offers a BAA with no premium pricing and no sales call. That difference decides more real builds than any accuracy benchmark.
We’re gmware, a custom software development firm in Austin, TX with engineering centers in Bangalore and Mohali, India, and we integrate speech-to-text into healthcare products as part of our delivery history. What follows is the comparison we run with clients before picking an ASR: where the two are genuinely even, where they diverge, and the case where neither is the right answer.
One opinion up front. Stop choosing on the vendor’s published word error rate. Both companies benchmark on their own test audio, and both win on it. The only accuracy number that means anything is the one you measure on your own recordings, in your specialties, with your microphones and your room noise. Everything else is marketing.
A note on pricing. Speech-to-text rates are low, tiered, and change often. The figures below trace to each vendor’s current published pricing at writing, but treat them as patterns, not quotes. Verify current pricing and current BAA terms with each vendor before you build a cost model or sign anything.
| What you’re comparing | Deepgram | AssemblyAI |
|---|---|---|
| Dedicated medical model | Nova-3 Medical | Medical Mode (add-on) |
| HIPAA BAA | Enterprise tier, ePHI | Included, no premium, no sales call |
| Deployment | Managed cloud, on-prem, VPC | Managed cloud |
| Streaming + batch | Both | Both |
| Speaker diarization | Yes (add-on) | Yes (add-on) |
| Custom vocabulary | Keyterm Prompting | Keyterms Prompting |
Where they’re genuinely even
Start with what won’t decide it, so you don’t waste a bake-off on it. Both vendors ship a medical-tuned model: Deepgram’s is Nova-3 Medical, built for clinical audio, and AssemblyAI’s is Medical Mode, which raises accuracy on medication names, procedures, conditions, and dosages. Both let you add custom terms to catch specialty vocabulary the base model misses. Both do speaker diarization so a transcript separates clinician from patient. Both run streaming for live capture and batch for uploaded files.
On capability, in other words, this is close to a coin flip, and that’s the point. The features that look like the comparison in a spec sheet are table stakes. The things that actually differ are deployment, BAA terms, the billing model, and how each performs on your specific audio.
Accuracy: ignore the benchmark, run your own
Both vendors publish medical accuracy numbers, and they’re impressive. Deepgram claims a median word error rate around 3.4% for Nova-3 Medical, with large stated gains over competitors. AssemblyAI publishes similarly strong figures for Medical Mode. Here’s the problem with both: a published word error rate is measured on the vendor’s own curated test set, and your clinic is not that test set.
Word error rate also undersells what goes wrong in medicine. A model can score a low overall error rate and still fumble the words that carry clinical weight, the drug names, dosages, and procedure terms where a single miss is the one that matters. That’s why both vendors report keyword accuracy separately from overall accuracy, and why your benchmark should weight the medical terms heavily. Pull a representative sample of your own recordings, across your specialties and accents and the microphone setup you’ll actually deploy, run all the candidates against it, and read the editing time, not the vendor’s slide. The model that wins a clean benchmark can lose badly on a noisy exam-room recording.
Streaming versus batch, and the billing trap
Both offer streaming and batch, so the question is which your workflow needs. Live use, an ambient scribe listening to a visit, a telemedicine call captioned in real time, needs streaming. Uploaded recordings and overnight processing can run batch, which is usually cheaper.
The trap is in how streaming is billed. AssemblyAI bills streaming on WebSocket session duration, which is open-to-close connection time rather than the audio you actually send. An idle connection left open is still on the meter, so a sloppy client that holds sockets open between utterances can run up a bill that has nothing to do with how much anyone said. Whatever you pick, model your real connection pattern, not just your audio minutes, and confirm each vendor’s current streaming and batch rates, because the streaming rate is typically several times the batch rate.
HIPAA and the BAA: the real divider
This is where the choice usually gets made. Both vendors will cover you under HIPAA, but the shape of that coverage differs, and it maps cleanly onto two kinds of teams.
Deepgram signs BAAs for Enterprise customers handling ePHI and offers on-premises and VPC deployment. If your requirement is that patient audio never leaves a boundary you control, that combination is the draw: HIPAA coverage plus the ability to run the model inside your own infrastructure. The cost is that the BAA lives on the Enterprise tier, so you’re in a sales conversation and a higher commitment to get it.
AssemblyAI offers a BAA with no premium pricing and no sales call. If your requirement is to move fast on managed cloud and get HIPAA coverage without a procurement cycle, that’s the draw. The trade is that it’s managed cloud: your audio goes to their infrastructure, which is fine for most builds and a dealbreaker for the ones with a hard data-residency or on-prem mandate.
Neither posture is better. They fit different constraints. Before you commit ePHI to either, confirm the current BAA terms directly, because this is exactly the kind of thing vendors revise. Whichever you pick, the BAA is one link in a chain, not the whole thing. Our HIPAA-compliant architecture guide covers why every vendor that touches PHI, not just the ASR, needs its own signed agreement.
When a specialized medical ASR beats both
Here’s the honest verdict the spec sheets won’t give you: sometimes neither general API is the right call. Deepgram and AssemblyAI are excellent general speech-to-text engines with a medical mode layered on. A specialized medical ASR is trained from the ground up on clinical speech, on the dense, jargon-heavy, drug-and-procedure-laden way clinicians actually dictate.
When does the specialist win? When the dictation is dense and specialty-specific, when the volume is high enough that a few points of accuracy turn into real editing hours, and when the cost of an editor fixing transcripts outweighs the higher price of the better model. The trade is usually a higher per-unit cost, fewer deployment options, or a heavier integration. The way to settle it is the same as everything else here: benchmark the domain-tuned option alongside Deepgram and AssemblyAI on your own audio, and let the editing-time difference, in real hours and real dollars, make the call. Don’t assume the general API is cheaper once you count the human cleaning up after it.
How gmware picks and integrates an ASR
We treat ASR selection as a measured decision, not a brand preference, inside our AI agents and LLM integration and healthcare software development practices. The process is the same every time: define the workflow (streaming or batch, real-time or overnight), set the HIPAA constraint (managed cloud is fine, or audio must stay on-prem), then benchmark the real candidates, Deepgram, AssemblyAI, and a domain-tuned medical option, on a representative sample of the client’s own recordings before anyone signs.
That benchmark is the whole game, and it’s where our Austin-based leads and Bangalore and Mohali engineers earn their keep: building the harness, weighting clinical-term accuracy, and reading editing time instead of vendor slides. If the transcription feeds clinical notes, our SOAP note automation cost guide covers what sits on top of the ASR, and our AI medical answering service work shows the same speech pipeline in a front-desk setting.
Tell us what the audio is and where it has to stay. Send us the shape of it and we’ll come back within 48 hours with a straight read on which ASR fits, what to benchmark, and the HIPAA terms to verify before you commit.