Deepgram
Speech recognition and voice-generation API that lets developers add accurate, fast transcription and natural-sounding speech to any app.
🔗 Visit DeepgramDescription
Turning spoken words into text — and text back into natural-sounding speech — sounds simple until you actually try building it: real conversations have background noise, accents, interruptions, and need to be transcribed in a fraction of a second for a live phone call to feel natural. Deepgram is a ready-made engine for both directions of that problem, built specifically so developers don't have to train their own speech models from scratch. Deepgram provides speech-to-text and text-to-speech APIs built for real-time and batch use, powering voice AI applications, contact centers and conversational AI products. Its Voice Agent API bundles transcription, text-to-speech and LLM orchestration into a single integration point, and the platform includes automatic language detection across roughly ten languages, audio intelligence features like summarization and sentiment analysis, and speaker diarization with redaction for compliance-sensitive use cases. Pricing is pay-as-you-go with $200 in free credit to start: speech-to-text runs roughly $0.0048-$0.0065 per minute, text-to-speech about $0.015-$0.030 per 1,000 characters, and the bundled Voice Agent API around $0.075 per minute, with a Growth plan starting near $4,000/year and custom Enterprise pricing above that.
💬 Our review
The short version: Deepgram is one of the more established, developer-first names in speech AI, and having both transcription and speech generation under one API — plus a bundled Voice Agent endpoint — reduces the number of vendors a team building a voice product has to stitch together.
It competes in a genuinely crowded 2026 field: AssemblyAI focuses more narrowly on transcription accuracy and audio intelligence, while newer specialists like Cartesia and Rime push harder on ultra-low-latency, highly expressive speech generation specifically for real-time voice agents. Deepgram's advantage is breadth — one vendor, one bill, for both STT and TTS — rather than being the single best option on either axis alone. The per-minute and per-character pricing is competitive but not obviously cheaper than specialists, so the honest reason to pick Deepgram over a best-of-breed combination is integration simplicity, not a clear cost or quality edge on either individual capability.
💰 Pricing
📊 Global score
🤖 AI-enriched data
Pros
Both speech-to-text and text-to-speech under one API and one bill
Bundled Voice Agent API combining STT, TTS and LLM orchestration
Audio intelligence features (summarization, sentiment, diarization) built in
Cons
Not the clear latency or expressiveness leader against specialists like Cartesia/Rime
Pricing is competitive but not a standout discount versus focused competitors
Real-world cost with Voice Agent API ($0.075/min) adds up for high call-volume products
❓ Frequently asked questions
- Does Deepgram do both speech-to-text and text-to-speech?
- Yes — it offers both directions under one API, plus a bundled Voice Agent endpoint that combines transcription, speech generation and LLM orchestration for full voice pipelines.
- How fast is Deepgram's transcription?
- It supports real-time streaming transcription in addition to batch processing, which is what makes it usable for live phone calls and voice agents, not just after-the-fact transcripts.
- Can Deepgram detect who's speaking in a call?
- Yes — speaker diarization is included, along with redaction features for removing sensitive information from transcripts, useful for compliance-sensitive industries.
- Is there a free way to try Deepgram?
- Yes, new accounts get $200 in free credit, and pricing beyond that is pay-as-you-go per minute or per character rather than a mandatory subscription.
- Is it worth the money compared to alternatives?
- Per-unit pricing is competitive but not dramatically cheaper than specialists like Cartesia or Rime. The value is in not needing separate vendors for STT and TTS — worth it if integration simplicity matters more than squeezing out the absolute best latency or voice quality on one axis.
- Which tool should you pick for your case?
- Want one vendor for both transcription and speech generation: Deepgram. Need the lowest possible speech-generation latency for a voice agent: Cartesia. Want the most natural, expressive voices at scale: Rime. Need the deepest transcription accuracy and audio intelligence specifically: AssemblyAI.
