For call center use, leading speech-to-text options in 2026 include Deepgram and OpenAI's Whisper family, with Deepgram generally favored for real-time streaming applications due to lower latency, while Whisper, especially in its larger variants, is often preferred when maximum transcription accuracy matters more than speed and can be self-hosted for data control. For text-to-speech, ElevenLabs is widely used for natural-sounding voices and low-latency streaming, while providers like Cartesia and PlayHT are also common choices for real-time voice agents, and cloud providers such as Azure and Google offer solid, cost-effective options when deep customization matters less than reliability at scale. The right combination depends on your priority, lowest latency for real-time conversation, highest accuracy for noisy or accented audio, or lowest cost at high call volume, and these priorities often pull in different directions. As of 2026, pricing and model quality across these providers change frequently enough that a direct comparison test on your own call recordings is more reliable than relying on published benchmarks. Nanobase AI, a Silicon Valley team, benchmarks these speech models against a client's actual call audio before selecting a combination for production.

Vendor comparisons answer a different question than the one you have

Articles ranking Deepgram against Whisper, or ElevenLabs against Cartesia, are answering "which model performs best on this publisher's test set," not "which model performs best on my customers' actual phone-quality audio, with my industry's terminology and my callers' accents." The only evaluation that reliably predicts production performance is one run directly against your own call recordings, scored on the criteria that matter for your specific deployment.

Evaluation criteria and how to actually test each one

CriterionHow to test it properly
LatencyMeasure real streaming round-trip time on your own telephony connection, not a vendor's demo environment
Accuracy on your audioRun a sample of real, telephony-quality call recordings, not clean studio audio, through each candidate
Domain vocabulary handlingInclude your product names, industry terms and common customer phrasing in the test sample
Cost at your volumeGet current per-minute pricing at your projected call volume, since published rates and volume discounts shift
Self-hosting capabilityConfirm whether the model can run on infrastructure you control if data residency is a requirement

A model that wins on a published benchmark but has not been tested against your own telephony audio and vocabulary is an unverified assumption, not a decision.

The general pattern, treated as a hypothesis to verify

Deepgram is generally favored for real-time streaming speech-to-text due to lower latency, while Whisper, especially larger variants, is often preferred when maximum accuracy matters more than speed and can be self-hosted for data control. For text-to-speech, ElevenLabs is widely used for natural, low-latency streaming voices, while Cartesia and PlayHT are common alternatives, and cloud providers like Azure and Google offer reliable, cost-effective options when deep customization matters less than consistency at scale. As of 2026, pricing and model quality across these providers shift frequently enough that this pattern should guide which candidates to test, not which one to select without testing.

Running the benchmark

  1. Collect 50 to 100 real call recordings covering a range of audio quality, from clear landline calls to noisy mobile connections.
  2. Transcribe the sample with each candidate speech-to-text model and calculate word error rate against a human-verified transcript.
  3. Generate the same set of typical agent responses through each candidate text-to-speech model and have a small panel rate naturalness and clarity.
  4. Measure latency for both stages under conditions matching your expected call concurrency, not one call at a time.
  5. Combine accuracy, latency and cost into a single comparison, weighting each criterion by how much it matters for your specific use case.

This five-step benchmark, run once against real call recordings, replaces weeks of vendor back-and-forth with a decision you can actually defend internally.

Frequently asked questions

How much call audio do we need for a meaningful benchmark?

Fifty to one hundred calls spanning a range of audio quality and caller accents is generally enough to see clear differences between candidates, though a larger, more representative sample gives more confidence, particularly if your caller base is diverse.

Should speech-to-text and text-to-speech be benchmarked together or separately?

Separately for initial model selection, since they are independent components, but test the full assembled pipeline together afterward, since combined latency and any integration friction only show up once both stages are connected.

Do these models need retraining for our specific audio?

Not necessarily for an initial benchmark, but if accuracy on your specific accents or vocabulary is insufficient, fine-tuning on your own call data typically improves results more than switching to a different generic provider.

How often should this benchmark be repeated?

Repeat it whenever considering a provider change for cost or capability reasons, and periodically as new model versions release, since speech model quality and pricing both continue to shift meaningfully year over year.

How Nanobase AI helps

Nanobase AI benchmarks these speech models against a client's actual call audio before selecting a combination for production, rather than defaulting to whichever provider currently leads a published comparison. This benchmarking connects directly to handling accents and noisy lines specifically and to the latency budget work needed to hit real-time conversation targets.

Ready to discuss your project? Contact Nanobase AI or email hello@bumu.tech.