What is speech to text?
Speech to text (STT) is a technology that converts spoken language into written text. In a contact center context, if a customer says “I need to check the status of my order from last Tuesday,” STT captures that audio and produces a text transcript. That transcript is then the input for every downstream system: the AI agent that responds, the analytics platform that tracks contact drivers, the agent assist tool that surfaces relevant knowledge, the QA system that scores the interaction.
STT is infrastructure that enables much of contact center technology. A contact center that has invested in AI voice agents, conversation intelligence or real-time agent assist has implicitly invested in STT, whether the vendor makes that explicit or not. The quality of the STT layer directly determines the quality of everything built on top of it.
How speech to text works
Modern STT systems process audio through several stages before producing a transcript:
- Audio preprocessing: Background noise is reduced, volume normalized and, for two-speaker calls, the agent and customer audio streams are separated so each speaker is transcribed independently.
- Acoustic modeling: A neural network maps the audio’s acoustic features to phonemes, the basic units of spoken language.
- Language modeling: A statistical model uses knowledge of how words and phrases typically co-occur to disambiguate between phonetically similar options and produce the most likely word sequence. “Two” versus “too” versus “to,” for example, is resolved by context.
- Transcript output: The final text is produced with timestamps and, in most enterprise systems, confidence scores that downstream analytics platforms use to flag lower-certainty segments for review.
In real-time speech to text, the transcript is available to an AI agent or agent assist system in near real-time, before the next turn in the conversation. In post-call transcription, the same process runs without latency constraints, which typically produces higher accuracy.
Real-time vs. post-call speech to text
STT operates in two modes in contact centers, each serving different use cases:
Real-time STT produces transcripts as the conversation unfolds. This is required for AI voice agents (which need to understand the caller before responding), real-time agent assist (which surfaces guidance during the call) and live sentiment monitoring. Speed is the priority, so the system transcribes with less context than post-call processing. This can result in lower accuracy.
Post-call STT transcribes recorded audio after the interaction ends. Without the need to instantly parse input, the system can apply more sophisticated language models, reprocess ambiguous segments and produce a higher-accuracy transcript. This mode is used for conversation intelligence, Auto QA, compliance monitoring and coaching workflows.
Most enterprise contact center platforms run both: real-time STT powering in-call capabilities, post-call STT powering analytics and QA at higher accuracy.
What affects STT accuracy in contact centers?
Word error rate (WER), the percentage of words in the transcript that differ from what was actually said, is the standard accuracy metric for STT. Modern systems achieve 3–6% WER in clean, controlled conditions. In real contact center environments, WER typically increases to 8–15% due to:
- Telephony audio compression: Phone calls have a lower sampling rate, 8kHz compared to the typical 44kHz used in high-quality audio recording. Compression removes acoustic detail that helps distinguish phonetically similar sounds.
- Accent and dialect variation: STT systems trained primarily on standard American English underperform on regional accents, non-native speakers and dialectal variation.
- Background noise: Customers calling from cars, public spaces or noisy home environments introduce audio quality challenges that directly increase error rates.
- Domain vocabulary: Product names, account identifiers, industry-specific terminology and company-specific language require domain-tuned models to transcribe accurately.
- Cross-talk: Speakers talking simultaneously creates overlapping audio that degrades accuracy for both speakers’ channels.
Domain-tuned STT models, trained on actual contact center audio in a specific industry vertical. substantially outperform general-purpose models in production environments. When evaluating voice AI vendors, ask how their speech to text layer handles your specific vocabulary, customer demographic and audio conditions.
Speech to text vs. text to speech
These are inverse processes that together form the voice layer of any AI voice agent:
- Speech to text (STT): Converts customer audio into text so the AI can understand what was said.
- Text to speech (TTS): Converts the AI’s text response into spoken audio so the customer can hear it.
STT and TTS quality both affect the perceived naturalness of the interaction, but STT failures are more consequential. One transcription error can produce a wrong or nonsensical response later in the interaction, regardless of how good the NLU and TTS layers are.
How Capacity uses speech to text
STT is a foundational layer beneath Capacity’s voice capabilities. The AI Voice Agent uses real-time STT to understand callers in natural conversation, with no menu navigation or keyword constraints. Transcription converts call audio into searchable, timestamped transcripts for real-time agent assist, allowing agents to help customers more seamlessly. The same STT powers automated call summaries, 100% QA scoring and conversation intelligence, allowing contact centers to automate the entire customer journey with one platform.
See how Capacity’s unified CX automation platform works →
Frequently asked questions about speech to text
Yes, though model quality varies significantly by language. Major world languages like Spanish, German or Japanese have well-developed models. Less widely spoken languages have fewer training data resources and typically higher error rates. For contact centers serving multilingual customer populations, verifying STT accuracy across all target languages in real call conditions is an important vendor evaluation step.
The standard metric is word error rate (WER) divided by the total number of words in the reference. A WER of 10% on a 600-word call means approximately 60 words were transcribed incorrectly. It’s important to consider not only how often errors occur but where, such as in product names, account numbers, accents, et cetera.
It depends on the platform and configuration. Some STT systems transcribe audio in real time and discard the raw audio immediately; others store recordings alongside transcripts for QA review and coaching. Data retention policies, storage locations and access controls vary by vendor and are governed by a combination of the vendor’s data practices and the contact center’s own compliance requirements. For industries with data retention regulations, like financial services or healthcare, reviewing a vendor’s recording and retention architecture is a compliance requirement before deployment.
Voice recognition, or voice biometrics, identifies who is speaking based on the unique acoustic characteristics of their voice. STT identifies what was said, not who said it. A contact center might use both: STT to understand the content of what a caller says and speaker recognition to verify their identity before giving them access to account information.