Text-to-Speech Comes of Age: Deepgram Launches Conversation-Native Speech
Flux TTS, the newest model in Deepgram’s Flux family, brings conversation-native speech to enterprise voice agents in production as voice becomes the preferred interface for AI and the company surpasses $100M in ARR
SAN FRANCISCO--(BUSINESS WIRE)-- As voice agents move from impressive demos into business-critical interactions, enterprises need systems that can do more than sound natural. They must maintain context, adapt when conversations take an unexpected turn, and reliably complete critical tasks at production scale without requiring extensive human intervention. Deepgram, the real-time AI infrastructure company underpinning the Voice AI economy, today released Flux TTS, a conversation-native text-to-speech (TTS) model purpose-built as the conversation engine for enterprise voice agents. To experience Flux TTS firsthand, listen to voice samples and explore the developer documentation at https://deepgram.com/product/text-to-speech/flux.
Flux TTS joins Flux speech-to-text (STT) within Deepgram’s broader Voice AI platform and is backed by the same enterprise-grade runtime and deployment flexibility organizations already use to process speech with Deepgram at scale. Through Deepgram’s Voice Agent API, enterprises can orchestrate speech recognition, agent reasoning, and speech generation through a single API, reducing integration complexity, latency, and potential failure points that come with stitching together separate speech models and vendors.
Companies including Decagon, Sierra, Vapi, and Granola already use Deepgram’s voice infrastructure to power conversational AI experiences at scale. As Deepgram surpasses $100 million in annual recurring revenue, Flux TTS builds on that production foundation and Deepgram’s existing TTS capabilities with more expressive, conversation-native speech generation designed specifically for live interactions.
Built for Conversations that Don’t Follow a Script
Traditional TTS models were not built to sustain an ongoing conversation. Each request is treated as static output: text comes in, speech goes out, and the model resets. Real conversations are stateful: people pause, change direction, refer back to earlier turns, share complex information, and interrupt one another. When a voice agent loses that continuity–forgets a newly requested time midway through rescheduling an appointment, loses track of which troubleshooting steps a customer has already tried, or forces the customer to repeat themselves–it leads to failed tasks, human escalation, and an inability to deploy the agent in higher-stakes workflows.
Enterprises also face a broader systems challenge. Building a voice agent often requires connecting separate models for listening and speaking, and then managing the latency, orchestration, and potential failure points between them. Flux TTS is designed to reduce that burden.
“The market has spoken–literally–and voice is becoming the preferred interface for AI. But the next phase of this market will not be won by the model that sounds best in a demo. It will be won by systems enterprises can trust to complete business-critical work when conversations become unpredictable. Flux TTS was built for that standard: it begins responding in as low as 80 milliseconds while maintaining context and adapting to interruptions, moving text-to-speech from an audio feature to business-critical infrastructure,” said Scott Stephenson, CEO and Co-Founder of Deepgram.
Designed for Production
Flux TTS brings three core system behaviors to the speaking layer:
- Stays consistent from one turn to the next: Rather than treating each response as a standalone line of speech, Flux TTS carries the conversation forward, bringing context along for the entire conversation. Prior turns inform how the next response is spoken, helping the voice maintain consistent tone, pacing, and emotional register throughout the exchange. This helps the agent stay aligned with the exchange without additional prompt engineering, SSML, or style tags.
- Keeps pace with a live conversation: Flux TTS begins responding in as low as 80ms, even under production load. Persistent state, native interruption handling, and an explicit turn lifecycle allow the agent to adapt when someone interrupts or changes direction.
- Holds up under enterprise conditions: Flux TTS accurately communicates consequential information such as drug names, account numbers, and alphanumerics. Deployment options in a customer’s cloud or on premises also support scale, security, compliance, and data-residency requirements.
How Flux TTS Delivers Value for Customers
IBM and Coval are two examples of how Flux TTS is being used across the Voice AI ecosystem, from powering enterprise-grade agent experiences to helping teams evaluate how voice agents perform in real-world conversations. These examples underscore what production Voice AI increasingly demands: infrastructure that can maintain consistency, handle complexity, and perform reliably as agents take on more consequential customer interactions.
“Enterprises are moving from scripted automation to AI agents that can handle more complex customer interactions. Deepgram’s Flux TTS gives watsonx Orchestrate customers access to a text-to-speech model designed to maintain context and consistency across an entire conversation,” said Suzanne Livingston, VP of watsonx Orchestrate at IBM. “Flux TTS strengthens voice as a core part of the enterprise agent stack, and ultimately helps enterprises build authentic voice experiences better suited for real-world deployment across industries.”
“We evaluate voice agents every day at Coval, and the failure modes on the TTS side are remarkably consistent: tone that flattens by minute three, expressiveness that requires heavy SSML tuning, and inconsistency across turns that breaks the sense of a real conversation,” said Brooke Hopkins, Founder & CEO of Coval. “Deepgram Flux TTS is the first model we’ve seen that tackles those issues directly, treating the conversation as the unit rather than the individual line. That gives voice agent teams a TTS model that can hold up under the same rigorous, scenario-based evaluation we apply to the rest of the stack, and greater confidence in how their agents will perform in production.”
This reliability is critical across interactions like account management, where an agent needs to accurately communicate alphanumeric codes; sales and scheduling, where calls might be interrupted by new requests; restaurant ordering, where changes happen midway through the interaction; and technical assistance, where losing context or accuracy can force a customer to start over, abandon the task, or escalate to a human.
Deepgram Flux TTS is now generally available. Through September 12, 2026, developers can build with Flux TTS free with up to 45 concurrent streaming connections globally (5 in EU/AU). Standard pricing applies beginning September 13, 2026.
About Deepgram
Deepgram is the real-time AI infrastructure company underpinning the Voice AI economy. Today, more than 200,000 developers and 1,400 organizations are Powered by Deepgram. Its Voice AI platform offers speech-to-text (STT), text-to-speech (TTS), and full speech-to-speech (STS) capabilities, all powered by an enterprise-grade runtime. Deepgram’s voice-native foundation models, accessed through cloud APIs or as self-hosted/on-premises APIs, deliver unmatched accuracy, low latency, and competitive pricing. Customers include technology ISVs building voice products or platforms, co-sell partners working with large enterprises, and enterprises solving internal use cases. Having processed over 50,000 years of audio and transcribed over 1 trillion words, there is no organization in the world that understands voice better than Deepgram. To learn more, please visit www.deepgram.com, read its developer docs, or follow @DeepgramAI on X and LinkedIn.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260812602613/en/
PR Contact:
Launchsquad PR
[email protected]
Source: Deepgram
Serious News for Serious Traders! Try StreetInsider.com Premium Free!
You May Also Be Interested In
- McCarthy Breaks Ground on National Employee Development Campus
- PACE Equity Finance, Lone Star PACE Facilitate $1.7 Million C-PACE Project for North Texas Medical Office Building
- Crypto News: Pepeto Announces $10.626M Raised While the Bitcoin Price Prediction Targets $126,000
Create E-mail Alert Related Categories
Business Wire, Press ReleasesSign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share