Modulate Raises 25 Million Dollars To Teach AI To Hear What Transcripts Miss

Mike Pappas and Carter Huffman
Image credit: Modulate
Modulate, a Boston based voice intelligence company, has raised 25 million dollars to expand a platform that analyses speech directly from audio rather than relying on text transcripts.
The round was led by Future Ventures, with participation from Hyperplane and Lakestar. Lakestar has been involved for some time, having led a 30 million dollar Series A in 2022. PitchBook data cited by TechCrunch indicated Modulate had raised 41 million dollars before this round at a 170 million dollar valuation.
Modulate was founded in 2017 by Mike Pappas and Carter Huffman, who met as physics students at MIT. Huffman is now chief executive, having taken over from Pappas, who serves as chairman. The company first became known in gaming, where it began by offering voice modulation before moving into voice chat moderation. Its ToxMod system has been used since 2023 to moderate voice chat in Call of Duty, and its models still look for harassment and child grooming on social and gaming platforms.
The rise of voice AI has widened the opportunity. Modulate's main product, Velma, analyses audio signals to pick up emotion, tone, intent, emphasis and whether a voice is synthetic. Those signals can be combined to flag higher level events such as a fraud attempt, a harassment case or a customer losing patience with an automated agent, and some of that analysis runs live while a conversation is still under way.
The core argument is that transcription, while widely available, discards much of what matters in spoken conversation. Huffman has said many companies are doing transcription, but few offer the full nuance and understanding of a conversation, which matters when one person is talking to another. He has also said that voice is becoming a primary interface for AI, and that this creates problems a transcript cannot solve on its own. A sarcastic remark, a shaky voice or a synthetic caller can look identical to a genuine one once speech has been reduced to text.
Rather than depending on a single large model, Velma uses what Modulate calls an Ensemble Listening Model. It selects and combines more than 100 smaller, specialised audio models for each task, split broadly into signal extraction models that read vocal emotion, tone, language and synthetic voice markers, and analysis models that assess intent, rule violations and scam attempts. Modulate says this design is up to 1,000 times more efficient than a single large model, and that Velma is twice as accurate as general purpose language models at spotting real problems while producing seven times fewer false alarms. Those comparisons come from the company itself and have not been independently verified.
Public benchmarks offer a firmer check. In July, Modulate's transcription models placed first among 88 entries on Hugging Face's Open ASR Leaderboard, and its deepfake detector has ranked first on Hugging Face's Speech Deepfake Arena since March with an equal error rate of 1.1 percent. Batch transcription is priced at 3 cents an hour, a figure aimed at developers who need audio understanding at scale.
Deepfake defence is becoming a bigger part of the business as cloned voices turn into a routine fraud tool. Modulate says hospitals use its models to defend against deepfake callers, that its systems now process more than 10 million hours of audio a month, and that more than 600 million hours have been analysed in total. It also offers voice masking to protect staff in high risk roles. Beyond fraud, enterprise customers use the platform alongside their existing voice stacks to monitor AI agent compliance in regulated industries, understand why customer service calls succeed or fail, and track voice based cyberattacks. Detection of AI generated music is another capability listed in its suite.
Steve Jurvetson, co founder of Future Ventures, said the company had gained a significant technical lead in audio native AI and that demand was spreading well beyond gaming into AI agents, security and customer experience.
The new funding will go into research and engineering, along with a push to win developers through new SDKs, APIs and partner integrations, so that teams building voice products can buy audio understanding rather than train their own models. Modulate reportedly has 40 to 45 employees and plans to add about 10 more, while expanding on premises and on device deployment options to address privacy concerns.
The company operates in a crowded and fast moving space. Investors have been backing voice AI companies that make synthetic voices sound more human, while a separate group of startups focuses on detecting the intent behind human conversation and on protecting people from cloned voice fraud. Modulate's bet is that a single platform combining transcription, emotional analysis, deepfake detection and policy enforcement will appeal to enterprises that would otherwise stitch those capabilities together from several vendors, particularly as more customer conversations are handled by AI agents that themselves need supervision.
Topics
Stay informed
Startup news in your inbox
Get important funding rounds, founder stories, and startup updates.
No spam - only important startup updates.





