# Modulate Raises $25M for Voice AI That Listens to Audio, Not Transcripts. Here's What It Sells and Which Claims You Can Check

The Boston company behind Call of Duty's ToxMod voice moderation is pitching its Velma models to teams running AI voice agents and fighting deepfake calls. Some of its numbers are public benchmarks; others are its own.

By Priya Nair (AI Industry Reporter) — published 2026-09-29, updated 2026-09-29
Source: https://aiscoutdaily.com/news/modulate-raises-25m-audio-native-voice-ai

## TL;DR
Modulate raised $25 million on September 28 in a round led by Future Ventures, with Hyperplane and Lakestar, and says its total funding is now $60 million. Its Velma platform analyzes voice audio directly for emotion, intent, synthetic speech, and policy violations, using more than 100 small specialized models instead of one large one. Its first place on Hugging Face's Open ASR Leaderboard and Speech Deepfake Arena can be checked publicly; its accuracy and efficiency multiples versus LLMs are its own claims. Batch transcription costs $0.03 an hour.

Most voice AI money in 2026 has gone to making machines talk. Modulate is raising on the other half of the conversation: understanding what a caller actually means, whether the voice is real, and whether an AI agent on the line is doing its job. The company, founded in 2017 by MIT physics graduates Mike Pappas and Carter Huffman, started in game voice chat and now sells to contact centers, banks, insurers, and healthcare organizations.

## The round

- Amount: $25 million, announced September 28, 2026.
- Investors: led by Future Ventures, with Hyperplane and Lakestar (which led Modulate's $30 million Series A in 2022).
- Total raised: $60 million, according to Modulate. TechCrunch, citing PitchBook, reported $41 million raised at a $170 million valuation before this round; the two figures don't reconcile, and no new valuation was disclosed.
- Team: 40-45 employees, with plans to add about 10, mainly in model building, per TechCrunch.
- Use of funds: research, engineering, new SDKs and APIs, developer relations, partner integrations, and more on-premises and on-device deployment options.

## What Modulate sells

Its Velma platform works on the audio itself rather than a transcript. Signal models pick up emotion, tone, language, accent, emphasis, and whether a voice is synthetic; analysis models combine those signals to spot higher-level events such as a fraud attempt, harassment, a frustrated customer, or a voice agent misunderstanding someone. It can run in real time, so a system can step in during a call. Modulate calls the architecture an Ensemble Listening Model, which picks from more than 100 small models per task rather than sending audio through one large model.

- Fraud and deepfake detection for contact centers, banks, and hospitals.
- Supervision of AI voice agents: whether the agent followed policy and whether the customer left satisfied.
- Trust and safety: ToxMod has moderated Call of Duty voice chat since 2023, and the models also detect grooming and harassment.
- Voice masking to protect staff in high-risk roles.

## Which claims you can check

| Claim | Source | Checkable? |
| --- | --- | --- |
| #1 on Hugging Face's Open ASR Leaderboard (out of 88 entries in July, per TNW) | Public leaderboard | Yes |
| #1 on Hugging Face's Speech Deepfake Arena, 1.1% equal error rate (per TNW) | Public leaderboard | Yes |
| 98.9% deepfake detection accuracy on public benchmark data | Modulate | Partly - depends on the dataset |
| 2x the accuracy of LLMs at true positives, 7x fewer false positives | Modulate | No - internal comparison |
| Up to 1,000x more efficient than a single large model | Modulate | No - internal comparison |
| 10 million+ hours analyzed per month, 600 million+ in total | Modulate | No - company figure |

## Pricing

Modulate lists batch transcription at $0.03 per hour of audio, with pay-as-you-go API access and enterprise plans. Its other APIs (deepfake detection, emotion, accent, language, audio event detection, PII redaction) are priced on its pricing page rather than in the announcement.

## What this means if you run voice agents

Transcript-only analytics miss things that matter on calls: a caller who stays polite but is clearly unhappy, or a cloned voice reading a convincing script. If you're deploying an AI voice agent or handling payments over the phone, an audio-native layer like Modulate's is worth piloting next to your existing transcription. Judge it on your own calls: the public leaderboard wins are a good sign for transcription and deepfake detection, but the headline multiples against LLMs are Modulate's own measurements.

## Sources
- [Modulate: Modulate Raises $25M to Scale Its Lead in Frontier Audio-Native AI (Sep 28, 2026)](https://www.modulate.ai/press-releases/modulate-raises-25m-to-scale-its-lead-in-frontier-audio-native-ai)
- [TechCrunch: Modulate raises $25M for its voice models and analysis suite](https://techcrunch.com/2026/09/28/modulate-raises-25m-for-its-voice-models-and-analysis-suite/)
- [The Next Web: Modulate raises $25M to sell AI that listens to voice calls, not just transcripts](https://thenextweb.com/news/modulate-25m-funding-audio-native-ai)
