The short version
- Modulate announced $25 million in new funding led by Future Ventures, with Hyperplane and Lakestar participating.
- The company says its models analyze raw audio rather than relying only on transcripts, allowing them to use signals such as emotion, tone, intent and synthetic speech.
- Modulate says more than 10 million hours of audio are processed through its systems each month and that lifetime processing has passed 600 million hours.
Modulate announced $25 million in new funding led by Future Ventures, with Hyperplane and Lakestar participating. Voice software has traditionally converted speech into text and then handed that transcript to a language model. That approach works for many commands, but it also discards information contained in the original audio. Tone, emotion, timing and other acoustic characteristics can matter when software is trying to understand what a person actually meant.
The case for listening to raw audio
The company’s Velma platform uses an Ensemble Listening Model architecture built from more than 100 specialized audio models.
Modulate’s product material explains its audio-first approach and its focus on raw audio rather than treating the transcript as the whole signal.
Modulate is taking a different route by building models that work directly with audio. The company says its systems can identify signals including emotion, tone, intent and synthetic speech. That makes the technology relevant to applications where the audio itself is part of the security or product experience.
- Tone, emotion, timing and other acoustic characteristics can matter when software is trying to understand what a person actually meant.
- The company says its models analyze raw audio rather than relying only on transcripts, allowing them to use signals such as emotion, tone, intent and synthetic speech.
- The company says its systems can identify signals including emotion, tone, intent and synthetic speech.
Modulate said it plans to expand research, engineering, developer relations, APIs, SDKs and partner integrations.
Audio-native processing matters because converting a conversation to text can discard information that is present in the original signal. A transcript can record the words while losing vocal stress, timing, background characteristics and evidence that a synthetic voice may have been used.
Modulate came from the gaming and voice-moderation market, where systems have to process large quantities of speech and identify harmful or suspicious behavior quickly. That background helps explain why the company is emphasizing live analysis rather than only post-conversation transcription.
The new funding gives Modulate more resources to bring that capability to developers. The developer market is important because audio intelligence becomes more useful when it can be embedded into games, communications products, moderation systems and other applications rather than remaining a standalone demonstration.
The developer focus is significant because audio models are most useful when applications can integrate them directly. APIs, SDKs and deployment options determine whether the technology becomes a component inside another product rather than a standalone demonstration.
The commercial challenge is accuracy in difficult conditions. Real audio contains accents, background noise, interruptions and deliberately manipulated speech. Developers will need predictable performance across those conditions before audio-native models can become a routine component of production software.