
ElevenLabs has released Dubbing v2 as an API, enabling developers and enterprises to embed the compa...
The AMW Read
API launch is a meaningful extension of an existing model, but not a paradigm shift; impacts the generative-media segment and enterprise localization workflows.
ElevenLabs has released Dubbing v2 as an API, enabling developers and enterprises to embed the company's audio-to-audio dubbing model directly into their own products and workflows. The announcement follows the model's initial UI-only debut on May 28, with the API opening a path to integration for external applications. Dubbing v2 converts speech directly to speech rather than relying on the traditional speech-recognition-to-translation-to-synthesis pipeline, which the company says better preserves the original performance and emotional expression. The model supports 92 languages, including Korean and Japanese, and is designed to match speaker voice and delivery timing to video, though it does not include lip-sync or real-time streaming capabilities.
The move expands ElevenLabs' enterprise reach by monetizing its dubbing technology via API, a strategy that aligns with the company's broader push into media localization and its reported $330M ARR in 2025. By offering high-quality dubbing at a fraction of the traditional cost (which the company says can run hundreds of dollars per minute for film and TV), ElevenLabs is positioning itself as a scalable alternative for video creators, game studios localizing dialogue, consumer brand marketing, and streaming services deploying content across regions. This API launch also intensifies competition in the AI dubbing space, where rivals like Hudson AI are also enhancing emotion-preserving features.
For developers and media companies, this API lowers the barrier to integrating multilingual, emotion-aware dubbing into their own platforms, potentially disrupting conventional localization workflows and making global content distribution more accessible. The focus on preserving emotional nuance could differentiate ElevenLabs in a market where synthetic voices often lose expressiveness. As the company continues to scale—recently raising $500M at an $11B valuation—this API release represents a concrete step toward turning its technology into a platform-level service, though it stops short of real-time or lip-synced offerings, leaving room for future iterations.
