StepFun launches StepAudio 3, claiming top Artificial Analysis voice ranks
The AMW Read
Incremental product expansion by a known CN foundation-model lab into a full voice stack with third-party leaderboard leads, updating the player map without resolving a named open debate.
StepFun launches StepAudio 3, claiming top Artificial Analysis voice ranks
On September 15, StepFun released its StepAudio 3 speech-model family—Realtime, ASR, TTS, Gen, and Music—spanning full-duplex conversation, speech recognition, human-like synthesis, multi-element audio generation, and interactive music creation. Per the report, StepAudio 3 Realtime scored 98.9% and ranked first on Artificial Analysis Conversational Dynamics and also led Speech Reasoning; StepAudio 3 ASR recorded a 1.7% word error rate, tying for first on non-streaming recognition accuracy. The lineup is live on StepFun's open platform and is positioned beyond single-task ASR or TTS toward voice understanding, realtime reasoning, tool calls, and audio content production.
That framing matters because foundation-model competition is widening from text and vision into a full voice stack where conversational dynamics, interrupt handling, and speech reasoning sit beside generation quality. Third-party leaderboard leads give buyers a shared scoreboard for realtime voice APIs just as Chinese labs intensify productization of listen-speak-create pipelines for assistants, media, and interactive apps. StepFun, founded in 2023 with $3.8B in total funding per the AI Market Watch index (coverage of ~5,000 companies, not a census), is using that scale to cover recognition, dialogue, synthesis, scene audio, and music in one release rather than a narrow niche model.
For builders and investors, the practical bar is moving: evaluate whether open-platform access, streaming TTS, asynchronous tool use during live talk, and natural-language multi-track audio generation convert these Artificial Analysis wins into sticky API and content-workflow usage—not only whether WER or conversational scores stay at the top of a weekly chart.
#StepFun #StepAudio3 #SpeechAI #FoundationModels #ArtificialAnalysis #VoiceModels

