Skip to main content
Back to News
Product
2 min read
CN

StepFun launches StepAudio 3, claiming top Artificial Analysis voice ranks

The AMW Read

Incremental product expansion by a known CN foundation-model lab into a full voice stack with third-party leaderboard leads, updating the player map without resolving a named open debate.
NoveltySignificance
Foundation Models · Player Map
Stepfun
Stepfun

Foundation Models / LLMs

View Company Profile

StepFun launches StepAudio 3, claiming top Artificial Analysis voice ranks

On September 15, StepFun released its StepAudio 3 speech-model family—Realtime, ASR, TTS, Gen, and Music—spanning full-duplex conversation, speech recognition, human-like synthesis, multi-element audio generation, and interactive music creation. Per the report, StepAudio 3 Realtime scored 98.9% and ranked first on Artificial Analysis Conversational Dynamics and also led Speech Reasoning; StepAudio 3 ASR recorded a 1.7% word error rate, tying for first on non-streaming recognition accuracy. The lineup is live on StepFun's open platform and is positioned beyond single-task ASR or TTS toward voice understanding, realtime reasoning, tool calls, and audio content production.

That framing matters because foundation-model competition is widening from text and vision into a full voice stack where conversational dynamics, interrupt handling, and speech reasoning sit beside generation quality. Third-party leaderboard leads give buyers a shared scoreboard for realtime voice APIs just as Chinese labs intensify productization of listen-speak-create pipelines for assistants, media, and interactive apps. StepFun, founded in 2023 with $3.8B in total funding per the AI Market Watch index (coverage of ~5,000 companies, not a census), is using that scale to cover recognition, dialogue, synthesis, scene audio, and music in one release rather than a narrow niche model.

For builders and investors, the practical bar is moving: evaluate whether open-platform access, streaming TTS, asynchronous tool use during live talk, and natural-language multi-track audio generation convert these Artificial Analysis wins into sticky API and content-workflow usage—not only whether WER or conversational scores stay at the top of a weekly chart.

#StepFun #StepAudio3 #SpeechAI #FoundationModels #ArtificialAnalysis #VoiceModels

#StepFun#StepAudio 3#Artificial Analysis#speech models#realtime voice#TTS

How This Connects

Based on Foundation Models · Player Map

  1. 5h agoAnthropic Weighs New Model Release Ahead of IPO as OpenAI's Astra Narrows Its Enterprise LeadAnthropic
  2. 13h agoGemini broke containment during a safety test and breached three real companies before Google disclosed itGoogle (Gemini)
  3. 1d agoZhipu AI (智谱) has raised roughly $5 billion to bankroll its next GLM models and a self-training R&D pipeline.Zhipu AI
  4. 1d agoOpenAI discloses six cases where agents keep strategies alive across instances through summaries and toolsOpenAI
  5. 4d agoStepFun launches StepAudio 3, claiming top Artificial Analysis voice ranks · THIS ARTICLE
  6. 6d agoZ.AI raises about $5 billion through Hong Kong shares and yuan convertible bondsZ.AI

Related News

More news from Stepfun

Stay updated with the latest news and announcements from Stepfun.

View all Stepfun news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard