Fish Audio
Category: Voice / Speech AI
AI voice generation platform offering expressive text-to-speech, voice cloning, and speech-to-text with fine-grained emotional controls, challenging ElevenLabs with open-source backing and enterprise features. Fish Audio was founded in 2025. The company is led by Rissa Cao. Based in Palo Alto, California, United States. Team size: 11-50. Total funding raised: $52M. Latest round: Seed. Key investors include Coreline Ventures, Capital Today, 359 Capital, Parable, Play Time, Alphaist Partners, Bayhouse Ventures, Carya Venture Partners, HF0, 645 Ventures.
- Founded
- 2025
- Headquarters
- Palo Alto, California, United States
- Team size
- 11-50
- Total funding
- $52M
Value proposition
Most expressive, emotionally controllable real-time voice model with fine-grained word-level emotion controls via 15,000+ natural language tags; open-source community trust with 31K+ GitHub stars; 67% listener preference in blind tests vs competitors; 11x lower cost than ElevenLabs; voice cloning from 5-second audio in 15 seconds; 83 languages supported
Products and solutions
Fish Audio S2.1 Pro (flagship expressive TTS model), Fish Speech (open-source TTS with 31K+ GitHub stars), Voice Cloning (from 5-second audio sample), Speech-to-Text (with multi-speaker diarization & emotion tags), Voice Agent (end-to-end voice agent solution), Text-to-Speech API (ultra-low latency, streaming), 2M+ user-uploaded voice library, 83 language support, 15,000+ natural language emotion controls
Unique value
Combines open-source transparency with enterprise-grade quality — the only voice AI platform offering both a top-ranked open-source model (Fish Speech) and a proprietary flagship (S2.1 Pro) that outperforms closed competitors in blind listening tests at a fraction of the cost
Target customer
Content creators (YouTubers, audiobook narrators, game developers), indie developers, and enterprise customers in regulated industries (healthcare, financial services) needing voice AI for customer support, AI avatars, and conversational agents
Industries served
Content Creation (video voiceovers, audiobooks), Gaming (character voices), Customer Support (conversational chatbots), AI Avatars, Healthcare (HIPAA-compliant deployments), Financial Services
Technology advantage
Dual-Autoregressive (Dual-AR) architecture — 4B-parameter Slow AR for semantic prediction + 400M-parameter Fast AR for acoustic generation, structurally isomorphic to standard LLMs enabling SGLang acceleration (Continuous Batching, Paged KV Cache, CUDA Graph); trained on 10M+ hours of audio across 80+ languages; GRPO reinforcement learning alignment; open-sourced S2 model under Fish Audio Research License; S2.1 Pro available via API only
How they differentiate
1) Open-source DNA: Fish Speech repo has 31K+ GitHub stars, building developer trust and community contributions; 2) Fine-grained emotional control: 15,000+ natural language tags for word-level emotion/prosody control (unique in the market); 3) Cost efficiency: 11x cheaper than ElevenLabs while winning 67% of blind listening tests; 4) Dual-AR architecture (4B param semantic + 400M param acoustic) enabling LLM-native inference optimizations; 5) 83 languages from a single model; 6) Free S2.1 Pro API tier for developers
Main competitors
ElevenLabs (well-funded competitor, raised $500M at $11B valuation), Cartesia (Sonic 3.5 model, 40ms latency, leads controlled voice arena), WellSaid (enterprise TTS), Speechify, Async (formerly Podcastle), Krisp, Murf AI, PlayHT
Key partnerships
HeyGen (AI avatars using Fish Audio voices), Sanas (enterprise voice AI), Retell AI (voice agent integrations), LiveKit (real-time voice infrastructure), Telnyx (TTS provider integration), OpenArt (partner/customer)
Notable customers
HeyGen, Sanas, Retell AI, LiveKit, OpenArt, Telnyx
Major milestones
2023: Fish Speech open-source project started by Shijia Liao as a bedroom project on a single GPU, 2025: Fish Audio (Hanabi AI Inc.) founded / commercial company launch, Fish Audio S1 model released, 2025: Reached $5M ARR with 20K active developers, 2025: Reached $10M ARR, 2026: Launched S2 open-source model (31K+ GitHub stars), 2026: Launched S2.1 Pro flagship model, 2026: Scaled team to 22 people; 5 models shipped in 1 year, 2026: Reached 8M+ users, $21M ARR, July 28, 2026: Raised $52M seed round co-led by Coreline Ventures and Capital Today
Growth metrics
8M+ users; $21M ARR; 31K+ GitHub stars; 2M+ voices in library; 5 models launched in 1 year; team of 22; ~66-67% blind test preference vs leading competitors
Market positioning
Formidable challenger to ElevenLabs in the AI voice generation market, differentiated by open-source community trust, superior cost efficiency (11x cheaper), and fine-grained emotional controls. Positioned as the developer-friendly, cost-effective alternative for both creators and enterprises, with 8M+ users and $21M ARR. Competes on expressiveness, multilingual support (83 languages), and price-performance ratio.
Geographic focus
Global (headquartered in Palo Alto, CA with strong presence in US, Japan, and Asia-Pacific markets; Tokyo taxi advertising campaign indicates Japan market focus)
Patents and IP
Fish Audio name and logos are trademarks of Hanabi AI Inc.; Fish Speech model licensed under Fish Audio Research License (research & non-commercial use permitted free of charge); no publicly disclosed patents found
About Rissa Cao
Ex-Meta (Meta Reality Lab Product), Ex-Amazon (Senior Product Manager-Technical), Ex-Petuum Inc., Co-founder of PixAI/Mewtant; Columbia University graduate
Latest news about Fish Audio
More Voice / Speech AI companies
Official website: https://fish.audio