Skip to main content

Fish Audio

Category: Voice / Speech AI

AI voice generation platform offering expressive text-to-speech, voice cloning, and speech-to-text with fine-grained emotional controls, challenging ElevenLabs with open-source backing and enterprise features. Fish Audio was founded in 2025. The company is led by Rissa Cao. Based in Palo Alto, California, United States. Team size: 11-50. Total funding raised: $52M. Latest round: Seed. Key investors include Coreline Ventures, Capital Today, 359 Capital, Parable, Play Time, Alphaist Partners, Bayhouse Ventures, Carya Venture Partners, HF0, 645 Ventures.

Founded
2025
Headquarters
Palo Alto, California, United States
Team size
11-50
Total funding
$52M

Value proposition

Most expressive, emotionally controllable real-time voice model with fine-grained word-level emotion controls via 15,000+ natural language tags; open-source community trust with 31K+ GitHub stars; 67% listener preference in blind tests vs competitors; 11x lower cost than ElevenLabs; voice cloning from 5-second audio in 15 seconds; 83 languages supported

Products and solutions

Fish Audio S2.1 Pro (flagship expressive TTS model), Fish Speech (open-source TTS with 31K+ GitHub stars), Voice Cloning (from 5-second audio sample), Speech-to-Text (with multi-speaker diarization & emotion tags), Voice Agent (end-to-end voice agent solution), Text-to-Speech API (ultra-low latency, streaming), 2M+ user-uploaded voice library, 83 language support, 15,000+ natural language emotion controls

Unique value

Combines open-source transparency with enterprise-grade quality — the only voice AI platform offering both a top-ranked open-source model (Fish Speech) and a proprietary flagship (S2.1 Pro) that outperforms closed competitors in blind listening tests at a fraction of the cost

Target customer

Content creators (YouTubers, audiobook narrators, game developers), indie developers, and enterprise customers in regulated industries (healthcare, financial services) needing voice AI for customer support, AI avatars, and conversational agents

Industries served

Content Creation (video voiceovers, audiobooks), Gaming (character voices), Customer Support (conversational chatbots), AI Avatars, Healthcare (HIPAA-compliant deployments), Financial Services

Technology advantage

Dual-Autoregressive (Dual-AR) architecture — 4B-parameter Slow AR for semantic prediction + 400M-parameter Fast AR for acoustic generation, structurally isomorphic to standard LLMs enabling SGLang acceleration (Continuous Batching, Paged KV Cache, CUDA Graph); trained on 10M+ hours of audio across 80+ languages; GRPO reinforcement learning alignment; open-sourced S2 model under Fish Audio Research License; S2.1 Pro available via API only

How they differentiate

1) Open-source DNA: Fish Speech repo has 31K+ GitHub stars, building developer trust and community contributions; 2) Fine-grained emotional control: 15,000+ natural language tags for word-level emotion/prosody control (unique in the market); 3) Cost efficiency: 11x cheaper than ElevenLabs while winning 67% of blind listening tests; 4) Dual-AR architecture (4B param semantic + 400M param acoustic) enabling LLM-native inference optimizations; 5) 83 languages from a single model; 6) Free S2.1 Pro API tier for developers

Main competitors

ElevenLabs (well-funded competitor, raised $500M at $11B valuation), Cartesia (Sonic 3.5 model, 40ms latency, leads controlled voice arena), WellSaid (enterprise TTS), Speechify, Async (formerly Podcastle), Krisp, Murf AI, PlayHT

Key partnerships

HeyGen (AI avatars using Fish Audio voices), Sanas (enterprise voice AI), Retell AI (voice agent integrations), LiveKit (real-time voice infrastructure), Telnyx (TTS provider integration), OpenArt (partner/customer)

Notable customers

HeyGen, Sanas, Retell AI, LiveKit, OpenArt, Telnyx

Major milestones

2023: Fish Speech open-source project started by Shijia Liao as a bedroom project on a single GPU, 2025: Fish Audio (Hanabi AI Inc.) founded / commercial company launch, Fish Audio S1 model released, 2025: Reached $5M ARR with 20K active developers, 2025: Reached $10M ARR, 2026: Launched S2 open-source model (31K+ GitHub stars), 2026: Launched S2.1 Pro flagship model, 2026: Scaled team to 22 people; 5 models shipped in 1 year, 2026: Reached 8M+ users, $21M ARR, July 28, 2026: Raised $52M seed round co-led by Coreline Ventures and Capital Today

Growth metrics

8M+ users; $21M ARR; 31K+ GitHub stars; 2M+ voices in library; 5 models launched in 1 year; team of 22; ~66-67% blind test preference vs leading competitors

Market positioning

Formidable challenger to ElevenLabs in the AI voice generation market, differentiated by open-source community trust, superior cost efficiency (11x cheaper), and fine-grained emotional controls. Positioned as the developer-friendly, cost-effective alternative for both creators and enterprises, with 8M+ users and $21M ARR. Competes on expressiveness, multilingual support (83 languages), and price-performance ratio.

Geographic focus

Global (headquartered in Palo Alto, CA with strong presence in US, Japan, and Asia-Pacific markets; Tokyo taxi advertising campaign indicates Japan market focus)

Patents and IP

Fish Audio name and logos are trademarks of Hanabi AI Inc.; Fish Speech model licensed under Fish Audio Research License (research & non-commercial use permitted free of charge); no publicly disclosed patents found

About Rissa Cao

Ex-Meta (Meta Reality Lab Product), Ex-Amazon (Senior Product Manager-Technical), Ex-Petuum Inc., Co-founder of PixAI/Mewtant; Columbia University graduate

Latest news about Fish Audio

More Voice / Speech AI companies

Official website: