Fish Audio's $52M Bet on AI Voices
A voice model that started on a single GPU now has 8 million users and a thorny consent problem to solve.
Synthetic voices have long had a flatness problem. They can read your words back, but they rarely sound like they mean any of it. A former NVIDIA researcher got tired of that, trained a voice model on a single graphics card, and put it online for free. Roughly a year later, that side project has grown into Fish Audio, a Palo Alto startup that just raised a hefty seed round.
What it is
Fish Audio builds AI voice models, meaning software that turns text into spoken audio and can clone or generate a voice on demand. The company was started by Shijia Liao, who was frustrated by the wooden synthetic voices on the market. His open-source project, called Fish Speech, now has more than 31,000 stars on GitHub, the coding site where developers bookmark projects they like. It is used by indie developers, game designers, and creators.
The pitch is flexibility. Fish Audio offers more than 15,000 natural language controls, which are essentially dials for tuning how a voice sounds and behaves. That matters because different customers want different things. A company making AI avatars wants realism. A gaming studio wants expressive characters. A voice agent handling phone calls wants natural, low-latency speech that does not lag mid-sentence.
The traction is real. Fish Audio says more than 8 million people use its open-source or hosted models, and it reports $21 million in annual recurring revenue, the yearly value of its subscriptions. It has shipped five models in a year, four for speech generation and one for speech-to-text, open-sourcing three while keeping its newest, S2.1 Pro, behind a paid API. Customers reportedly include HeyGen and Sanas.
Why it matters
On Tuesday, Fish Audio said it raised $52 million in a seed round led by Coreline Ventures and Capital Today, with a long list of other investors joining. That is a large sum for a seed, the earliest formal funding stage, and it signals how crowded and competitive the AI voice space has become. Fish Audio is up against well-funded rivals including ElevenLabs, Cartesia, Speechify, and Krisp.
Interestingly, CEO and co-founder Rissa Cao said the company did not actually need the money when it was running as an open-source project with creator plans. It raised because it wants to build more advanced models and serve enterprise customers, and because investors were knocking.
The consent question
Here is the part worth watching. Fish Audio built part of its voice library by asking users to submit their own voices for training, and paying them if a voice gets used. That community approach is clever, but it ran into trouble. Some creators alleged their voices were uploaded without consent.
The company had a DMCA takedown process, the legal mechanism for requesting removal of copyrighted material, but Cao acknowledged the takedowns were slow. Fish Audio says it has now automated the system. A creator can submit a short voice sample or a contract to prove ownership, and the voice comes down in under three minutes.
That is faster, but it does not fix the underlying gap. Nothing stops someone from uploading an artist's voice in the first place, and it stays live until the artist notices and files a complaint. Osuke Honda of Coreline Ventures, an investor, was candid about this. A community model, he said, only works if creators trust the platform, which means consent, transparency, and attribution need to be built into the product rather than bolted on later. He also floated verified voice ownership and revenue sharing as where the industry should head.
What's next
Fish Audio plans to release an audio understanding model this year, plus a speech-to-speech model that would convert one voice directly into another. Investors are betting that fine-grained controls and cheap model training let a lean team compete with bigger labs.
The technical gap between robotic and human-sounding voices is closing fast. The harder problem, and the one that will separate durable platforms from cautionary tales, is proving that the voices being sold actually belong to the people selling them. Speed of takedown is a start. Genuine consent by default is the real finish line.