The market for AI-generated speech models is huge. Creative use cases require AI voice models to be more expressive, while companies looking to automate customer service and sales operations need them to be more manageable. Palo Alto-based Fish Audio wants to address all of those use cases with its library of more than 15,000 natural
The market for AI-generated speech models is huge. Creative use cases require AI voice models to be more expressive, while companies looking to automate customer service and sales operations need them to be more manageable.
Palo Alto-based Fish Audio wants to address all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open source or hosted versions of its models, and now generates annual recurring revenue of $21 million.
To continue building on that traction, the startup said Tuesday it had raised $50 million in a seed round led by Coreline Ventures and Capital Today. The financing also included participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners and HF0.
Fish Audio began as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by the non-expressive synthetic voices available on the market, trained a speech generation model on a single GPU, which was open sourced. The Fish Speech repository on GitHub now has over 31,000 stars and is used by independent game developers, designers, and creators.
The company launched five models last year: four speech-generating models and one speech-to-text model. It has three of its speech generation models open source, but its latest S2.1 Pro model is only available through its paid API.
Fish Audio offers paid monthly plans suitable for creators and teams that unlock a set number of generation minutes, as well as voice cloning features. The company also offers an enterprise version of its API and platform, and says organizations like HeyGen, Sanas, and Plaud are already using it.
“Every company has different use cases and different preferences. For example, companies like HeyGen, which uses our voices to power AI avatars, want realism in the voices; a game studio would want expressive voices for its characters; and voice agent companies like LiveKit want more natural-sounding, low-latency voices that are expressive enough for calls,” Cao said.
One way the startup has built its voice library is by simply asking users to submit their own voices to train its models and compensating them if their voices are used. However, that resulted in some problems a few months ago, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA takedown process to address such concerns, but the takedown itself took a long time.
Fish Audio CEO and co-founder Rissa Cao told TechCrunch that the company has now automated the removal process. Creators can easily submit a short voice sample or contract to prove that the uploaded voice belongs to them, and their voice will be removed from the startup’s platform in less than 3 minutes, it said.
Still, that doesn’t stop someone from uploading an artist’s voice without their knowledge. And until the artist finds out, their voice will continue to be used on the platform until they request its removal.
Oskue Honda, partner at Coreline Ventures, said a community-driven model only works when creators trust the platform.
“A community-focused approach can only become a lasting advantage if creators trust the platform. That means consent, transparency and attribution need to be built into the product rather than treated as an afterthought. I think the industry needs to move towards verified voice ownership, clear licensing terms, simple reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially,” he said.
Cao said that when the startup only offered its product as an open source project with plans for creators, it ran efficiently and didn’t need money. But it wanted to develop more advanced models and also wanted to adapt to companies as investor interest increased, which led it to seek capital.
Looking ahead, Fish Audio plans to release an audio understanding model this year. It is also building a voice-to-speech model.
The speech generation market is crowded, with companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp competing for the wallets of creators and businesses.
According to Rico Mallozzi, partner at 359 Capital, detailed controls for developers and cost-effective model training will help Fish Audio better compete with large AI labs.
“I think what they’ve been able to build, state-of-the-art models, with the equipment they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in bridging the gap between artificial sound and human voices,” Mallozzi told TechCrunch during a call.
When you buy through links in our articles, we may earn a small commission. This does not affect our editorial independence.
For more tech updates, stay tuned to our blog.

















