The market for AI-generated voice models is huge. Creative use cases require AI voice models to be more expressive, while companies looking to automate customer support and sales operations need AI voice models to be more usable.
Palo Alto-based Fish Audio hopes to address all of these use cases with a library of over 15,000 natural language controls. Since its founding last year, the startup now has more than 8 million people using open source or hosted versions of its model and generates $21 million in annual recurring revenue.
Adding to its momentum, the startup announced Tuesday that it has raised $52 million in a seed round led by Coreline Ventures and Capital Today. 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0 also participated in the funding.
Fish Audio started as a small project by former Nvidia researcher Shijia Liao. Frustrated with the lack of expressive synthetic speech available on the market, he trained a speech generation model on a single GPU and open sourced it. The Fish Speech repository on GitHub currently has over 31,000 stars and is used by indie developers, video game designers, and creators.
The company launched five models last year: four speech generation models and one speech-to-text model. The company has open sourced three of its voice generation models, but the latest S2.1 Pro model is only available through a paid API.
Fish Audio offers paid monthly plans suitable for creators and teams that unlock limited time generation and audio cloning features. The company also offers an enterprise version of its API and platform, which organizations such as HeyGen and Sanas are already using.
“Every company has different use cases and different preferences. For example, companies like HeyGen, which uses our voices to power their AI avatars, want realism in their voices. Game studios want expressive voices for their characters, and voice agent companies like LiveKit want more natural-sounding, low-latency voices with enough expressiveness for their calls,” said Rissa Cao, CEO and Co-Founder of Fish Audio. says Mr.
One of the ways the company built its audio library is by asking users to submit their own audio to train the model, and compensating them if their audio is used. However, this caused some trouble a few months ago when some creators claimed that their voices were uploaded to Fish Audio without their consent. The startup had introduced a DMCA removal process to address such concerns, but the removal itself took a long time.
Cao told TechCrunch that the company has automated the removal process. Creators can submit a short audio sample or contract to prove that the uploaded audio is theirs, and the audio will be removed from the startup’s platform within three minutes, she said.
Still, nothing can prevent an artist’s voice from being uploaded without their knowledge. And until artists find out, their voices will continue to be used on the platform unless they request removal.
Ouke Honda, partner at Coreline Ventures, said a community-driven model only works if creators trust the platform.
“A community-centric approach will only be a lasting benefit if creators trust the platform, which means consent, transparency, and attribution need to be built into the product, rather than treated as an afterthought. I think the industry needs to move to verified audio ownership, clear licensing terms, easy reporting and removal processes, and ultimately revenue-sharing models where creators benefit financially when their audio is licensed or used commercially,” he said.
Cao said that when the startup was simply offering its product as an open source project with plans for creators, it operated efficiently and didn’t need outside capital. However, the company wanted to develop more advanced models, wanted to cater to businesses as investor interest was growing, and was seeking capital.
Looking ahead, Fish Audio plans to release a speech understanding model this year. We are also building a speech recognition model.
The audio generation market is crowded, with companies like ElementalLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp competing for creator and enterprise budgets.
According to Rico Mallozzi, partner at 359 Capital, granular control for developers and cost-effective model training will allow Fish Audio to be more competitive with large AI labs.
“Compared to other well-funded AI labs and AI companies, I think it’s incredible that they were able to build a cutting-edge model with the team they have. This shows their technical acumen in bridging the gap between artificial-sounding voices and human-like voices,” Mallozzi told TechCrunch by phone.
This article has been updated to reflect that the company has raised a $52 million seed round.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
