Fish Audio secures $52 million to enhance AI voice technology
technology
innovative
controversial

Fish Audio secures $52 million to enhance AI voice technology

11
(Update: )
American artificial intelligence research organization
  • Fish Audio has over 8 million users and generates $21 million in annual recurring revenue.
  • The startup raised $52 million in a seed round to enhance its AI voice models.
  • The funding will help Fish Audio expand its offerings and compete in a crowded market.
Share opinion
1

Story

In the United States, specifically Palo Alto, Fish Audio has made significant strides in the AI-generated voice model market since its launch last year. The startup has attracted over 8 million users to its open-source and hosted voice models, generating an impressive annual recurring revenue of $21 million. To further its growth and development, Fish Audio announced a successful seed funding round, raising $52 million, which was led by Coreline Ventures and Capital Today, with additional participation from several other investors. This funding will enable the company to expand its library of over 15,000 natural language controls and enhance its offerings for both creators and enterprises. The demand for AI voice models is rapidly increasing, driven by diverse use cases ranging from creative applications to enterprise needs for customer support automation. Fish Audio aims to cater to these varied requirements by providing a range of voice models that can be expressive for creative projects and steerable for business applications. The company has also open-sourced three of its speech generation models, while its latest S2.1 Pro model is available exclusively through a paid API. This dual approach allows Fish Audio to serve both individual creators and larger organizations effectively. However, the company has faced challenges, particularly regarding the consent of voice submissions from creators. Some users alleged that their voices were uploaded without permission, prompting Fish Audio to implement an automated take-down process. This new system allows creators to quickly remove their voices from the platform if they can provide a short sample or a contract proving ownership. This move is crucial for maintaining trust within the community and ensuring that creators feel secure when contributing their voices for model training. As the speech generation market becomes increasingly competitive, with notable players like ElevenLabs and WellSaid vying for market share, Fish Audio's innovative approach and community-driven model may provide a significant advantage. The company’s ability to develop state-of-the-art models with a relatively small team, compared to larger, well-funded AI labs, has garnered attention and praise from industry experts. With the recent funding, Fish Audio is well-positioned to continue its growth trajectory and meet the evolving demands of both creators and enterprises in the AI voice technology landscape.