Twenty-one million in recurring revenue in twelve months, with no outside capital. And one limitation money cannot fix.
In brief
- Fish Audio, legally Hanabi AI Inc., announced a $52 million seed round on 28 July 2026, co-led by Coreline Ventures and Capital Today.
- In its first year the company reached $21 million in annual recurring revenue and more than 8 million users — having raised no outside capital before this round.
- Its business model inverts the industry norm: model weights are released openly, revenue comes from APIs and enterprise contracts. That is also its weak point, and the company acknowledges it.
A seed round the size of a Series B
A $52 million seed is an unusual number: it is Series B territory, not a first institutional cheque. The explanation sits in the traction behind it. Fish Audio, headquartered in Palo Alto and legally registered as Hanabi AI Inc., announced the round on 28 July 2026, on its first anniversary — twelve months in which it went from zero to $21 million in annual recurring revenue and more than 8 million users across creators, developers and enterprises, without outside capital.
Coreline Ventures and Capital Today co-led, joined by a syndicate including 359 Capital, Parable, Play Time, HF0, 645 Ventures, Carya Venture Partners, Bayhouse Ventures and Alphalist Partners, plus undisclosed angels. No valuation was announced.
| Company card | |
|---|---|
| Company | Fish Audio (Hanabi AI Inc.) |
| Headquarters | Palo Alto, California |
| Round | $52M seed, announced 28 July 2026 |
| Lead investors | Coreline Ventures, Capital Today |
| Reported ARR | $21 million |
| Users | over 8 million |
| Product | text-to-speech, voice cloning, voice agents |
| Named customers | OpenAI, HeyGen, Retell, LiveKit, Telnyx, Sanas |
| Valuation | not disclosed |
From one GPU in a bedroom to eight million users
The origin story is the kind investors like to tell, and in this case it is documented. Shijia Liao, co-founder and chief scientist, was a video researcher at Nvidia before Fish Audio. A lifelong fan of Japanese animation, he was frustrated by how flat available synthetic voices sounded: he trained a voice generation model on a single GPU and open-sourced it. That repository — Fish Speech — now has more than 31,000 GitHub stars and is used by independent developers, game studios and creators. The company is led by CEO Rissa Cao.
Over twelve months Fish Audio has released five models: four for speech generation and one for speech recognition. Three of the generation models were open-sourced; the most recent, S2.1 Pro, is available only through the paid API. In blind listening tests run by the company, S2.1 Pro was preferred by 67% of listeners over competing models, with support for a stated 83 languages.
The business model: give away the weights, sell the infrastructure
The strategy is what the industry calls open-weights: base model weights are freely downloadable, and the repository doubles as a developer acquisition channel. Revenue comes not from the model but from everything around it — high-throughput API access, hosted infrastructure, enterprise contracts.
It is a two-tier structure that sharply lowers customer acquisition cost — the community is the funnel — and shifts competition from the model to the service. The customer list shows where the value sits: AI avatar platforms such as HeyGen, voice agent infrastructure such as LiveKit and Retell, telecom services such as Telnyx, alongside Sanas and OpenAI. Over time the company has moved from its initial target of game developers and creators toward regulated industries like healthcare and financial services, offering on-premises deployments with zero data retention and compliance with US health privacy standards.
The limitation $52 million cannot fix
Here the analysis parts ways with the press release. The open-weights model carries a structural downside the company itself acknowledges: once downloaded, the model runs entirely locally, on the hardware of whoever took it. Any consent verification or takedown mechanism implemented on the platform does not reach copies already distributed.
That distinction matters in a market where the principal risk is not technical but legal and reputational: cloning someone’s voice without permission. On the hosted service the company can process removal requests and consent checks ; on downloaded weights it can do nothing. Capital accelerates model development, sales hiring and commercial expansion — it does not close that door, which is a consequence of the architectural choice, not an implementation gap.
For anyone evaluating these tools inside a company, that is the first question to ask: do the governance guarantees cover the cloud service, or also the model anyone can download?
The competitive context: David, Goliath and eleven billion
The market Fish Audio is entering with institutional capital already has a dominant player. In February 2026 ElevenLabs closed a $500 million Series D at an $11 billion valuation, led by Sequoia Capital. Cartesia and WellSaid compete on the same ground.
The difference is not scale, it is architecture. The main competitors run proprietary, walled ecosystems; Fish Audio offloads part of the optimisation and adoption work to a global developer community. It is a bet: that in synthetic voice, as in other layers of software infrastructure, the open model gains distribution faster than the closed model gains margin.
The new capital targets three stated fronts: voice-native large language models and direct speech-to-speech translation, deeper integrations with ecosystem partners, and building an enterprise sales team. It is the classic transition from spontaneous developer adoption to multi-year contracts — the point where many open-source companies have stalled.
Frequently asked questions
Who is Fish Audio? A voice AI startup based in Palo Alto, legally Hanabi AI Inc., founded by former Nvidia researcher Shijia Liao and led by CEO Rissa Cao.
How much did it raise, and from whom? $52 million in a seed round announced on 28 July 2026, co-led by Coreline Ventures and Capital Today with a syndicate of other funds and angels. No valuation was disclosed.
What does “open-weights” mean? That the model weights — the parameters learned during training — are downloadable and freely usable. The company monetises the services around the model, not the model itself.
Is Fish Audio profitable? Not stated. The disclosed figure is different and should be reported precisely: $21 million in annual recurring revenue reached in its first year without outside capital.
Who are its competitors? Primarily ElevenLabs, which raised $500 million at an $11 billion valuation in February 2026, alongside Cartesia and WellSaid.
Sources
- Fish Audio official press release, PRNewswire, 28 July 2026
- TechCrunch — “Fish Audio raises $52M seed to build AI voice models for creators and enterprises”
- SiliconANGLE — “Fish Audio makes a splash after raising $52M seed funding for AI voices”
- TechTimes — analysis of the structural limitation of open-weight models
- MLQ News, AI Weekly — ARR, user and enterprise customer data



