Co-founder and Chief Scientist Shijia Liao, a former NVIDIA researcher, launched the company after growing frustrated with the robotic quality of existing synthetic voices. By training models on a single gaming GPU, he created the open-source project Fish Speech, which garnered over 31,000 stars on GitHub. That initial momentum fueled the platform's current capability to clone voices from five-second clips in roughly 15 seconds, supporting over 83 languages with nuanced emotional control.
CEO Rissa Cao noted that the company’s growth stems from a commitment to human-sounding output that remains accessible to everyone from independent creators to large-scale enterprises. The startup's latest model, S2.1 Pro, outperformed leading competitors in 67% of blind listening tests, helping the firm secure partnerships with companies like HeyGen and Retell.

Comments (0)
No comments yet. Be the first!