A text-to-speech model released under Apache 2.0, with text as inputs and audio as outputs. It ranks in first place on TTS Arena with over 7,000 votes.