Suno has launched a public beta of Speech, a feature that generates spoken-voice audio tracks with optional AI-generated background music, The Verge reported on October 2.
The launch expands Suno beyond its core AI music service into the competitive AI-generated speech market, where platforms such as ElevenLabs, Adobe, and DeepMind have been active for years, The Verge noted.
Suno’s chief product officer, Jack Brody described Speech as “the first audio model that generates voice and music together as one cohesive track,” The Verge quoted.
Users can turn off the background music with a toggle to obtain clean speech, according to the report.
The feature offers Simple and Advanced modes. Simple mode lets users describe a scenario, such as “a pirate captain rallying his crew.” Advanced mode accepts a custom script and provides controls for gender, speech style, and generation variety. Each track can run up to around eight minutes.
The company cautioned that the beta is unfinished. “British accents can wander off to Australia and back. Dramatic pauses may be very dramatic,” Brody said, according to The Verge. Suno added that it will refine Speech based on user feedback.
The announcement did not include pricing or details on whether independent performance tests have been conducted.