Suno has expanded its generative audio capabilities beyond music, launching a public beta for a new feature called Speech. The tool allows users to generate spoken voiceovers based on scripts or prompted descriptions, marking a significant diversification for the AI music platform.

What Happened

The Speech feature is now available across Suno’s web and mobile platforms. According to the company, it is the first audio model designed to generate voice and music together as a single cohesive track. Users can choose between two modes: 'Simple,' which generates audio from a descriptive prompt, and 'Advanced,' which accepts custom scripts. In Advanced mode, users can adjust the gender of the AI voice, the speech style, and the level of variety in each generation.

A key differentiator for Suno’s implementation is the ability to simultaneously generate background music to accompany the spoken words. This feature is optional and can be toggled off for clean speech. Suno suggests use cases such as calming soundtracks for poetry or energetic scores for dramatic voiceovers. The maximum duration for a generated speech track is approximately eight minutes.

Why It Matters

This launch positions Suno in the competitive text-to-speech market, joining established players like ElevenLabs, Adobe, and DeepMind. The move appears to be a strategic effort to diversify the platform’s offerings, particularly as its core music generation service has faced multiple lawsuits. By integrating voice and music generation, Suno aims to offer a unified audio creation workflow that distinguishes it from specialized speech synthesis tools.

Jack Brody, Suno’s chief product officer, emphasized that while music remains central to the company’s vision, the expansion into speech represents a broader approach to human expression. The company acknowledges the feature’s limitations, noting that beta-stage imperfections such as accent inconsistencies and exaggerated dramatic pauses may occur. Suno states that it will continue to refine the model based on user feedback.

The Bottom Line

Suno’s entry into AI-generated speech with simultaneous music accompaniment represents a strategic pivot to diversify its product suite amidst legal challenges. The public beta offers a cohesive audio generation tool with a cap of eight minutes per track, though the company admits the technology is still in early stages of development.