ACTIVE · VERIFIED AGAINST THE PRODUCT

Text to Dialogue

Build multi-speaker audio from written dialogue. Assign a different voice to each speaker and control pauses between lines.

Verified operating specifications

Real inputs, limits and controls

Speakers
2–26 speakers labeled A> through Z>
Timing
Per-speaker speed 0.5–1.5; segment delay 0–5 seconds
Voice mix
Different connected voice source for each character
Extras
Pronunciation dictionary, AI expression tags and optional transcript

Core capabilities

What this tool does

Multiple speakers
Voice per character
Pause controls
Unified voice library

Connected services

Related voice sources, services or models

ElevenLabsConnected service
MiniMaxVoice source
VbeeVoice source
Fish AudioVoice source
Microsoft Edge TTSConnected service

Product workflow

From input to finished output in three steps

01

Format the script

Write lines with A>, B>, C> labels so each segment maps to a character.

02

Cast each speaker

Assign a voice and speed per character, then set the delay between segments.

03

Generate one conversation

Create the mixed dialogue and optionally export its timed transcript.

Use cases

Good fits

  • Podcast conversations
  • Training role-play
  • Character scenes
  • Multi-speaker localization

Developer API

Integrate this workflow

Generation tasks are asynchronous: receive a task_id, then poll or use a webhook.

POST/v3/text-to-speech/dialogue
Read the public API guide
Product-verified information

This page only describes capabilities visible in the current application with a corresponding processing path at the latest verification.