ACTIVE · VERIFIED AGAINST THE PRODUCT

Text to Speech

Turn scripts and long-form text into natural speech. Process up to 1,000,000 characters per task with voice sources from ElevenLabs, MiniMax, Vbee, Fish Audio or low-cost Microsoft Edge voices.

Focused TTS guides

Explore quality, measured production speed, long-form scale and cost

Dedicated pages explain voice sources, a recent BullMQ production benchmark, long-form limits, low-cost calculations and dated pricing.

Verified operating specifications

Real inputs, limits and controls

Input
Text, SRT, TXT, ZIP or folder batch
Task limit
1,000,000 characters; SRT timeline up to 5 hours
Voice controls
Speed 0.5–1.5, pronunciation rules and AI expression tags
Quality tiers
High quality or Low cost (50% fewer credits in the current UI)
Low-cost voice library
318 Microsoft Edge voices across 140 locale codes

Core capabilities

What this tool does

Up to 1,000,000 characters per task
SRT timelines up to five hours
High-quality and low-cost tiers
Pronunciation and transcript controls

Voice / reference-source labels

Voice-source choices across High quality and Low cost modes

ElevenLabsVoice / reference-source label
MiniMaxVoice / reference-source label
VbeeVoice / reference-source label
Fish AudioVoice / reference-source label
Microsoft Edge TTSLow-cost voice library

Core High quality TTS uses a separate OpenSpeaker synthesis bridge; these labels do not guarantee the named provider’s model or API.

Product workflow

From input to finished output in three steps

01

Add the script

Paste text or SRT, or import TXT, SRT, ZIP files and folders for sequential batch work.

02

Choose the voice

Select a source, speed, quality tier, pronunciation dictionary and optional expression controls.

03

Generate and export

Run the queued task, review the audio and download the result with an optional transcript.

Use cases

Good fits

  • Audiobooks and long narration
  • YouTube and course voiceovers
  • SRT-timed localization
  • High-volume batch production

Developer API

Integrate this workflow

Generation tasks are asynchronous: receive a task_id, then poll or use a webhook.

POST/v3/text-to-speech
Read the public API guide
Product-verified information

This page only describes capabilities visible in the current application with a corresponding processing path at the latest verification.