Production benchmark · Single-speaker + dialogue · 2 August 2026

153K characters to 2h 25m of high-quality speech—in 2m 38s.

This was one real single-speaker OpenSpeaker production job using an ElevenLabs voice source, not a lab estimate: 153,072 characters became 2h 25m 46s of audio in 2m 38s, with transcript generation off and RTF 0.0181—55.2× faster than playback. Two 100K+ multi-speaker dialogue jobs reached 56.1× and 58.9× real time. Long-form is not an edge case here; it is what OpenSpeaker is built to do.

153Ksingle-speaker inputElevenLabs voice source
2h 25mfinished audioMP3 · 44.1 kHz · stereo · 128 kbps
2m 38sactive processingFull production workflow
55.2×faster than playbackRTF 0.0181 · with_transcript=false
OpenSpeaker long-form TTS benchmark: a 153K-character single-speaker job generated 2h 25m of audio in 2m 38s at RTF 0.0181, alongside two 100K+ dialogue jobs above 55 times real time
Shareable production benchmark covering one 153K single-speaker job and two 100K+ dialogue jobs. It publishes aggregate measurements only—no user content or identifiers.

Short answer

OpenSpeaker is a strong fit for audiobooks, podcasts, courses and long video scripts that need both serious throughput and a voice people will want to keep listening to. Submit up to 1,000,000 characters as one task, choose a High quality voice or a multi-speaker cast, and let OpenSpeaker handle segmentation, synthesis, merging and delivery. The production evidence here shows three 100K+ jobs generating audio 55.2–58.9× faster than playback, with transcript generation disabled.

100K+ production proof

Three 100K+ production jobs. Every one generated audio at more than 55× real time.

The audit found one 153K single-speaker narration job and two 100K+ multi-speaker dialogue jobs. All three had transcript generation explicitly disabled, complete timing fields and final audio that could be measured.

Dialogue A · multi-speakerwith_transcript=false
101,481 characters
Final audio
96.64 min
Active processing
103.321s
RTF
0.0178
Playback multiple
56.12×
Dialogue B · multi-speakerwith_transcript=false
103,074 characters
Final audio
102.56 min
Active processing
104.451s
RTF
0.017
Playback multiple
58.91×

This is full workflow speed—not a cherry-picked model-latency number. Active processing includes orchestration, synthesis, merge and post-processing, upload and durable task completion. RTF divides that time by the final MP3 duration measured with ffprobe.

More than one standout job

45 long-form jobs. 26.4 hours of audio. Every one faster than real time.

The clean cohort contains 45 single-speaker jobs from 5,263 to 153,072 characters. Every job explicitly recorded with_transcript=false, had valid worker timing and yielded a measurable final MP3; ffprobe failed on none of them.

45/45faster than real timeEvery qualifying job
0.0346median RTFLower is faster
28.9×median speedVersus playback length
26.4hfinal audio measured1,385,427 total characters
Active processingfinishedOn − processedOn
RTFactive seconds ÷ final audio seconds
Audio durationffprobe(tasks.output_uri)

Speed in listening terms

More than two hours of audio, generated while you make coffee

The 153,072-character single-speaker job produced 8,746.475 seconds of finished audio—2 hours, 25 minutes and 46 seconds—in 158.396 seconds of active processing. That is RTF 0.0181: roughly 55 minutes of listening generated for every minute of processing.

Character counts vary by language and writing style. RTF compares processing time with the final audio duration, so it tells the long-form story more honestly. Across all three 100K+ samples, RTF ranged from 0.0170 to 0.0181, equivalent to 55.2–58.9× real time.

  • 153,072 characters → 2h 25m 46s audio
  • 2m 38s active processing
  • RTF 0.0181 · 55.2× real time
  • Two 100K+ dialogues · 56.1–58.9× real time

Long-form without the busywork

One manuscript in. One finished production task out.

Bring a chapter, podcast episode, course or full script—OpenSpeaker accepts up to 1,000,000 characters in one task. Behind the scenes it creates engine-sized segments, processes batches with engine-specific concurrency, restores the original order and delivers the merged audio.

Use a single narrator for an audiobook or explainer, or build a multi-speaker dialogue for a podcast, audio drama or localized conversation. You follow one job instead of babysitting dozens of provider-sized requests.

  • Up to 1,000,000 characters per task
  • Single-speaker narration or multi-speaker dialogue
  • Automatic segmentation and ordered merge
  • TXT, SRT, ZIP, folder and API workflows

Made to be listened to

Fast generation only matters when the voice deserves the runtime

OpenSpeaker’s High quality mode puts premium voice choices from ElevenLabs, MiniMax, Fish Audio and Vbee in the same long-form workflow. Choose a warm narrator, a crisp explainer or distinct character voices without rebuilding the production stack around separate tools.

RTF measures speed, not perceptual quality. OpenSpeaker pairs that speed with the controls that matter for long listening: voice selection, language filters, pacing and pronunciation dictionaries. Test one representative passage, lock the voice and recurring terms, then scale with confidence.

Why OpenSpeaker

Premium voice choice, serious long-form capacity and a low-cost way in

One workspace combines high-quality voice options, a separate budget tier, million-character tasks, pronunciation control, batch imports and API access. Packages start at $5 for 1,000,000 premium credits, so creators can test a real chapter before committing a full production budget.

OpenSpeaker does not force every project through one model. Pick the voice and cost profile that fit the job, then keep the entire long-form workflow in one place. This benchmark proves OpenSpeaker’s observed production speed; it is not a same-script race against every provider.

Measured, not estimated

Real production jobs. Final audio duration. Transparent RTF.

RTF is active processing time divided by final MP3 duration; lower is faster. Active processing covers orchestration, synthesis, merge and post-processing, upload and durable task completion. Transcript generation was explicitly false for all three 100K+ jobs.

The samples come from completed jobs retained by the live production queues, so they are evidence—not a universal guarantee. Active processing uses one worker clock; recorded end-to-end can inherit sub-second cross-host skew. The benchmark measures speed, not MOS or a head-to-head competitor result, and publishes no text, user identifiers, task IDs or output URLs.

Private data stays private

The article and JSON endpoint publish aggregates only. Input content, user IDs, task IDs, output URLs and credentials were not included in the aggregate or published.

Frequently asked questions

TTS speed, audiobooks and benchmark scope

What is the best fast long-form TTS for audiobooks?

For creators who want premium voice choice, million-character tasks, a low-cost way to start and speed measured against the finished audio, OpenSpeaker deserves a place at the top of the shortlist. The dated production proof here includes a 153K-character single-speaker run at 55.2× real time and a 45-job long-form cohort with a 28.9× median speedup. Audition the exact voice and language you need before committing a full book.

Can OpenSpeaker handle a 100K+ character single-speaker audiobook job?

Yes. In this production snapshot, one single-speaker job turned 153,072 characters into 2h 25m 46s of finished audio in 158.396 seconds of active processing. Its RTF was 0.0181, or 55.2× faster than playback. This is an observed result, not a guaranteed completion time.

What is RTF, and what did OpenSpeaker achieve?

Real-time factor divides active processing time by final audio duration; lower is faster. The three 100K+ production samples measured 0.0170–0.0181 RTF, equivalent to generating audio 55.2–58.9× faster than real time.

Does this speed benchmark prove that the voices sound high quality?

The benchmark measures speed, not perceptual quality or MOS. OpenSpeaker offers a High quality mode with premium voice choices from ElevenLabs, MiniMax, Fish Audio and Vbee. Audition a representative passage in the voice and language you plan to use; then lock pacing and pronunciation before generating the full project.

Did the measured jobs include transcript generation?

No. Transcript generation was explicitly disabled for the 153K single-speaker job and both 100K+ dialogue jobs, so their RTF reflects the speech-production workflow without a separate transcription step.

Is OpenSpeaker an ElevenLabs alternative for long-form TTS?

Yes. OpenSpeaker is an independent, lower-cost long-form workspace with high-quality voice choices from ElevenLabs, MiniMax, Fish Audio and Vbee, plus a separate budget tier and tasks up to 1,000,000 characters. This dataset is not a same-script speed comparison with ElevenLabs.

Were manuscript text, audio files or user details published?

No. The public benchmark contains aggregate measurements only. It does not publish input text, prompts, user IDs, task IDs, credentials, output URLs or private audio.