ACTIVE · VERIFIED AGAINST THE PRODUCT

Speech to Text

Transcribe audio with ElevenLabs Scribe v2 and export readable text, SRT subtitles or structured timing data.

Verified operating specifications

Real inputs, limits and controls

Upload limit
200 MB per audio file
Formats
AAC, AIFF, OGG, MP3, Opus, WAV, WebM, FLAC and M4A
Outputs
SRT segments plus JSON segment and word timing
Detection
Optional audio-event tagging; multi-file batch and ZIP download

Core capabilities

What this tool does

Scribe v2
SRT export
Structured timing data
Audio upload

Connected services

Related voice sources, services or models

ElevenLabsConnected service

Product workflow

From input to finished output in three steps

01

Upload the recordings

Add one or more files in any of the nine supported formats.

02

Choose transcription detail

Enable audio-event tags when non-speech sounds matter.

03

Export timed text

Download SRT or structured JSON individually or as a batch ZIP.

Use cases

Good fits

  • Subtitle creation
  • Podcast transcripts
  • Interview analysis
  • Word-level timing data

Developer API

Integrate this workflow

Generation tasks are asynchronous: receive a task_id, then poll or use a webhook.

POST/v1/task/speech-to-text
Read the public API guide
Product-verified information

This page only describes capabilities visible in the current application with a corresponding processing path at the latest verification.