Audio, video, and link transcription

MP3 to SRT. Turn audio into timed subtitles.

Upload a recording, transcribe the speech, review text and timing, and download SRT subtitles.

Multi-language source detectionWords, speakers, SRT, full textSource media normally deletes after about 3 hours
podcast-ep-12.mp3TranscribingAuto language: EN+ZH
00:00:01,000 --> 00:00:04,120Speaker 1

Welcome back — today we're talking about subtitles.

00:00:04,480 --> 00:00:07,900Speaker 2

Every word lands with its own timestamp.

00:00:08,260 --> 00:00:11,540Speaker 1

Export SRT the moment the last cue settles.

Words 2,184Speakers 2Cues 46SRT ready
Free daily allowance: up to 5 audio minutes; video uses credits faster. Sign in to keep the result in your workspace.

How does MP3 become SRT?

MP3 stores sound; SRT stores text with start and end times. Renaming the extension cannot convert between them. Speech recognition first transcribes the audio, then aligns the words with the recording.

From upload to download

Choose an audio file above and submit your MP3. When processing finishes, replay important passages, check names, technical terms and numbers, then review cue timing and download SRT. Import the subtitle file alongside your media in an editor that supports SRT.

How to get more accurate subtitles

Use a clear voice recording with limited background music. For meetings, check overlapping speech and speaker changes carefully. Compression artifacts, accents and specialized vocabulary can affect recognition; automatically generated subtitles still benefit from review.

Transcription versus format conversion

MP3 to SRT requires speech recognition and uses transcription allowances and credits. Check pricing and the submission screen for applicable limits. Converting an existing SRT file to VTT does not require transcribing the audio again.

Keep your original audio

SRT is a separate text file and does not contain sound. Keep your original MP3 for playback, editing and future review after downloading your subtitles.

Free daily run

Free daily allowance: up to 5 audio minutes; video and link minutes use more credits.

Short media retention

Uploaded source media normally expires after about 3 hours unless a feature offers longer retention.

Secure checkout

The payment provider identified at checkout handles card payment for global users.

7 source languages

Auto detection, or pin English, Chinese, Japanese, Korean, Spanish, or French.

How it works

00:01

Drop media or paste a link

Audio, video, or a public video URL — the workbench accepts all three.

00:02

We route by language and media shape

Task shape, language, and duration decide the path. You never compare vendors.

00:03

Export every layer you need

SRT, word timestamps, speaker turns, and clean full text — separately or as one ZIP.

One job, four deliverables.

Every completed transcription exposes the same four layers, ready to copy or download.

12
00:00:48,120 --> 00:00:51,400
Every cue lands on time.

Sentence timestamps

SRT subtitles built from complete utterance timestamps.

{"w":"every","s":0.42,"e":0.61}
{"w":"word","s":0.63,"e":0.85}
{"w":"字","s":0.88,…}

Word timestamps

One JSON line per word or character, with start/end time.

S1 00:00:04 → 00:00:11
  "Welcome back to the show."
S2 00:00:11 → 00:00:19

Speaker segments

Sentence-level speaker turns with text and word ranges.

Welcome back to the show. Today,
we're talking about subtitles — and
why timing is everything.

Punctuated full text

The punctuated full transcript.

Pay for minutes, not seats.

1 credit equals 1 audio minute. The USD $20 starter pack covers about 20 audio hours.

$20

1,200 credits

See full pricing

Frequently asked

How accurate are the transcripts?

Media is routed by its language and shape. Available word timestamps, punctuation, and speaker turns are normalized into one reviewable transcript.

What can I do for free?

Free daily allowance: up to 5 audio minutes; video uses credits faster. Sign in to keep the result in your workspace.

What happens to my uploads?

Source media normally deletes after about 3 hours unless a feature offers longer retention. Transcripts stay in your workspace.

Which languages are supported?

Auto detection handles mixed-language media. You can also pin English, Chinese, Japanese, Korean, Spanish, or French as the source.

Can my scripts or agents use it?

Yes. Create an API key billed at $1 per audio hour, or hand your AI agent the public skill file — it inherits your account permissions.