12 00:00:48,120 --> 00:00:51,400 Every cue lands on time.
Sentence timestamps
SRT subtitles built from complete utterance timestamps.
Upload a recording, transcribe the speech, review text and timing, and download SRT subtitles.
Welcome back — today we're talking about subtitles.
Every word lands with its own timestamp.
Export SRT the moment the last cue settles.
MP3 stores sound; SRT stores text with start and end times. Renaming the extension cannot convert between them. Speech recognition first transcribes the audio, then aligns the words with the recording.
Choose an audio file above and submit your MP3. When processing finishes, replay important passages, check names, technical terms and numbers, then review cue timing and download SRT. Import the subtitle file alongside your media in an editor that supports SRT.
Use a clear voice recording with limited background music. For meetings, check overlapping speech and speaker changes carefully. Compression artifacts, accents and specialized vocabulary can affect recognition; automatically generated subtitles still benefit from review.
MP3 to SRT requires speech recognition and uses transcription allowances and credits. Check pricing and the submission screen for applicable limits. Converting an existing SRT file to VTT does not require transcribing the audio again.
SRT is a separate text file and does not contain sound. Keep your original MP3 for playback, editing and future review after downloading your subtitles.
Free daily run
Free daily allowance: up to 5 audio minutes; video and link minutes use more credits.
Short media retention
Uploaded source media normally expires after about 3 hours unless a feature offers longer retention.
Secure checkout
The payment provider identified at checkout handles card payment for global users.
7 source languages
Auto detection, or pin English, Chinese, Japanese, Korean, Spanish, or French.
Audio, video, or a public video URL — the workbench accepts all three.
Task shape, language, and duration decide the path. You never compare vendors.
SRT, word timestamps, speaker turns, and clean full text — separately or as one ZIP.
Every completed transcription exposes the same four layers, ready to copy or download.
12 00:00:48,120 --> 00:00:51,400 Every cue lands on time.
SRT subtitles built from complete utterance timestamps.
{"w":"every","s":0.42,"e":0.61}
{"w":"word","s":0.63,"e":0.85}
{"w":"字","s":0.88,…}One JSON line per word or character, with start/end time.
S1 00:00:04 → 00:00:11 "Welcome back to the show." S2 00:00:11 → 00:00:19
Sentence-level speaker turns with text and word ranges.
Welcome back to the show. Today, we're talking about subtitles — and why timing is everything.
The punctuated full transcript.
1 credit equals 1 audio minute. The USD $20 starter pack covers about 20 audio hours.
$20
1,200 credits
Media is routed by its language and shape. Available word timestamps, punctuation, and speaker turns are normalized into one reviewable transcript.
Free daily allowance: up to 5 audio minutes; video uses credits faster. Sign in to keep the result in your workspace.
Source media normally deletes after about 3 hours unless a feature offers longer retention. Transcripts stay in your workspace.
Auto detection handles mixed-language media. You can also pin English, Chinese, Japanese, Korean, Spanish, or French as the source.
Yes. Create an API key billed at $1 per audio hour, or hand your AI agent the public skill file — it inherits your account permissions.