Audio, video, and link transcription

Audio to SRT. Subtitles ready to review.

Upload audio or video, or provide a public video link you are authorized to process. Get timestamped transcripts, speaker segments, SRT, and clean text without choosing between transcription engines.

Multi-language source detectionWords, speakers, SRT, full textSource media normally deletes after about 3 hours
podcast-ep-12.mp3TranscribingAuto language: EN+ZH
00:00:01,000 --> 00:00:04,120Speaker 1

Welcome back — today we're talking about subtitles.

00:00:04,480 --> 00:00:07,900Speaker 2

Every word lands with its own timestamp.

00:00:08,260 --> 00:00:11,540Speaker 1

Export SRT the moment the last cue settles.

Words 2,184Speakers 2Cues 46SRT ready
Free daily allowance: up to 5 audio minutes; video uses credits faster. Sign in to keep the result in your workspace.

Free daily run

Free daily allowance: up to 5 audio minutes; video and link minutes use more credits.

Short media retention

Uploaded source media normally expires after about 3 hours unless a feature offers longer retention.

Secure checkout

The payment provider identified at checkout handles card payment for global users.

7 source languages

Auto detection, or pin English, Chinese, Japanese, Korean, Spanish, or French.

How it works

00:01

Drop media or paste a link

Audio, video, or a public video URL — the workbench accepts all three.

00:02

We route by language and media shape

Task shape, language, and duration decide the path. You never compare vendors.

00:03

Export every layer you need

SRT, word timestamps, speaker turns, and clean full text — separately or as one ZIP.

One job, four deliverables.

Every completed transcription exposes the same four layers, ready to copy or download.

12
00:00:48,120 --> 00:00:51,400
Every cue lands on time.

Sentence timestamps

SRT subtitles built from complete utterance timestamps.

{"w":"every","s":0.42,"e":0.61}
{"w":"word","s":0.63,"e":0.85}
{"w":"字","s":0.88,…}

Word timestamps

One JSON line per word or character, with start/end time.

S1 00:00:04 → 00:00:11
  "Welcome back to the show."
S2 00:00:11 → 00:00:19

Speaker segments

Sentence-level speaker turns with text and word ranges.

Welcome back to the show. Today,
we're talking about subtitles — and
why timing is everything.

Punctuated full text

The punctuated full transcript.

Pay for minutes, not seats.

1 credit equals 1 audio minute. The USD $20 starter pack covers about 20 audio hours.

$20

1,200 credits

See full pricing

Frequently asked

How accurate are the transcripts?

Media is routed by its language and shape. Available word timestamps, punctuation, and speaker turns are normalized into one reviewable transcript.

What can I do for free?

Free daily allowance: up to 5 audio minutes; video uses credits faster. Sign in to keep the result in your workspace.

What happens to my uploads?

Source media normally deletes after about 3 hours unless a feature offers longer retention. Transcripts stay in your workspace.

Which languages are supported?

Auto detection handles mixed-language media. You can also pin English, Chinese, Japanese, Korean, Spanish, or French as the source.

Can my scripts or agents use it?

Yes. Create an API key billed at $1 per audio hour, or hand your AI agent the public skill file — it inherits your account permissions.