Audio, video, and link transcription

Turn media into accurate transcripts and SRT.

Upload audio or video, or paste a public video link. Get timestamped transcripts, speaker segments, SRT, and clean text without choosing between transcription engines.

Multi-language source detectionWords, speakers, SRT, full textSource media auto-deletes in 3 hours
podcast-ep-12.mp3TranscribingAuto language: EN+ZH
00:00:01,000 --> 00:00:04,120Speaker 1

Welcome back — today we're talking about subtitles.

00:00:04,480 --> 00:00:07,900Speaker 2

Every word lands with its own timestamp.

00:00:08,260 --> 00:00:11,540Speaker 1

Export SRT the moment the last cue settles.

Words 2,184Speakers 2Cues 46SRT ready
Free daily check: one audio or video transcription under 5 minutes. Sign in to start and keep the result in your workspace.

Free daily run

Free daily check: one media job under 5 minutes after sign-in.

3-hour retention

Uploaded source media expires after 3 hours unless explicitly retained.

Stripe checkout

Stripe handles card payment for global users.

7 source languages

Auto detection, or pin English, Chinese, Japanese, Korean, Spanish, or French.

How it works

00:01

Drop media or paste a link

Audio, video, or a public video URL — the workbench accepts all three.

00:02

We route it to the best-fit engine

Task shape, language, and duration decide the path. You never compare vendors.

00:03

Export every layer you need

SRT, word timestamps, speaker turns, and clean full text — separately or as one ZIP.

One job, four deliverables.

Every completed transcription exposes the same four layers, ready to copy or download.

12
00:00:48,120 --> 00:00:51,400
Every cue lands on time.

Sentence timestamps

SRT subtitles built from complete utterance timestamps.

{"w":"every","s":0.42,"e":0.61}
{"w":"word","s":0.63,"e":0.85}
{"w":"字","s":0.88,…}

Word timestamps

One JSON line per word or character, with start/end time.

S1 00:00:04 → 00:00:11
  "Welcome back to the show."
S2 00:00:11 → 00:00:19

Speaker segments

Sentence-level speaker turns with text and word ranges.

Welcome back to the show. Today,
we're talking about subtitles — and
why timing is everything.

Punctuated full text

The provider's full transcript with punctuation.

Pay for minutes, not seats.

1 credit ≈ 1 audio minute. The $20 starter pack covers about 20 audio hours.

$20

1,200 credits

See full pricing

Frequently asked

How accurate are the transcripts?

Media is routed to the engine that best fits its language and shape. Word timestamps, punctuation, and speaker turns come from the same pass, so every layer agrees.

What can I do for free?

Free daily check: one audio or video transcription under 5 minutes. Sign in to start and keep the result in your workspace.

What happens to my uploads?

Source media auto-deletes after 3 hours unless you explicitly retain it. Transcripts stay in your workspace.

Which languages are supported?

Auto detection handles mixed-language media. You can also pin English, Chinese, Japanese, Korean, Spanish, or French as the source.

Can my scripts or agents use it?

Yes. Create an API key billed at $1 per audio hour, or hand your AI agent the public skill file — it inherits your account permissions.