12 00:00:48,120 --> 00:00:51,400 Every cue lands on time.
Sentence timestamps
SRT subtitles built from complete utterance timestamps.
Upload audio or video, or paste a public video link. Get timestamped transcripts, speaker segments, SRT, and clean text without choosing between transcription engines.
Welcome back — today we're talking about subtitles.
Every word lands with its own timestamp.
Export SRT the moment the last cue settles.
Free daily run
Free daily check: one media job under 5 minutes after sign-in.
3-hour retention
Uploaded source media expires after 3 hours unless explicitly retained.
Stripe checkout
Stripe handles card payment for global users.
7 source languages
Auto detection, or pin English, Chinese, Japanese, Korean, Spanish, or French.
Audio, video, or a public video URL — the workbench accepts all three.
Task shape, language, and duration decide the path. You never compare vendors.
SRT, word timestamps, speaker turns, and clean full text — separately or as one ZIP.
Every completed transcription exposes the same four layers, ready to copy or download.
12 00:00:48,120 --> 00:00:51,400 Every cue lands on time.
SRT subtitles built from complete utterance timestamps.
{"w":"every","s":0.42,"e":0.61}
{"w":"word","s":0.63,"e":0.85}
{"w":"字","s":0.88,…}One JSON line per word or character, with start/end time.
S1 00:00:04 → 00:00:11 "Welcome back to the show." S2 00:00:11 → 00:00:19
Sentence-level speaker turns with text and word ranges.
Welcome back to the show. Today, we're talking about subtitles — and why timing is everything.
The provider's full transcript with punctuation.
1 credit ≈ 1 audio minute. The $20 starter pack covers about 20 audio hours.
$20
1,200 credits
Media is routed to the engine that best fits its language and shape. Word timestamps, punctuation, and speaker turns come from the same pass, so every layer agrees.
Free daily check: one audio or video transcription under 5 minutes. Sign in to start and keep the result in your workspace.
Source media auto-deletes after 3 hours unless you explicitly retain it. Transcripts stay in your workspace.
Auto detection handles mixed-language media. You can also pin English, Chinese, Japanese, Korean, Spanish, or French as the source.
Yes. Create an API key billed at $1 per audio hour, or hand your AI agent the public skill file — it inherits your account permissions.