Whisper-quality SRT subtitles, straight from your video.
Whisper transcribes audio with sentence- and word-level timestamps, which is exactly what subtitle files need. Upload a video or audio file, run Whisper, edit any cues that need a fix, and export an SRT or VTT you can drop straight into YouTube, Vimeo, Premiere, Final Cut, DaVinci Resolve, or your LMS. No timecoding by hand. No SubRip syntax to remember.
How to export Whisper subtitles as SRT
- 01
Upload your video or audio file
MP4, MOV, MP3, M4A, WAV — anything with a dialogue track works. You don't need to extract the audio first.
- 02
Run the transcription
Whisper transcribes the dialogue and produces word- and segment-level timestamps in one pass.
- 03
Review and fix the cues
Open the transcript in the editor, correct proper nouns, jargon, and any misheard words. Adjust cue breaks if a sentence runs long.
- 04
Decide on speaker labels
For interviews keep them in; for monologue, turn them off. Toggle once before export.
- 05
Export as SRT or VTT
Download the file and upload it to YouTube, Vimeo, your LMS, or drop it into Premiere/Final Cut/DaVinci alongside the video.
Who this is for
- •YouTubers who need accessible, search-indexed captions on every upload
- •Course creators producing video lessons that require captions
- •Video editors who'd rather fix a few cues than time-code from scratch
- •Documentary and short-form producers shipping multilingual subtitles
- •Marketing teams making social-first video that plays sound-off
Who this is not for
- •Live broadcasts where you need real-time captioning (use a live-caption tool instead)
- •Karaoke-style word-by-word highlighting (look for a dedicated lyric-sync tool)
- •Translated subtitles that require certified review (you'll still want a human translator)
Platform recipes for an exported SRT
Once you have the file, here's where it goes — the exact menu path on each platform we get asked about most.
| Platform | Where to upload the SRT |
|---|---|
| YouTube | Studio → Subtitles → Add → Upload file → SRT |
| Vimeo | Video Manager → Distribution → Subtitles → Add SRT |
| HTML5 video on a website | <track src="captions.vtt" kind="subtitles"> — use VTT |
| Premiere Pro | Window → Captions → Import sidecar SRT, then style and render |
| Final Cut Pro | File → Import → Captions, choose the SRT |
| DaVinci Resolve | Right-click the timeline → Import Subtitle |
| Loom / Riverside / Descript | Caption settings → Upload SRT |
| LMS platforms (Canvas, Moodle, Thinkific) | Add captions field on the video — accepts SRT or VTT |
Cue conventions worth following
An SRT file is just numbered cues with start and end timecodes. Whisper handles the timecodes for you, but a few editorial rules separate readable captions from a wall of text on screen:
- Keep each cue to one or two lines of around 32–42 characters per line.
- Show a cue for at least one second, no more than seven.
- Break cues at natural pauses, not mid-clause.
- Keep speaker labels short — "Anna:" reads better than "Speaker 1 (Anna):".
- Spell proper nouns and product names the way the brand spells them; fix once in the editor before exporting.
More Whisper guides
For a longer walk-through of subtitling videos, see the video transcription & subtitles guide. If you'd rather avoid the install path entirely, see use Whisper without coding, or run it straight in the browser with OpenAI Whisper online. For interview-style videos, Whisper speaker detection covers how labels flow into your SRT. Comparing tools? See WhisperAI vs TurboScribe and the alternatives hub.
Frequently asked questions
What is an SRT file and why do I need one for my video?
SRT (SubRip Subtitle) is the most widely supported subtitle format. It's a plain-text file with numbered cues, each with a start and end timecode and the line of text to show. YouTube, Vimeo, TikTok (via third-party tools), Premiere Pro, Final Cut, DaVinci Resolve, and most LMS platforms all accept SRT directly. You need one whenever you want captions to display on a video — for accessibility, for sound-off viewing, or to satisfy a platform's caption requirement.
What's the difference between SRT and VTT?
Both are caption formats that share the same basic idea — timecoded lines of text. SRT is older and slightly simpler; VTT (WebVTT) is the modern HTML5 standard and supports a few extras like styling and positioning hints. For most YouTube uploads either works. For HTML5 video on a website, prefer VTT. WhisperAI exports both, so you can pick whichever your platform asks for.
How accurate are Whisper's subtitle timecodes?
Whisper produces word-level and segment-level timestamps that are accurate enough for most published video work, especially on clean dialogue. On music-heavy tracks, overlapping speakers or very fast speech you'll usually want to scrub through the cues and nudge a few. A good editor lets you correct timing without re-running the whole transcription.
Can I edit the SRT before exporting?
Yes. A hosted Whisper tool lets you edit the transcript line by line in the browser, then export the corrected SRT. That's the workflow you want — fix typos, proper nouns and punctuation once, then export, instead of opening the SRT in Notepad and editing raw timecodes.
Will the SRT include speaker labels?
If you enable speaker detection before exporting, your SRT cues can be prefixed with labels like "Speaker 1:" or names you've assigned. Not every video platform displays these gracefully — YouTube, for example, treats them as part of the caption text. For interview-style videos it's almost always worth keeping them. For monologue, switch them off.
What's the line-length limit for SRT cues?
There's no hard maximum, but two readable conventions are: keep each cue under ~42 characters per line and split long sentences across two cues rather than one wall of text. A Whisper tool with a real editor will already chunk cues sensibly; if you're hand-rolling SRT, aim for cues of 1-2 lines that stay on screen for 1-7 seconds.
Can I burn subtitles into the video instead of using a sidecar SRT?
That's a separate step — "hardcoded" or "burned-in" subtitles get rendered into the pixels. WhisperAI exports the SRT/VTT file; you then load it into your video editor (Premiere, Final Cut, DaVinci Resolve) or a free tool like HandBrake to burn it in. The advantage of a sidecar SRT is that viewers can turn captions on and off.
Ship a captioned video this afternoon.
Upload, edit, export. Your SRT will be ready before your render finishes.