Short answer: drop an audio or video file and this page transcribes it to an .srt subtitle file using Whisper running in your browser, the file never uploads anywhere. First run downloads the speech model (~40MB, cached after that). English model; clear speech transcribes best, and every line is editable before you download.
Why local transcription is the right default
Voiceovers and meeting audio are some of the most sensitive files you own. Cloud caption services process them on their servers, usually fine, never verifiable. Browser-side Whisper (via transformers.js) means the audio bytes stay in this tab. The trade-off is honest: the small on-device model is a touch less accurate than the largest cloud models and it's English-only here, proofread names and jargon before shipping.
Using the SRT
- YouTube: upload the .srt under Subtitles, better than auto-captions because you proofread it.
- LinkedIn/social: most feeds want burned-in captions instead, see recording for LinkedIn.
- Players: name it like the video (
demo.srtnext todemo.mp4) and VLC picks it up.
Captions burned into the video?
Girato records your screen, auto-captions your mic with the same on-device approach, and renders the captions into the export.
Try Girato free No download · no card · nothing uploaded