Short answer: you don't need a cloud service, or a per-minute credit meter, to caption a screen recording. Modern speech models run on your own machine: Girato downloads a ~40MB model once, then transcribes your mic or voiceover on-device, free and unlimited, producing timed cues you can edit like text. Your audio never uploads.
Why captions matter more than ever
Most feeds autoplay muted. A demo without captions is a silent film for the majority of first-time viewers. LinkedIn, X, Reels, docs pages. Captions aren't accessibility garnish; they're the difference between "watched" and "scrolled past".
The two ways tools do this
Cloud transcription (Tella, FocuSee's AI tools, Descript): audio uploads to their servers. Quality is good; the costs are per-minute pricing, credit meters, and your voice on someone else's machine. On-device transcription (Girato): the model runs locally. First run needs internet once to fetch the model; after that it works on a plane. No meter, no upload.
How to do it in Girato
- Record with the mic on (or upload a voiceover afterwards).
- Open Captions → Auto-captions from mic. First run downloads the model (~40MB, once).
- Proofread the cue list, every cue is editable text with its own timing.
- Export. Captions render into the video, in every aspect ratio.
Tips for cleaner transcripts
- Speak a touch slower than feels natural; pause at sentence ends, the model uses silence to time cues.
- Name your product once, clearly, then correct any mishears with one edit per cue.
- Cut dead air first (Girato's silence auto-cut) so cue timing stays tight.
Frequently asked questions
Can captions really be generated without uploading audio?
Are local captions accurate enough?
Why do some tools charge credits for captions?
See it on your screen
Girato records, auto-zooms and exports right in your browser, free for 7 days.
Start free No download · no card · nothing uploaded