Short answer: record your voiceover after editing the picture, keep music around 30–40% of voice volume, choose instrumental tracks, and let interaction sounds (clicks, typing) sit just under the voice. Screen videos with layered, controlled audio feel produced; single-layer audio feels like a screencast.
Voiceover: after picture, not during
Narrating while driving the demo gives you "um"-laden audio and mistimed clicks, you're doing two jobs badly. Instead: record the screen silently, edit it (trim, zooms, ramps), then narrate against the finished cut. Your sentences land exactly on the visuals, and a flubbed line costs a ten-second retake instead of a re-record.
- Any quiet room + a decent USB or headset mic beats a great mic in an echoey room.
- Write beats, not scripts, one line per visual moment reads as natural speech.
- Keep voice as the loudest element in the mix, always.
Music: the 35% rule
Background music exists to remove silence-discomfort on longer videos, not to be heard. Practical settings: instrumental only (lyrics fight narration), steady energy (no big drops mid-explanation), volume at roughly a third of the voice. If you notice the song while listening to the words, it's too loud.
For sub-60-second social clips with no voiceover, invert the rule: music carries the energy at full volume and the interaction sounds punch through it.
The three-layer mix, in order
- Voice, 100%, the reference layer.
- Interaction sounds, clicks/typing just below voice; they're texture, not events.
- Music, 30–40% of voice, faded in and out at the ends.
In Girato this is two uploads and two sliders: add a voiceover file (plays from clip start) and a music file (auto-loops), set volumes, and the whole mix, including the generated click and typing sounds, renders into the export on your machine.
See it on your screen
Girato records, auto-zooms and exports right in your browser, free for 7 days.
Start free No download · no card · nothing uploaded