
Write ad copy for video ads as captions first and voiceover second, because most of your audience will watch with the sound off. That means the copy has to carry the argument in two-to-four word chunks that a person can read at arm's length while their thumb hovers, not in sentences that sound good read aloud. If the ad still makes its case with the audio muted and the captions covered by a comment overlay, the copy is finished. If it does not, no amount of voice talent will save it.
Why the caption is the copy
A video ad has three copy layers: the primary text above the video, the voiceover, and the burned-in captions. Only the third one is guaranteed to be read. Primary text gets truncated after a couple of lines, and the voiceover only exists for the minority who tap for sound.
So the working order changes. Write the caption cards, get the argument standing on its own, then write a voiceover that says roughly the same thing in a way a human would actually speak it. Doing it the other way round produces captions that are a transcript, and a transcript is a bad caption track: it is too long, it breaks in the wrong places, and it puts the verb on the card after the one where the viewer stopped reading.
The practical test is brutal and takes ten seconds. Mute your ad, play it once, and write down what you think it was selling. Most drafts fail that on the offer, not the hook.
How to chunk captions
Two to four words per card, one idea per card, and never split a phrase across cards in a way that changes its meaning. Here is the same sentence chunked badly and well.
| Version | Cards | Problem |
|---|---|---|
| Transcript style | "Our new reconciliation tool pulls both of your systems into one" / "view so month-end close takes three days instead of nine" | Two cards, both unreadable at speed, the number arrives after the viewer has gone |
| Sentence style | "Month-end close took nine days." / "Now it takes three." | Readable, but the product never appears |
| Chunked | "NINE DAYS." / "MONTH-END CLOSE." / "NOW THREE." / "TWO LEDGERS, ONE VIEW." / "FREE TO CONNECT." | Each card is one beat, each survives being read alone, the offer is on card five |
Rules we apply without arguing about them:
- Numbers get their own card. A figure buried mid-sentence is not read.
- The verb stays with its object. "Cancel" and "any time" on separate cards reads as a threat.
- No card is a fragment that needs the next card to make sense, with one exception: a deliberate two-card reveal, used once per ad at most.
- The offer gets its own card, in the first half of the ad.
- Punctuation carries tone the voice cannot. A full stop between chunks reads slower than a line break.
Reading speed at phone size
The constraint is not how fast people read, it is how fast they read while doing something else on a screen the size of a hand. Typical working numbers, from writing a lot of these rather than from a study:
| Chunk length | Minimum time on screen | Where it works |
|---|---|---|
| 1 to 2 words | 0.6 to 0.8 seconds | Hook cards, numbers, single-word emphasis |
| 3 to 4 words | 0.9 to 1.2 seconds | Body of the ad, the default |
| 5 to 7 words | 1.4 to 1.8 seconds | Testimonial quotes, where the voice carries it |
| 8 or more | Do not | Nowhere in a 15 second ad |
Two more things that decide whether copy gets read at all. First, the caption block belongs in the middle third vertically, not the bottom: on 9:16 placements the bottom 15 percent or so of the frame sits under the interface, and copy there is simply gone. Second, keep the type big enough that you can read it on your own phone at desk distance with the screen at half brightness. If you are squinting on a good screen, a commuter on a cracked one is not reading it. How we handle captions and text overlays covers the placement and safe-area side of this, and the auto caption generator for ads does the chunking and timing rather than dumping a transcript on the frame.
Voiceover copy and caption copy are different drafts
They carry the same argument in different registers. Treat them as two drafts of one script, not one asset used twice.
| Voiceover copy | Caption copy | |
|---|---|---|
| Unit | The sentence | The card, 2 to 4 words |
| Length for 15s | 35 to 40 words of English | 20 to 28 words across 7 to 10 cards |
| Tone tools | Pace, pause, breath, emphasis | Capitals, line breaks, full stops, one colour change |
| Handles | Nuance, warmth, a caveat | Numbers, names, the offer, the ask |
| Cannot do | Be read silently | Sound like a person |
| Fails as | Copy that reads flat on screen | Copy that sounds robotic when spoken |
A caveat is the clearest example of the split. A voiceover can say "most people see this in about a fortnight, some take longer" and sound honest. On a caption card it reads as hedging and costs you the sale. Put the caveat in the voice and the number on the card.
Testing caption-only against voiced
Run this as a real test, not an assumption. Build one version with voiceover plus captions and one silent version with captions and music or ambient sound only, keep the shots and the offer identical, and split them evenly.
What the result usually means:
- Silent version wins on hook rate, voiced version wins on click-through: your captions are doing the stopping and your voice is doing the persuading. Keep both, and shorten the voice.
- Silent version wins on both: the voiceover is fighting the captions. Two competing lines of copy at once is a common cause, and it is fixable by cutting the voice down to the cards.
- Voiced version wins on both: your captions are probably a transcript. Rewrite them as chunks and retest.
- Neither moves: the offer is the problem, and no copy layer is going to rescue it.
This is worth doing early because the answer holds for a while, and creative decays fast enough that you do not want to re-learn it every month. Our 2026 fatigue benchmark reports CTR falling 15 to 20 percent in a creative's first two weeks and week three landing 45 to 70 percent below launch, so the copy format you settle on gets reused across a lot of variations.
That reuse is what Genyad, our product, is for: your footage is transcribed and tagged once, then each variation gets its own script, shot selection, voiceover and caption set in 9:16, 4:5, 1:1 or 16:9. If you would rather write the argument yourself and generate the variants around it, the AI ad script generator takes the offer and the proof and writes both layers. We do not make static banners, so if the campaign needs display sizes from the same copy, that is a different tool.
Frequently asked questions
Should video ad copy be written for sound on or sound off?
Sound off, then adapted for sound on. Write the caption cards until the ad makes its case muted, then write a voiceover that says the same thing in speakable sentences. Assuming the audio will be heard is the most expensive assumption in short-form video.
How many words per caption card?
Two to four words for the body of the ad, with 1 to 2 word cards for hooks and numbers, held on screen for 0.9 to 1.2 seconds at the default length. Anything over seven words on a card in a 15 second ad will not be read.
Where should captions sit in the frame?
In the middle third vertically. On 9:16 placements the lower part of the frame is covered by the platform interface, and copy placed there disappears on the placements you care most about.
Is the primary text above the video part of the ad copy?
Yes, but treat it as the least reliable layer. It gets truncated, so put the offer in the first line and never rely on it to explain something the video leaves out.