A scripting document laid out in a timing column and a spoken line column, with the first row highlighted and a struck-through setup line above it

A TikTok ad script opens on the claim, in the first spoken line, with no setup in front of it. The structure that works is claim, evidence, objection, ask, delivered in 9 to 21 seconds, and the single most common fix we make to a submitted script is deleting its first sentence. Setup exists for the writer's benefit, not the viewer's: the frame already supplies the context the setup was explaining.

The template, with timings

This is a 15 second script, which sits in the middle of the range most in-feed offers land in. Stretch or compress the middle rows, never the first.

Time Beat What it does Length
0.0 to 1.5s Claim The outcome or the problem, flat delivery, no wind-up 6 to 12 words
1.5 to 4.0s Evidence Shows it being true, ideally without narration One visual, one line
4.0 to 8.0s Objection Names the reason not to buy, then answers it 15 to 25 words
8.0 to 12.0s Proof Price, number, comparison or someone else's verdict 10 to 15 words
12.0 to 15.0s Ask One action, named specifically, once 6 to 10 words

Two notes on using it. The objection row is what separates a script from a slogan, and it is the row people cut first when shortening. Cut the proof row instead: a 12 second version with the objection intact beats a 12 second claim, price, buy.

And the ask is a single sentence. Two calls to action in the last three seconds splits the response. Our hook, body, CTA glossary entry covers why the three-part shape holds across platforms even though the timings do not.

Why the setup always gets cut

Every first draft arrives with a version of this line in front of the claim: "So I've been looking for a way to keep my white trainers clean for ages." It reads naturally, it is how a person starts talking, and it is why the ad fails. Three reasons it goes.

The frame already did that work. The viewer can see the trainers. Narrating what is visible is the most reliable way to spend a second telling someone nothing.

It moves the claim past the decision point. TikTok's median hook rate is about 33 percent according to the platform medians in our 2026 fatigue benchmark, so roughly two thirds of impressions do not survive the opening. If the claim lands at second two, most of the people you paid for never hear it.

It is a comfort structure for the writer. Setup makes a script feel logically ordered, and short-form does not reward logical order. It rewards the interesting part first and the reasoning afterwards, which feels wrong to write and right to watch.

The test is mechanical: delete your first sentence and read the script again. If it still makes sense, and it will roughly nine times out of ten, that sentence was setup. If it genuinely stops making sense, you have a product that needs a demo, and the fix is to open on the demo rather than on an explanation of it.

Where setup is allowed: as the second beat. Claim, then one line of context, then evidence, is a legitimate order for anything high-consideration. Context before claim is not.

Spoken or caption-carried?

Sound is on by default on TikTok, so the default is that speech carries the argument and captions reinforce it. That is the opposite of Meta feed, where you have to assume muted playback and the caption carries the claim.

The division we use:

  • Speech carries the objection, the concession, the tone, anything that needs to sound like a person rather than a graphic. Nuance dies in text.
  • Captions carry anything mishearable or checkable: the price, the product name, a number, a percentage, a timeframe. "Fourteen pounds" is easy to mishear and impossible to misread.
  • Both carry the claim, but not word for word. Identical spoken and written text makes the viewer read rather than listen, and reading is slower.

Placement matters as much as content. Captions belong in the middle third of the frame, because TikTok's interface eats roughly the bottom 480px of a 1080x1920 canvas plus about 140px down the right edge. Lower-third captions in brand typography are a Meta convention and get partially covered here. The TikTok ad specs post has the safe area measurements.

One more caption rule: never let a caption be the only place a required qualifier appears. Put it in speech too, because captions get cropped in ways your review tool will not show you.

Two worked scripts

A physical product, 15 seconds

The shot list is deliberately plain, because this should be shootable on a phone in one location.

Time Shot Spoken On-screen
0.0 to 1.5s Hands rubbing a stained white trainer, close "This pen took a two week old coffee stain out of white leather." STAIN PEN
1.5 to 4.0s Same shot, stain visibly lifting, no cut Silence, just the sound of rubbing none
4.0 to 8.0s Person to camera, arm's length, kitchen behind "I assumed it would bleach the leather. It did neither, and I've done four pairs with the same pen." none
8.0 to 12.0s Trainer before and after, side by side "Eleven pounds, and it replaced the forty pound kit I never used." £11
12.0 to 15.0s Person to camera "Tap the link, it lives in my hallway drawer now." none

Notice what is not in there. No greeting, no brand introduction, no second CTA, no music carrying the emotional weight. The objection beat at second four is doing the selling.

A B2B SaaS product, 20 seconds

Longer, because a workflow needs showing.

Time Shot Spoken On-screen
0.0 to 1.5s Screen recording, a rota being dragged into place "Our shift rota took four hours a week. It now takes twenty minutes." 4 HRS TO 20 MIN
1.5 to 5.0s Continued screen recording, one drag, conflicts flagging red "It flags the conflicts before I publish, so I stopped getting the Sunday night phone calls." none
5.0 to 11.0s Person to camera, office, handheld "We did this in a spreadsheet for two years. The problem was never the grid, it was that nobody knew when it changed." none
11.0 to 16.0s Screen recording, notification hitting staff phones "Staff get the change on their phone. Twelve managers, two sites, no printouts since." 12 MANAGERS, 2 SITES
16.0 to 20.0s Person to camera "Free for the first rota, link's below." none

The B2B version breaks one rule deliberately: the first line is a number, not a benefit. For an operational buyer, four hours to twenty minutes is the benefit, and abstracting it into "save time on scheduling" is weaker.

Where the script writing sits in production

Genyad is our product, so read the following with that in mind. Scripting is where the variation lives: you upload footage you already own, it transcribes and tags every clip, and each variation is a fresh script with its own shot selection, voiceover and caption set drawn from that library rather than the same timeline with new text over it. That matters for the structure above, because a re-cut cannot change where the claim sits and a new script can. The AI ad script generator page covers the scripting side, and the TikTok ad maker page the platform build.

Scripts are written natively in English, German, French, Spanish, Italian and Hindi rather than translated from an English master, which matters for this template: a claim-first opening translated word for word from English often lands as an odd construction in German, where a native draft reorders it.

The limits, plainly. No predicted performance score, so nothing will tell you which of five scripts wins. No AI avatars, so a script needs footage of a real person to be delivered by one. And no direct publishing to TikTok, so you export the file and set the campaign up in Ads Manager yourself.

Frequently asked questions

How long should a TikTok ad script be?

Write to 9 to 21 seconds of screen time, roughly 25 to 60 spoken words. A 15 second script with claim, evidence, objection, proof and ask is the standard shape. If it will not fit, cut the setup and the second proof point rather than the objection beat.

Should the product be named in the first line of a TikTok ad script?

Name what it does, and name the product only if the name means something to a stranger. "This pen took a coffee stain out of white leather" beats a brand nobody knows, and it beats "struggling with stains?" by a wide margin, because a claim invites a question while a question invites a no.

Do TikTok ad scripts need a voiceover, or can captions carry them?

Sound is on by default on TikTok, so speech should carry the argument and captions should reinforce numbers, prices and product names. Caption-only scripts work but give up the tone that makes an objection sound honest. On Meta feed the opposite holds, because muted viewing puts the claim in the caption.

What should the call to action in a TikTok ad script say?

One action, named specifically, in one sentence, in the last two to three seconds. Attach a reason to it rather than giving a bare instruction, and do not stack a second ask behind it, because two calls to action split the response.