
A scroll-stopping hook does not stop a scroll. It survives a glance of roughly 0.4 seconds, which is 12 frames at 30 fps and about 1.2 readable words, and then it either earns the next two seconds or it does not. The measurable version of the phrase is your hook rate divided by the median hook rate for that placement: 31 percent is a loss on TikTok, where the median is 33 percent, and a comfortable win on YouTube in-stream, where it is 22 percent. Same file, two verdicts, and the phrase "scroll-stopping" cannot tell them apart.
What actually fits inside a 0.4 second glance
Two budgets, both small.
Frames: at 30 fps a glance is 12 frames. At 24 fps it is 10 and at 60 fps it is 24. Whatever the frame rate, the whole event is over before a one-second logo animation has finished resolving.
Words: we budget burned-in text at about three words per second of screen time, slower than speech because the viewer is also looking at the picture. Three words a second across 0.4 seconds is 1.2 words. One word. Possibly two if one of them is "the".
That single arithmetic result is the deflation. If a glance carries about one word, then no line of copy is what survives it. What survives a glance is a shape, a face angle, a colour block or a direction of movement. The line matters enormously at 1.2 seconds and not at all at 0.2.
There is a physical version of the test that takes two minutes and settles most arguments. Export frame one as a still. Scale it to 15 percent, roughly its size in a feed on a phone held at arm's length while somebody scrolls. Show it to a colleague for four tenths of a second. Then ask them one question: what was it. If the answer is a noun, the frame works. If the answer is "an ad", it does not.
Why motion beats text in frame one
Text has a scale problem and a time problem at once. A wordmark or a six-word headline that is crisp on a 27-inch review monitor is a grey smudge at 15 percent scale, and even legible it needs 2 seconds to be read against a 0.4 second budget. Motion has neither problem: a hand moving across the frame is a hand moving across the frame at any scale, and it is identifiable in three frames.
That is why the sequence matters more than the content. Movement first, so the frame reads as something happening. Then the object, so the viewer can name it. Then, at around 0.6 to 1.2 seconds, the four-word claim, which is now being read by someone who has already stopped.
Text in frame one is not banned, it is just limited to one word doing one job: a number, a price, a place, a negation. "Wrong." works. "The three biggest mistakes people make with their morning routine" does not, and the reason is not that it is a bad line. It is that it is 11 words against a 1.2 word budget, and it will still be unread when the audience has already halved.
Three hooks that read well in a doc and fail in a feed
The review surface is the problem. Hooks get approved as rows in a spreadsheet, where a line either reads well or it does not, and the 12-frame constraint is invisible.
"Have you ever struggled with X?"
In a doc this is a clean problem statement. In a feed it is six words at a 1.2 word budget, and the 12 frames show a person with their mouth open and nothing else in shot. The glance returns "someone talking", which is the most common thing in the feed and therefore the least interesting.
The fix keeps the line and moves it. Let the voiceover ask the question while frame one shows the thing failing: the drawer that will not close, the dashboard with the flat line, the shirt with the stain. The glance returns a noun, and the question lands at 1.0 seconds to the people who stayed.
The clean brand title card
This one fails the scale test rather than the time test. At 15 percent scale a wordmark is illegible, so the glance returns no noun at all, which is worse than returning the wrong one. It also spends the frames with the largest audience of the entire ad on information that answers no question the viewer has.
The fix is not to hide the brand. It is to attach it to a shape: the logo on the packaging in a hand, the product name on the interface in a screen recording. Both survive scaling because the recognisable object carries them.
The static "three reasons your X is not working" list frame
A list frame has no motion vector. Across 12 frames nothing changes, so the frame is indistinguishable from a still image that ended up in a video slot, and the viewer's eye has nothing to track.
The fix is to make reason one an action and let the count arrive as an on-screen counter at about 1.2 seconds. Same script, same argument, but the glance now returns movement instead of typography.
The wider fix is procedural. Review hooks as 12-frame clips at feed scale, not as lines in a doc. If you are cutting for TikTok specifically, a TikTok ad maker workflow that outputs finished 9:16 files is what makes that review possible at all, because you cannot glance-test a script.
The measurable replacement: hook rate over the placement median
Drop the adjective and use a ratio. Divide your hook rate by the median for the placement, which our 2026 fatigue benchmark, a synthesis of published platform and agency figures rather than our own measurement, puts at 28 percent for Meta Reels and Meta feed, 33 percent for TikTok and 22 percent for YouTube in-stream.
| Reported hook rate | Placement | Placement median | Hook rate over median | Call |
|---|---|---|---|---|
| 31% | TikTok | 33% | 0.94 | Below median. Replace it. |
| 31% | Meta Reels | 28% | 1.11 | Modest winner. Keep. |
| 31% | YouTube in-stream | 22% | 1.41 | Top of the set. Scale it. |
| 24% | Meta feed | 28% | 0.86 | Below median. Replace it. |
| 24% | TikTok | 33% | 0.73 | Well below. Kill it. |
| 24% | YouTube in-stream | 22% | 1.09 | Marginal. Judge on CTR. |
Two of those rows are 31 percent and three verdicts come out of them. That is what the phrase "scroll-stopping" hides.
The bands are not a matter of taste either. A hook rate read on roughly 2,000 impressions can settle a gap of about 3 percentage points at a 28 percent base rate, and 3 points over 28 is 0.11. So a ratio between 0.90 and 1.10 is the same as the median as far as your sample can tell, below 0.90 is a real loss, and above 1.10 is a real win. The resolution of the decision is set by how much media you bought, not by how the hook felt in the review.
How much glance you actually own
The frequency ceilings put a limit on the whole exercise that is worth seeing written down. Multiply the weekly frequency you can run before decline sets in by the 0.4 second glance.
| Placement | Weekly frequency ceiling | Glances per person per week | Glance time per person per week | Does the glance model apply |
|---|---|---|---|---|
| Meta Reels | 2.5 | 2.5 | 1.0s | Yes |
| Meta feed | 2.5 | 2.5 | 1.0s | Yes |
| TikTok | about 3.0 | 3.0 | 1.2s | Yes, and the fastest of the three |
| YouTube in-stream | 3 to 7 per week | 3 to 7 | 1.2 to 2.8s | No: five seconds are compulsory before the skip |
One second per person per week on Meta prospecting. That is the budget the phrase is really describing, and it explains why hook variety beats hook polish: you are not trying to win one glance, you are trying to be worth a noun on three separate occasions with a fatigue clock running. Meta internal research cited in the benchmark report puts the CTR drop at 45 percent after a fourth exposure to the same creative, and 2.5 a week reaches the fourth exposure in about eleven days.
The bottom row is the honest exception. In-stream is not a glance format. The first five seconds are forced, which is why "scroll-stopping" is a category error there and why the median hook rate is still the lowest of the four despite the compulsory window.
Where Genyad fits, and what it will not do
Genyad is our product, so read this as disclosure. Nothing above requires our tool: the glance test needs an export and a colleague, and the ratio needs a spreadsheet. What we solve is the step in between, which is producing enough finished openings to have something to glance-test. Upload footage once, Genyad transcribes and tags every clip, and each variation comes out as a fresh script, shot selection, voiceover, caption set and export from that library rather than a re-cut of one timeline. One variation costs 1 credit and the free plan gives 5 with no card.
What it does not do. No AI avatars or synthetic presenters, so a face you never filmed is out of reach. No static banner formats, no product-URL import, no product-feed or CSV-driven template rendering, no predicted performance scores, and no direct publishing to Meta or TikTok. And it cannot manufacture motion that is not in your footage. A library of static product photography will not produce a frame one that passes the glance test, whatever is written over it.
Frequently asked questions
What makes a hook scroll-stopping?
Nothing, as an inherent property. A first frame either returns a recognisable noun when shown at feed scale for about 0.4 seconds or it does not, and the result you can act on is your hook rate divided by the median for that placement. Above about 1.10 is a real winner, below about 0.90 is a real loss, and anything between is inside the noise of a 2,000-impression read.
How long is the glance a hook has to survive?
About 0.4 seconds, which is 12 frames at 30 fps and roughly 1.2 words of readable on-screen text. That budget is why a shape or a movement can survive a glance and a headline cannot, and why the copy belongs at 0.6 to 1.2 seconds rather than at frame one.
Is thumbstop rate the same as hook rate?
They measure the same thing, video plays past a threshold divided by impressions, under two names. The one difference that matters is the threshold: the thumbstop rate label is more often attached to a two-second read and hook rate to a three-second read, so confirm which you are looking at before comparing two numbers.
Should the first frame have any text at all?
One word, doing one job: a number, a negation, a place. More than that is unread at a 1.2 word budget, and it occupies the frames with the largest audience the ad will ever have. Put the four-word claim in around 0.6 to 1.2 seconds, once motion has already earned the attention.