An illustration of one bright shape interrupting a stream of identical ones

The first 3 seconds of a video ad is 90 frames at 30 fps, and those frames are not worth the same amount. Weight each frame by how many impressions are still watching it and the opening half second carries about 24 percent of the hook's total attention while the closing half second carries about 9 percent. That ratio, not a theory about attention spans, is what decides where the subject goes, where the claim goes, and why a half-second logo animation is the single most expensive thing you can put in an ad.

Below is the frame budget worked out, with the model stated so you can swap in your own numbers.

How much is each frame in the first three seconds worth

Take Meta Reels' median hook rate of 28 percent, from our 2026 fatigue benchmark, a synthesis of published platform and agency figures rather than our own measurement. That figure says 28 percent of impressions are still there at frame 90 and 72 percent have gone.

Assume for the arithmetic that they leave at a constant rate. That is a model, not a measurement: real retention curves are steeper at the front. Under it, 72 points spread across 90 frames is 0.8 points of audience lost per frame, so the audience at frame f is 100 minus 0.8f percent.

Multiply each frame by its audience and you get viewer-frames, which is the honest unit for a frame budget. There are 5,724 of them in a Reels hook, and here is where they sit.

Segment Frames at 30 fps Share of the 90 frames Impressions watching, start to end Share of the hook's total attention
0.00 to 0.50s 1 to 15 16.7% 100 to 88% 24.5%
0.50 to 1.00s 16 to 30 16.7% 88 to 76% 21.4%
1.00 to 1.50s 31 to 45 16.7% 76 to 64% 18.2%
1.50 to 2.00s 46 to 60 16.7% 64 to 52% 15.1%
2.00 to 2.50s 61 to 75 16.7% 52 to 40% 12.0%
2.50 to 3.00s 76 to 90 16.7% 40 to 28% 8.8%

Every row is one sixth of the runtime and none is one sixth of the value. The first half second is worth 2.8 times the last half second. Anything that has to be seen by the largest possible number of people belongs in row one, and anything that only needs to reach people who have already decided to stay belongs in row six.

What has to be on screen, second by second

The table below is the plan we brief against. Timecodes are for a 30 fps export; at 24 fps multiply the frame numbers by 0.8, and at 60 fps double them.

Timecode Frames Job of these frames What must not be here
0.00 to 0.20s 1 to 6 One subject, already mid-movement, filling most of the frame. The viewer must be able to name the object. Fade from black, logo, title card, establishing wide shot, anything static
0.20 to 0.60s 7 to 18 The movement resolves into an identifiable action: hands doing a thing, a face already talking, a product already being used. Text longer than two words, a cut to a second location
0.60 to 1.20s 19 to 36 The claim, burned in, four words or fewer. The audio may say more; the text carries it alone. The price, the CTA, the brand name in isolation
1.20 to 2.00s 37 to 60 The stake. What is wrong now, or what is at risk. One cut is allowed here, not three. A beauty shot that makes no argument
2.00 to 3.00s 61 to 90 The first piece of evidence from the body, and the cut that starts it, so frame 91 is continuous rather than a restart. End card, logo lockup, offer terms

The single most common failure in this table is not the opening. It is the join at frame 90. If the hook ends and the body begins with a hard change of location, subject and audio level, you have built a second first-impression at exactly the point where 28 percent of your audience is deciding whether this was worth it. Make frame 90 and frame 91 share either the subject or the audio.

Why a logo in frame one wastes the hook

Run a half-second brand animation across frames 1 to 15 and check it against the budget above. Those 15 frames are 16.7 percent of the runtime and 24.5 percent of the total attention in the hook. A one-second sting across frames 1 to 30 takes 33 percent of the runtime and 45.9 percent of the attention.

So a one-second logo open spends nearly half of everything the opening had, on an asset that answers a question nobody scrolling has asked. The trade is worse than the percentages suggest, because the frames it buys are also the only frames where the audience is at full size. There is no second chance at frame 1.

The brand argument for early logos is recall, and it is not a stupid argument. It is just that the placement is wrong. Put the brand where it costs 8.8 percent instead of 24.5 percent: on the packaging in shot, on the interface in a screen recording, in the caption line at 2.5 seconds, or in the end card that the 28 percent who stayed will actually see. If a brand team insists on frame-one presence, a corner watermark costs a few hundred pixels rather than fifteen frames.

Planning for muted playback, not hoping against it

Plan every hook as if the sound is off, because your reporting cannot tell you what share of impressions had it on. That is not a claim that nobody hears your ads. It is a statement about what you can verify before you spend money.

The practical constraint is reading speed, and it is tighter than most scripts assume. We budget burned-in text at about three words per second of screen time, which is deliberately slower than speech, because the viewer is also watching the picture. Run that against the frame table:

  • A four-word claim needs about 1.3 seconds, or 40 frames, on screen.
  • Introduce it at frame 19 (0.60s) and it finishes being read around frame 59, just before the two-second mark. That works.
  • Introduce a six-word claim at the same point and it is still being read at frame 79, inside the least valuable segment of the hook, competing with the first evidence shot.

Which gives the rule: the on-screen claim in the first three seconds is four words, five at a push, and every extra word costs you ten frames somewhere else. Long claims are a body problem, not a hook problem.

Two more muted-playback rules that follow from the same logic. First, never let the reveal be audio-only: if the voiceover says the number and the screen does not, the number did not ship. Second, put the caption where the platform chrome is not, because a claim behind a username overlay is a claim nobody read.

What the numbers look like on other placements

The frame budget shifts with the placement's median hook rate, since that is what sets the decay rate in the model. The thumbstop rate you report is the same measurement under a different name, so the same arithmetic applies to it.

Placement Median hook rate at 3s Audience lost per frame Still watching at 1.0s Second one worth this much more than second three
Meta Reels 28% 0.80 points 76% 2.2x
Meta feed 28% 0.80 points 76% 2.2x
TikTok 33% 0.74 points 78% 2.0x
YouTube in-stream 22% 0.87 points 74% 2.5x

The spread is narrower than the hook rates suggest, which is the useful finding. Whatever placement you are cutting for, the first second is worth roughly twice the third, so one frame plan serves all four. YouTube in-stream is the strictest because it loses audience fastest, and it is also the placement where a five-second pre-skip window means the join at frame 90 matters most.

Where Genyad fits, and what it will not do

Genyad is our product, so treat this as disclosure. The frame plan above is easy to write and slow to execute, because producing eight openings that each hit these marks means eight sets of shot choices, not eight trims of one timeline. You upload footage once, Genyad transcribes and tags every clip, and each variation comes out as a fresh script, shot selection, voiceover, caption set and export drawn from that library. Exports are 9:16, 4:5, 1:1 and 16:9, at 1080p on self-serve plans and 4K on Enterprise, with no watermark on any plan. One variation costs 1 credit and the free plan gives 5 with no card. Scripts are written natively in English, German, French, Spanish, Italian or Hindi rather than translated, which matters here because a four-word claim in English is rarely four words in German. The hook variation generator is the route for openings specifically.

What it does not do, plainly. No AI avatars or synthetic presenters, no static banner formats, no product-URL import, no product-feed or CSV-driven template rendering, no predicted performance scores, and no direct publishing to Meta or TikTok. It also cannot invent a frame you never filmed. If your library has no shot of a hand already in motion, no tool will build you one from it.

Frequently asked questions

How many frames is the first three seconds of a video ad?

Ninety frames at 30 fps, 72 at 24 fps, and 180 at 60 fps. Working in frames rather than seconds is worth the effort because the decisions are that fine: the difference between a logo on frames 1 to 15 and a logo on frames 76 to 90 is about 16 percentage points of your hook's total attention.

Should a logo appear in the first frame of a video ad?

No. On the frame budget above, a half-second brand animation consumes about 24 percent of the attention available in the whole three-second opening and a one-second sting consumes about 46 percent. Put brand presence on the product in shot, in the interface, or in the end card, where it costs a fraction of that.

How much on-screen text fits in the first three seconds?

Budget about three words per second of screen time, so four words is comfortable and six is not. A four-word claim introduced at 0.6 seconds finishes being read just before the two-second mark, which leaves the last second for the first evidence shot. Anything longer belongs in the body.

What should happen at the three-second mark itself?

Frame 90 and frame 91 should share either the subject or the audio, so the body reads as a continuation rather than a second opening. This is the join most teams ignore, and it sits exactly where the audience has shrunk to the placement's hook rate, around 28 percent on Meta feed and Reels, so a restart there costs you the people who already agreed to stay.