One long bar breaking into short segments that regroup vertically

Count usable clips, not hours. Around 30 usable clips is the point where ad variations stop repeating the same openings, and below roughly 15 you will be building the same ad with different words no matter how many scripts you write. An hour of long-form video typically yields 10 to 25 usable clips, so two or three hours of assorted source material is a realistic first library. Hours are a measure of what you filmed, and clips are a measure of what you can actually build with.

Why clip count is the right unit

A variation is assembled from clips: an opening shot, an evidence beat, a cutaway or two, a closing product shot. That means the number of distinct ads you can build is a function of how many clips you have in each role, not of total duration.

This is why footage libraries fail in a counter-intuitive way. A team films a beautiful two minute brand film and has, in practice, about four usable clips: a wide establishing shot, a product hero, one demo and one closing shot. Another team hands over 40 minutes of unglamorous rushes from a phone and has 30 clips across eight roles. The second library produces better ads and more of them.

The failure mode with a thin library is specific and recognisable: every ad starts with the same three seconds. You can write ten different scripts, and because there are only two shots strong enough to open on, the set reads as one ad with new captions. Frequency does the rest, and our benchmark reports CTR dropping 45 percent after a fourth exposure to the same creative.

What can you build at each clip count?

Usable clips Realistic variations What happens
Under 10 2 to 3 Every ad opens the same way, and it is obvious
10 to 20 5 to 8 Works, with visible repetition in openings and cutaways
20 to 30 8 to 12 Enough for a real prospecting set, some roles still thin
Around 30 12 to 15 Openings stop colliding, the set feels genuinely different
30 to 60 15 to 30 Comfortable, you can afford to test length and pacing too
Over 60 30 plus Curation becomes the constraint rather than supply

The 8 to 20 live variations that our benchmark reports as typical for an active campaign therefore need somewhere between 20 and 40 usable clips behind them. If you are aiming at the report's throughput finding, that brands shipping 15 to 50 variants a month see 3 to 5 times longer campaign lifespan than quarterly refreshers, plan on a library you keep adding to rather than a one-time shoot. Those figures come from our 2026 fatigue benchmark, which synthesises published platform and agency data rather than reporting our own measurement.

What counts as a usable clip?

Stricter than most people assume. A usable clip is two to six seconds of continuous, in-focus footage where one thing is clearly happening, framed loosely enough to crop for both 16:9 and 9:16, and with no burned-in graphics from a previous edit.

Things that disqualify a clip:

  • A cut inside it. Two shots joined is not one clip.
  • Someone speaking a sentence that only makes sense in context.
  • A logo, lower third or subtitle already burned into the picture.
  • Motion blur from a fast handheld pan, which looks fine at full length and terrible at two seconds.
  • Anything shot so tight that a vertical crop has nowhere to go.

Coverage per role matters more than total count. A library of 30 clips that is 25 talking-head shots and 5 product shots is not a 30 clip library, it is a 5 clip library with a lot of one thing. The roles we check depth against are hero product, product in use, hands, result or outcome, reaction, environment, and generic cutaway. Two or three per role gets you to a healthy 20 or so, and the gaps are usually in the same places: product in use, and the result.

Coverage beats polish, and it is not close

If you have a fixed budget and have to choose between one well-lit day and three scrappy ones, take the three scrappy ones. This is the single most consistent thing we see across accounts.

The reason is arithmetic rather than aesthetics. Polish improves how each clip looks by a margin most viewers in a feed will not register. Coverage increases the number of distinct ads you can build, and variation count is what predicts campaign longevity in the benchmark data: throughput, not talent, is the bottleneck it identifies.

There is a floor, and it is lower than production people like. Sharp focus, exposure that is not clipped, stable enough to watch, and clean audio if the clip needs to speak. A phone in daylight clears that floor. What matters after that is whether the shot shows something.

Where polish genuinely earns its cost: the hero product shot that closes the ad, and any shot with legible on-screen text. Those two are worth doing properly, and they are two clips, not two days.

The shots teams always come up short on

Product actually being used, by a person, with hands in frame. The result or outcome, filmed at the end when everyone has stopped caring. Multiple openings, meaning three or four different first shots for the same message. Nobody plans for the third one, and it is the one that decides whether your set repeats itself.

When to generate a missing clip instead of reshooting

Reshooting is the right call more often than tool vendors admit, and there is a clear band where generating is better.

Generate when the gap is short, abstract and non-human. A two or three second cutaway, a texture, a location you cannot access, an establishing shot. AI-generated video costs 2 credits per second against 1 credit for a whole variation, so a three second clip is 6 credits, and that only makes sense because the clip joins your library and gets reused across a dozen variations. Generate short, generate few, reuse constantly.

Reshoot when the gap involves hands, faces, several products in one frame, or your actual product doing its actual job. Generated footage is identifiable in exactly those cases, and hands are the first thing a viewer checks. A phone-shot clip of someone using the product genuinely beats a generated approximation of it.

Gap in the library Cheapest fix Why
Abstract 2 to 3 second cutaway Generate 4 to 6 credits, reusable, indistinguishable at that length
Hero product on plain background Animate an existing photo You already have the photography
Product in use with hands Shoot it, even on a phone Generation shows seams on hands
A customer speaking Film or record a real one We do not offer avatars, and they read as synthetic
Location or environment you cannot visit Generate Often the only option
A result or outcome shot Shoot it Needs to be your real product

Making the library usable once you have it

Volume without organisation is a folder nobody opens. The clips need to be findable by role, by product, by whether they contain speech and by whether the framing survives a vertical crop, otherwise every new ad reaches for whatever is at the top of the list.

Genyad is our product, so treat this as disclosure. Upload footage once and it transcribes and tags every clip into an indexed clip library, then each variation is built as a fresh script, shot selection, voiceover and caption set from that library rather than a re-cut of one timeline. Uploading footage, editing and re-exporting cost no credits, so growing the library is free and only the variations are billed at 1 credit each. Our conventions for tagging and structuring a library are in the organising your library documentation.

What it will not do: it cannot invent coverage you do not have. If the library has no product-in-use footage, the variations will all reach for the same hero shot, and no amount of scripting fixes that. It also has no AI avatars, no static banner formats, no product URL import, no product feed rendering, no predicted performance scores and no direct publishing to Meta, TikTok or Google Ads.

Frequently asked questions

How much footage do I need to make video ads?

Think in clips: around 30 usable clips is where variations stop repeating their openings, and 20 to 40 clips comfortably supports the 8 to 20 live variations a typical active campaign runs. Since an hour of long-form video yields roughly 10 to 25 usable clips, two or three hours of assorted material is a realistic starting library.

What makes a clip usable for an ad?

Two to six seconds of continuous, in-focus footage where one thing is clearly happening, framed loosely enough to crop for both wide and vertical, with no graphics burned in from a previous edit. Clips containing a cut, a sentence that needs context, or motion blur from a fast pan get discarded. Depth per role matters more than the raw count.

Is it better to shoot more footage or better footage?

More, in almost every case. Coverage decides how many distinct ads you can build, and variation count is what predicts campaign longevity, while polish improves something most viewers in a feed will not consciously notice. Spend the production quality on the hero product shot and anything with legible on-screen text, then prioritise coverage everywhere else.

Should I generate footage to fill gaps in my library?

For short abstract cutaways and locations you cannot shoot, yes. Generated video costs 2 credits per second against 1 credit for a full variation, so keep clips to 2 to 4 seconds and rely on reusing them across the set to justify the cost. For hands, faces, multiple products in frame or your product doing its real job, shoot it instead, even on a phone.