A dense lattice resolving into a few clean deliberate shapes

Ads that AI assembles from your own real footage perform like ads: some win, most do not, and the distribution looks the same as a human editor's. Ads built entirely from generated footage generally underperform on direct response, for a specific reason that has nothing to do with model quality. The product on screen is not your product, and the moment a viewer has to evaluate a purchase, that gap starts costing you conversions rather than views.

Assembled versus wholly generated

These two things get filed under the same headline and they are not the same product.

Assembly means a model writes the script, chooses which of your existing clips to use and in what order, generates the voiceover, sets the captions and exports the ratios. Every frame is real footage of the real product. Nothing about the visual truth of the ad has changed, only who did the editing. So the performance question collapses into an ordinary creative question: is this a good script over the right shots, and there is no reason a machine-assembled version should behave differently from a human-assembled one. In practice it behaves the same, and it exists in far greater quantity, which matters because our 2026 fatigue benchmark links brands shipping 15 to 50 variants a month to 3 to 5 times longer campaign lifespan than quarterly refreshers.

Wholly generated means the footage itself came out of a text-to-video or image-to-video model. Now the visual truth has changed, and three things follow.

The product is wrong in ways viewers cannot articulate but do notice. Label typography drifts, a seam sits in the wrong place, the bottle cap has one thread too many. Nobody consciously spots it. They just trust the ad slightly less, and trust is the mechanism direct response runs on.

Proof disappears. Most performance creative works by demonstration: the thing being used, the before and after, the texture of the fabric. A generated demonstration proves nothing, because it is a picture of a demonstration that never happened. You can generate a plausible shot of someone enjoying a product. You cannot generate evidence.

Hook rate and conversion rate come apart. Generated footage is often striking, so it can win the first two seconds. Then the click quality is worse and you see it in cost per acquisition rather than in CTR. If you evaluate only top-of-funnel metrics, wholly generated creative looks better than it is.

Our glossary entry on generative video covers what the models are actually doing, which explains most of the failure list below.

What we see in keep rates

We look at what survives human review, since that is the earliest honest signal. These are our own observations across accounts rather than a study, so treat the numbers as directional.

Assembled variations from a decently tagged library: most accounts keep somewhere around seven or eight in ten. Rejections are usually script problems, a clip that is technically fine but off-brand, or two variations landing too close to each other.

Generated clips used as peripheral b-roll, kept under about two seconds, with no product and no hands: keep rates look similar to real b-roll, roughly six or seven in ten, with the misses being tone rather than artefacts.

Generated clips where the product is the subject: keep rates fall off a cliff, closer to two or three in ten, and the ones that get kept tend to be products with simple geometry and no printed text. Anything with a label, a screen or a logo is usually a reshoot.

That spread is the whole story. Generated footage is a supply of texture and atmosphere, not a supply of product shots.

Where generated footage is genuinely invisible, and where it fails

Shot type Generated footage verdict Why
Establishing exterior, city, road, coastline Reliable No product, no people in close-up, viewers have no reference to check against
Sky, weather, water, smoke, fire Reliable Motion models handle fluid and particle behaviour well and there is nothing to get factually wrong
Abstract texture, macro fabric, gradient motion Reliable Used behind text or as a transition, it reads as design rather than as documentation
Background plates behind captions or an end card Reliable Attention is on the text, and any short peripheral shot survives that
Ambient people at distance, crowd, street Usually fine Faces below a certain size do not trigger the uncanny response
Hands using the product Fails Hands are the classic failure: finger count, joint direction, and grip contact with an object the model is also inventing
Face to camera, speaking Fails for direct response Lip sync, eye line and micro-expression are where viewers are most sensitive, and the result reads as an advert about nothing
Product hero shot or close-up Fails Logo, label typography, proportions, materials and colour will all be approximately right and specifically wrong
Anything with on-screen text, packaging copy or a UI Fails Models produce text-shaped glyphs. Legible, correct copy is not something to rely on
Continuity between two generated shots Fails Each generation is independent, so wardrobe, lighting and set change between cuts
Food and drink being poured, eaten or cut Mixed, leans fail Physics looks plausible in a one second cut and wrong by three seconds

The practical rule we work to: generated footage belongs in shots where the viewer is not being asked to believe anything. Short, peripheral, product-free, under two seconds. Used that way it fills the gaps in a thin library, which is what an AI b-roll generator is actually for. Used as your hero shot, it undermines the only thing the ad has to do.

Where a whole generated ad can work: brand or awareness campaigns with no product demonstration, category-level messaging, abstract or service businesses with nothing physical to show, and internal concept tests where you want to see a storyboard move before committing to a shoot. Those are real use cases. They are not the same as a purchase ad.

How to test this in your own account

Do not take our word for it, and do not take a vendor's. The test is cheap.

Build three cells with the same offer, same audience, same landing page, same length, same voiceover, same captions. Cell A is entirely your own footage. Cell B is your footage with two or three generated peripheral shots substituted for shots you do not have, none of them showing the product. Cell C is entirely generated. Give each cell the same budget and run until each has enough volume to say anything, which for most accounts means low hundreds of clicks per cell at minimum, and at least seven days.

Then read it in this order. Hook rate first, expecting Cell C to look competitive or better. Then CTR: our benchmark puts platform medians at 1.78 percent video CTR on Meta Reels, 1.62 percent on Meta feed, 0.84 percent on TikTok and 0.42 percent on YouTube in-stream, so you have a reference point. Then cost per acquisition, which is where the gap normally appears. Then, if you can, run a 30 day check on return rate and refund rate, because an ad that oversold a product it never showed accurately produces expensive customers.

One warning about study design. Do not compare "AI ads" against "our ads" as a blanket test, because you will be comparing four variables at once and learning nothing. Hold everything constant except footage provenance. If you want the fastest version of this, our creative testing workflow has the cadence.

Disclosure: Genyad is our product, and it does both. Assembly from your own library costs 1 credit per variation. Generated video is available as a separate feature at 2 credits per generated second, and we price it higher because it costs us more compute, not because it is better. We have no AI avatars or synthetic presenters, and we are not going to add them, because the failure mode above is not a bug we expect to fix. If you want the mechanics of the generation side, see how our AI video generation works.

Frequently asked questions

Do AI-generated ads get rejected by Meta or TikTok more often?

Not in our experience, as long as the creative follows the same policies as anything else. The rejections we see are ordinary: unsupportable claims, before-and-after imagery in restricted categories, and text problems. Generated footage does not attract extra scrutiny, but a generated product shot that misrepresents what you ship is a genuine policy and consumer-protection risk.

Is a fully AI-generated video ad ever the right choice?

For brand and category messaging with nothing physical to demonstrate, yes. For concept testing before a shoot, yes, and it is a cheap way to see whether an idea holds. For a conversion ad where the product needs to be believed, we would not run one, and the reason is that the product on screen is not the product you sell.

How much generated footage can I mix into a real ad before it hurts?

The limit is not a percentage, it is what the generated shots are doing. Two or three short peripheral cuts with no product and no hands go unnoticed. One generated hero shot of the product can undermine the whole ad, even if it is only 8 percent of the runtime.

Does AI-assembled creative beat human-edited creative?

On a per-ad basis, no, and we would not claim it does. The advantage is throughput: our fatigue benchmark puts most creative effectively dead within three weeks and a typical active campaign's need at 8 to 20 live variations, and assembly makes that cadence reachable. More shots at the target beats a better single shot when the target keeps moving.