Several paths rising from one origin, one clearly outpacing the rest

Creative testing only survives a real budget when you separate the cheap decision from the expensive one. Cut bad hooks on about 2,000 impressions per variation, then spend real budget on survivors until you have roughly 15 conversions per variation before reading cost. That is the framework: a hook gate, then a cost gate.

Why most creative tests fail before the creative does

We used to run one big test, wait for statistical confidence, then scale the winner. By the time the result came back, the ad was dead. Our 2026 creative fatigue benchmark puts the CTR decline at 15 to 20 percent in a creative's first two weeks, then a drop to 45 to 70 percent below the launch baseline by week three. Most creative is effectively dead within three weeks.

The report's operational conclusion matters more than the decay numbers. Throughput, not talent, is the bottleneck. The number of unique creative concepts a team ships per month predicts campaign lifespan better than the quality of any single ad. Brands shipping 15 to 50 variants a month see 3 to 5 times longer campaign lifespan than quarterly refreshers. That means a creative testing framework has to support high-volume testing weekly, not once a quarter.

The two thresholds that actually matter

At the hook gate, spend about 2,000 impressions per variation. Hook rate is the share of impressions that get past the first few seconds. Typical ranges from accounts we run are 15 to 25 percent on Meta feed and 20 to 30 percent on TikTok in-feed. The point of 2,000 impressions is not statistical perfection. It is enough to separate a hook in the top quarter from a hook in the bottom quarter, and cheap enough to run for ten or fifteen variants.

At the cost gate, wait for roughly 15 conversions per variation. Cost per acquisition is a ratio. With fewer conversions, a couple of late-attributed purchases can change the read from pass to fail. At 15 conversions the cost read is still directional, but it is no longer just noise. If you cannot tell whether a difference is real, use the creative testing calculator before you kill or scale.

Gate Threshold per variation What it tells you What to do
Hook gate About 2,000 impressions Whether people get past the first few seconds Kill obvious losers, keep survivors
Cost gate Roughly 15 conversions Whether the cost per action survives a real denominator Kill expensive survivors, scale cheap ones
Fatigue refresh 8 to 20 live variations in rotation Whether the account has enough new creative Replace before the week three cliff
Test structure One broad ad set Whether creative differences are actually comparable Avoid twenty tiny ad sets

Why one ad set beats twenty

Twenty ad sets mean twenty different auctions, twenty different delivery histories and twenty different audience shapes. You are no longer testing creative. You are testing which tiny ad set happened to find the right pocket of users.

Put all variations in one broad ad set with one budget and one targeting signal. The platform can allocate spend and the hook-rate read becomes a measure of the creative, not a measure of the audience pockets. It also concentrates delivery enough to get every variation to the roughly 2,000-impression hook gate before budget leaks away.

Genyad does not publish directly to Meta or TikTok and does not predict performance scores. The hook gate is your scoring model. You still run the read in your ads manager.

The week-by-week schedule

Assume you start with a footage library and enough budget for a real test, but not enough to waste on dead hooks.

Week Hook gate Cost gate Live variation target
Week 1 Launch 10 to 15 variations in one broad ad set. Spend until each has roughly 2,000 impressions. No cost reads yet. 10 to 15
Week 2 Kill the variants below the hook-rate floor. Replace them with fresh hooks. Push the top three to five survivors towards conversion volume. 8 to 20
Week 3 Keep the hook gate running. Replace anything flat before the fatigue cliff hits. Read cost on survivors with roughly 15 conversions. Kill the expensive ones. 8 to 20
Week 4 Leave a small test cell for fresh hooks. Scale the two to four proven variants. Watch frequency and CTR decay. 8 to 20

By week four you have not found one winner and stopped. You have built a rotation. A typical active campaign needs 8 to 20 live variations in rotation. The fatigue benchmark says most creative is dead within three weeks, so weekly testing is not a project. It is the schedule.

When to stop testing and scale

Stop testing a specific variation when it has passed both gates. That means hook rate at or above the typical floor for the platform and roughly 15 conversions with a cost per action at or below target. Then scale it.

Scaling does not mean turning off every other test. Scale the proven variant and keep a test cell running for replacements. The fatigue benchmark puts the start of decline at a weekly frequency of about 2.5 on Meta prospecting. If you scale by hammering the same audience, you are buying frequency, not growth.

Also stop testing if you cannot fill the hook gate. If the team can only produce one or two videos a quarter, the framework has nothing to triage. Our creative testing workflow shows how to turn an existing footage library into a steady queue of variations.

On Genyad, one video ad variation costs one credit. The free plan gives you five variations, enough for a first small hook read. Growth is €99 for 65 credits, which is built for a weekly schedule like this. You upload footage once, it is transcribed and tagged, and each variation is a fresh script, shot selection, voiceover, caption set and export from that library rather than a re-cut of the same timeline. If you need AI avatars or synthetic presenters, Genyad does not do that. Use a tool built for avatar generation instead.

Frequently asked questions

How many variations should I launch each week?

Launch enough to keep 8 to 20 live variations in rotation after killing the weak hooks. That usually means creating 5 to 10 new variations a week, depending on your kill rate. The exact number matters less than keeping the hook gate full.

Why not wait for 10,000 impressions before cutting a hook?

Because the fatigue curve works against you. By the time every variation has 10,000 impressions, the early ones are often deep into week-two CTR decline. The roughly 2,000-impression threshold is a triage point, not a final verdict.

What should I use as a hook-rate floor?

Typical ranges from experience are 15 to 25 percent on Meta feed and 20 to 30 percent on TikTok in-feed. Cut below the lower end first. The floor moves by account and product, so use the creative testing calculator if the gap is small.

When should I stop testing and just scale?

Scale a variation once it has passed the hook gate and reached roughly 15 conversions with an acceptable cost per action. Keep the overall test cell running for replacements. Fatigue is the default, not the exception, and the creative testing glossary defines the terms we use.

Does Genyad publish directly to Meta or TikTok?

No. Genyad exports video files in 9:16, 4:5, 1:1 and 16:9. You download the exports and set up the ad set yourself. It also does not make static banner formats or predict performance scores, so the hook gate runs on platform data, not on a pre-trained scorer.