Interlocking modules whose edges meet imperfectly, leaving visible gaps

To choose a video ad tool, give two or three candidates the same real brief, make each produce five outputs, and count how many of those five you would publish without touching them. That count, the keep rate, plus the elapsed minutes from login to five finished exports, will separate the tools more reliably than any feature comparison, and you can finish the whole exercise in an afternoon. We make one of the tools you might test (Genyad, which builds ads from footage you already own), and the protocol below is deliberately one we lose sometimes.

The five-step protocol

Step 1: write one brief, and make it real

Not a demo brief. Take a campaign you are actually shipping this month and write down eight things: the product, the audience, the single claim you want to lead with, the offer, the source assets you have, the placement, the aspect ratio and the duration. One page, fifteen minutes.

Every tool gets that identical brief, verbatim. The moment you start tailoring the brief to each tool's strengths, you are testing your own skill at prompting rather than the tool, and you will pick the one whose interface you happened to warm to first.

Step 2: pick candidates from different categories, not the same one

Three tools from three categories tells you something. Three prompt-to-video tools tells you almost nothing, because they will fail and succeed in the same places. Choose by what you already own: a template renderer if you have a locked brand template and structured data, an avatar tool if you own no footage, a library-driven tool if you have hours of footage and no concept. Our comparison hub sets out which tool sits in which category, including where we recommend someone else.

Two candidates is enough if your inputs are clear. Three is the practical maximum for one afternoon.

Step 3: produce five outputs each, and keep all five

Five, not one. One output tells you the tool's ceiling, which is the number every demo shows you. Five tells you its floor and its variance, which is what you will live with.

The discipline that makes this test worth running: no cherry-picking, no regenerating a bad one before you have scored it, no second prompt attempt. Start a timer at login and stop it when the fifth file is on your disk.

Step 4: score keep rate, ideally blind

Put all fifteen outputs in one folder with neutral filenames and score them the next morning, or have a colleague score them. For each, one of three verdicts: publish untouched, publish after one small edit, or discard.

Keep rate is the count of "publish untouched" divided by five. Score it against what you would genuinely put behind spend on a client or your own account, not against what is impressive for a machine. The second category matters too, but count it separately, because "one small edit" times twenty ads a month is a job.

Step 5: convert to cost per kept ad

The list price is not the price. If a tool costs five credits to produce five outputs and you keep two, you paid two and a half credits per usable ad. A cheaper tool with a keep rate of one in five is more expensive than a dearer tool that keeps three. Run the same arithmetic on your own time: our video ad cost calculator is built for this comparison, and our own per-output rates are on the Genyad pricing page.

What to record on the scoring sheet

What you record How to capture it What a good result looks like What a bad result is telling you
Minutes, login to five exports Stopwatch, one sitting Tens of minutes for an assembly tool, hours for a generator you edit yourself You bought a faster hand tool, not throughput
Keep rate, untouched Blind score out of five Two or three of five is a strong result on a first pass The tool's floor is below your publish bar
Keep rate after one small edit Scored separately Most of the remainder Fixes are structural, not cosmetic
Cost of a revision Change one script line and watch the bill Free, or one unit of output Every client note costs money
Cost per kept ad Total spend divided by untouched keepers Compare across tools, never against list price Cheap plan, expensive output
Ratios delivered per run Count the files All the placements you buy, in one pass Manual re-export multiplies everything
Language handling Ask for one non-English variant A script written for that market A translated English script, which reads like one

The two rows people skip are revision cost and ratios per run, and they are the rows that decide whether month two feels different from month one. A tool that produces one excellent 9:16 file and makes you re-export 4:5, 1:1 and 16:9 by hand has quietly handed you back the work you were buying your way out of.

Why keep rate beats every other number

Because it is the only number that already contains everything else: script quality, shot selection, pacing, caption placement, brand fit and your own taste. A tool with a keep rate of three in five and a mediocre interface will out-produce a beautiful tool that keeps one.

It also sets the honest volume expectation. Most accounts need 8 to 20 live variations per active campaign, and our 2026 fatigue benchmark puts week three at 45 to 70 percent below the launch baseline, with a creative averaging 38 percent below peak by week five. If you need twelve live ads and your keep rate is two in five, you need thirty outputs a month, and that is the number to price, not twelve. Work it out before you sign, not in week three.

Why feature checklists mislead

Feature parity is now close to universal at the checkbox level. Captions, vertical export, voice options, brand kits: everything has them, so a checklist scores everything equally and then you pick on price, which is how teams end up with a tool nobody opens after six weeks.

Three specific ways the checklist lies:

  • It cannot see variance. Two tools both tick "AI script writing". One writes something you would run, one writes something that mentions your product twice and says nothing. Same tick.
  • It counts features you will never use. Predicted performance scores demo well and change no decisions, because the model is guessing at your account. We do not ship one, and we would not weight it if we did.
  • It ignores the seams. Where the work actually accumulates is between features: a caption that needs repositioning per ratio, an export that needs a manual re-crop, a revision that costs a full regeneration.

Ask instead for one number no vendor publishes: of the last ten outputs your tool produced for an account like mine, how many went live untouched. The answer, or the refusal, is informative.

When the test is already decided

Do not spend the afternoon if your inputs make the answer obvious. If you have a signed-off template and a product feed with 300 rows, buy a template renderer and stop reading. If the brand is three weeks old and owns no footage, an avatar or generative tool is your only option, and we cannot help you: Genyad has no AI avatars, no product-URL import and no product-feed rendering. If you have footage and no concept, the test is worth running, and we would like to be in it.

Frequently asked questions

What is a good keep rate for a video ad tool?

Two or three of five outputs publishable untouched is a strong first-pass result, in our experience, and one in five is workable if the outputs are cheap enough per unit. What matters more than the absolute figure is comparing it across tools on the same brief and the same assets, since your publish bar is your own. Score the "publish after one small edit" bucket separately, because those edits are recurring work.

How long should evaluating a video ad tool take?

One afternoon for the generation and one morning for blind scoring. If a tool needs a sales call, an onboarding session and a two-week implementation before it produces a first output, that is itself a result worth recording, particularly for a small team. Free tiers make the test nearly free: ours gives 5 variations and one AI-generated video with no card required.

Should I choose based on price?

Choose on cost per kept ad, which is list price divided by keep rate, plus your own minutes. Entry prices among the tools in this category ranged from free to about $99 per month as publicly listed in August 2026 and they move often, so anchor on output cost rather than a headline monthly figure. A subscription is also a worse fit than per-output pricing if your monthly volume swings.

Do I need to test with my own footage?

Yes, and it is the single most common reason a trial misleads. Stock footage flatters every tool; your actual library, with its bad audio, missing close-ups and inconsistent grade, is what the tool will have to work with in production. If a vendor's demo only works on their assets, you have learned something useful.