A one hour video timeline with a dozen short segments highlighted along it, each labelled with a clip role such as claim, demo, proof or cutaway

A long YouTube video is an ad library nobody indexed. One hour of recorded footage typically yields 10 to 25 usable clips, which is enough source material for a full prospecting set, and the work is finding them rather than filming anything new. The counter-intuitive part is which moments to take: the passage where you explained the product most clearly is usually the worst ad in the file, because an ad needs a claim and a lesson is not a claim.

Which moments in a long video are worth extracting?

Scan the transcript before you scan the video. Reading is faster than scrubbing, and the moments that matter are almost always identifiable from words alone.

Moment in the long video Why it works as an ad Ad role
An unhedged claim, said once, fast Already sounds like a person, not a script Hook or opening line
A number stated out loud Specific, survives being quoted Claim or proof
A comparison with the old way Frames the decision for the viewer Body, objection handling
A moment of visible use or demo Shows rather than describes Evidence beat
A customer or guest saying the benefit Third party voice beats first party voice Social proof
A reaction: surprise, laughter, a pause Reads as unstaged, holds attention Hook or transition
Hands doing something with the product Universally reusable, no lip sync Cutaway
Rejected takes, fluffed lines, setup shots Nobody looks like they are performing B-roll, often the best in the file

That last row is worth dwelling on. The takes that got cut from the finished video are frequently the most useful footage in the whole archive, because the subject was between performances. Behind-the-scenes rushes, the ten seconds before someone realised the camera was rolling, the wide shot nobody used: those are the clips that make an ad look like it was filmed rather than assembled.

What to leave behind: anything with a chapter transition in it, anything that references the video it came from, anything where the point takes more than one sentence to arrive, and anything where the speaker is off-frame or looking at a script.

Why the clearest explanation makes the worst ad

This is the mistake almost everyone makes on their first pass. You know exactly where the good bit is: minute fourteen, where you finally explained the product properly. Two minutes of clean, patient, complete explanation. It will not work as an ad, and the reason is structural.

Long-form teaching is built to be understood by someone who has already decided to listen. It sequences, it qualifies, it handles edge cases, it earns each point before making the next one. An ad speaks to someone who has not agreed to listen at all, and has to assert rather than build. The claim comes first, the proof comes second, and any sentence beginning with "so what that means is" belongs to a different medium.

In practice, the ad-usable version of a two minute explanation is a nine second fragment from the middle of it, usually the sentence where you stopped being careful. Look for the moment the speaker gets slightly impatient and says the thing plainly. That sentence is the ad.

The second half of the same mistake is length. Teams cut a 45 second segment because the whole thought is in there, then wonder why retention collapses at second twelve. The thought was not the problem. The setup attached to it was.

How many ads does an hour of footage support?

Ten to 25 usable clips per hour is the range we see across webinar recordings, podcast video, product walkthroughs and unedited shoot rushes. Where a given file lands depends on three things: how much of it is one static shot, how much of it is dead air, and how many people are visible.

Clip count matters more than hours, because variations are assembled from clips. Around 30 usable clips is the threshold where variations stop repeating openings. Below that, every new ad you build starts to reach for the same three strong shots, and the set begins to feel like one ad wearing different hats even when the scripts genuinely differ. Above it, the combinations open up and a set of a dozen ads can have a dozen distinct first seconds.

So the practical arithmetic for a first library:

  • One hour of long-form video: 10 to 25 clips, enough for perhaps six to eight variations with some repetition.
  • Two to three hours from different sources: comfortably past 30 clips, and openings stop colliding.
  • Add product-in-use footage separately, because long-form recordings are heavy on faces and light on product.

That last point is the reliable gap. Talking-head archives give you claims, proof and reactions, and almost no clean product footage, which is exactly the material an ad needs for its evidence beat. The repurpose existing footage page goes into how mixed libraries get balanced.

Volume has a direct payoff. Our benchmark reports that brands shipping 15 to 50 creative variants a month see 3 to 5 times longer campaign lifespan than teams refreshing quarterly, and that 8 to 20 live variations is what a typical active campaign needs in rotation. An indexed back catalogue is the cheapest route to that number, and our 2026 fatigue benchmark sets out why the number matters more than the polish.

How do you reframe 16:9 into 9:16 properly?

Not with a centre crop. A 16:9 frame cropped to 9:16 keeps the middle 56 percent of the width, which is fine for a single centred speaker and destructive for everything else: two-person interviews lose a person, a product held at the edge of frame disappears, and any lower-third graphic gets sliced.

Reframing properly means three things.

Reposition per shot, not per timeline. Each cut needs its own framing decision, following whoever or whatever matters in that shot.

Punch in where you have the pixels. A 4K source cropped to 1080x1920 leaves room to reframe without softening. A 1080p source does not, so scale sparingly and accept a slightly wider look rather than a mushy one.

Rebuild the text. Captions written for a wide frame are the wrong length and the wrong place for vertical. Vertical caption text belongs in the middle third of the frame, clear of the interface overlays at the bottom and down the right edge.

One honest limitation: some shots cannot be reframed. A wide two-shot with meaningful action at both edges is a 16:9 clip, full stop. Use it for in-stream and find something else for Shorts rather than mangling it.

Audio is the limiting factor, not video

Most repurposing projects fail on sound. Webinar audio recorded through a laptop microphone, a conference room with a hard ceiling, a podcast guest on a bad connection: the picture is usable and the audio marks the ad as amateur in the first second, which is exactly where you cannot afford it.

What we do about it, in order of preference:

  1. Use the original audio if it is clean. Nothing beats a real person saying a real sentence.
  2. Keep the clip, replace the audio. Use the moment as visual and carry the argument in a fresh voiceover, with the original sound dropped under it or muted.
  3. Use it as B-roll only. A clip with unusable audio is still a perfectly good cutaway, and cutaways under three seconds carry no dialogue anyway.
  4. Cut it. If the moment only works because of what was said and the recording is poor, it is not a clip.

Option two is the workhorse. It means audio quality decides whether a clip can be a hook, not whether it can be used at all, and it is why a badly recorded archive is still worth indexing.

Indexing the archive

Genyad is our product, so read this as disclosure. It exists because of exactly this problem: footage that is already paid for and unusable at speed because nothing knows what is in it. Upload the long files once and it transcribes and tags every clip into an indexed clip library, then each variation is built as a fresh script, shot selection, voiceover and caption set drawn from that library rather than a re-cut of one timeline. One variation costs one credit, uploading footage and re-exporting cost nothing, and exports cover 9:16, 4:5, 1:1 and 16:9 with framing handled per ratio. There is a fuller walkthrough on the long video to short ads page.

Where it stops: it does not import from a YouTube URL, so you upload the source files yourself. It does not publish to any ad platform. There are no AI avatars, no static banners and no predicted performance scores. And it cannot rescue audio that is genuinely broken, only work around it by replacing the voice.

Frequently asked questions

How many ads can I get from one hour of video?

Expect 10 to 25 usable clips from an hour of long-form footage, which supports roughly six to eight variations before openings start repeating. Getting past around 30 usable clips, usually by combining two or three sources, is where a set of a dozen ads can have a dozen genuinely different first seconds. Talking-head sources will still be short of clean product footage.

Which part of a long video makes the best ad?

Not the part where you explained things well. Look for the sentence where the speaker stopped being careful and said the claim plainly, usually a few seconds long, often surrounded by material you will discard. Reactions, stated numbers and moments of visible product use are the other reliably useful fragments.

Can I crop a 16:9 video into a vertical ad?

Only by reframing shot by shot, not by cropping the centre of the timeline. A centre crop keeps about 56 percent of the width, which loses one person in an interview and any product held at the edge of the frame. Punch in only as far as your source resolution allows, and rewrite caption text for the vertical frame rather than reusing wide captions.

What if the audio on my old footage is bad?

Keep the clip and replace the audio. Carry the argument in a fresh voiceover with the original sound muted or ducked underneath, which turns a badly recorded moment into usable visual material. Clips with unusable sound still work as cutaways under three seconds, where there is no dialogue to hear anyway.