Transcription and clip search
Every clip is transcribed with word-level timing and given a short description of what is visible in it. Those two fields are what let a script be written against your actual footage, and the word timings are what make burned-in captions land on the right frame.
Last updated
What is stored per clip
| Field | Used for |
|---|---|
| Transcript with word timings | Script writing, caption sync, clip search |
| Visual description | Shot selection from a plain-language request |
| Role (clip type) | Which slot in the sequence the clip can fill |
| Scope (product or campaign) | Which campaigns may use it |
| Duration and resolution | Caption sizing, export framing, timing |
Why this matters for output quality
A generator with no view of your footage writes a plausible ad for your category and then needs an edit to fake the shots it assumed. Writing against the index removes that failure mode: if there is no shot of the product being unboxed, no script asks for one.
Frequently asked questions
Can I search my footage by what is said in it?
Yes. Clips are transcribed on upload and the transcript is used both for search and for shot selection during script writing.
How accurate is the transcription?
Good enough for selection and captioning on clear audio. Noisy or heavily accented recordings produce weaker transcripts, which is why audio quality matters more than image quality.
Are captions generated from the transcript or the voiceover?
From the generated voiceover's own word-level timings, so they match the read in the finished ad exactly.