
Prompt-to-video fails in six specific, predictable ways: hands, faces, product accuracy, on-screen text, continuity between shots, and quality decay as the clip gets longer. None of these are prompt problems, so no amount of prompt engineering fixes them, and all six are worth knowing by name because they map directly onto which shots you can safely generate for an ad and which you cannot. What it does well is atmosphere: exteriors, weather, water, abstract texture, and anything peripheral where the viewer is not being asked to verify a fact.
The six failure modes, and why each one happens
| Failure mode | Why it happens | What to do instead |
|---|---|---|
| Hands and fine manipulation | The model learned a distribution over hand-shaped pixels, not a five-fingered skeleton with joint limits. Contact with an object it is also inventing has no physical constraint to satisfy | Film hands. It is two minutes of phone footage |
| Faces, speech and eye line | Mouth shapes are generated as plausible texture, not aligned to phonemes, and human face perception is unusually sensitive. Errors of a few percent read as wrongness | Use real people, or keep faces small and distant, or cut before anyone speaks |
| Product accuracy | The model has never seen your SKU. It produces a category-average object, so label typography, proportions, seams and materials come out approximately right and specifically wrong | Image-to-video from a real photo, covered below |
| On-screen text and glyphs | Text is generated as visual texture in the same latent space as everything else. There is no character-level representation to be correct about | Add all text in the edit, as a caption or overlay |
| Continuity and identity drift | Each generation is an independent sample with no persistent scene state. Wardrobe, lighting, set dressing and faces change between two clips of the same nominal subject | Treat every generated clip as a separate location. Never cut two together as one scene |
| Length and temporal decay | The model conditions on a limited temporal window, so error accumulates. Motion slows, subjects morph, physics that looked right at one second is wrong by four | Generate two to three seconds. Cut, do not extend |
A few notes on the ones people argue about.
Hands are the canonical example and the reason is worth understanding, because it generalises. A generative model is not simulating a hand, it is producing something whose statistics match hands it has seen. Most of those references are ambiguous, partly occluded and mid-motion, so the learned prior covers a huge space of hand-like arrangements, and finger count is not a hard constraint anywhere in that space. Add a product to the grip and it gets worse, because the model is now inventing both objects and the contact between them with nothing enforcing that they occupy separate volumes.
Faces fail for a different reason: not that the model is worse at faces than at trees, but that you are much better at faces than at trees. A slightly wrong branch is a branch. A slightly wrong blink is a person who is not real. For direct response this is expensive, because trust is the mechanism the ad runs on.
Product accuracy is the one that matters most for advertising and the one vendors are quietest about. There is no path by which a model that has not been trained on your product can render your product. It will render the idea of your product. If your packaging has printed copy, both failure modes stack.
On-screen text deserves a hard rule rather than a caution. Short words on a plain background sometimes come out clean, and that occasional success is what makes people keep trying. Put your price, your offer or your legal disclaimer in a generated frame and you will eventually ship something misspelled. Add text in the edit, where it is deterministic.
Our glossary entry on how generative video models work has the mechanism in more detail, and every item above follows from it.
What prompt-to-video does reliably
The list is shorter than the marketing suggests and genuinely useful.
Exteriors and terrain: coastlines, mountains, deserts, forests, roads, cities at dawn. Weather and elements: rain on a window, snow, fog, waves, fire, smoke, steam. Abstract and macro texture: fabric weave, liquid in motion, gradients, particles, ink in water. Slow camera moves over static scenes, especially drone-style aerials, where there is no articulated motion to get wrong. Ambient people at a distance, where faces are too small to trigger anything. And background plates that will sit behind captions or an end card, where attention is on the text anyway.
The common property: nothing in frame is a claim. Nobody watches an establishing shot of a coastline and asks whether that coastline exists. That is exactly the shot to generate, and it is normally the shot missing from a footage library, which is why we treat generation as gap-filling rather than as production.
Runway remains the strongest general-purpose tool in this category if raw generation quality is what you need, listed at $15 a month in August 2026, and prices move so check before committing. We are not a substitute for it in that job, and we have written up where the two categories separate in our notes on choosing between Runway and a footage-first tool.
Image-to-video is the actual fix for product accuracy
If you need your product in a generated shot, do not describe it. Start from a photograph of it.
Image-to-video conditions the generation on a real first frame, so the model animates what is already there rather than inventing it. Your label stays your label, the proportions stay correct, the colour is the colour you shipped. This turns product shots from unusable into sometimes usable, which is a real change.
It is not a full fix, and the limits are predictable from the same mechanism.
- Keep the motion small. A gentle push in, a slow orbit of a few degrees, a parallax drift. Ask for a 180 degree rotation and the model has to invent the back of the product, which puts you straight back into the accuracy problem.
- Keep it short. Identity holds for two or three seconds and then drifts, because each frame is conditioned on the last rather than on your original photo.
- Avoid revealing text. If a camera move brings packaging copy from unreadable into readable, the model will produce glyph-shaped noise where the words should be.
- Use a clean, well-lit source frame. Artefacts in the input become animated artefacts in the output.
- Expect to discard some attempts. Budget two or three generations per usable clip and price it accordingly.
The practical version of this is in our image-to-video documentation. Used within those limits it covers a genuine gap: a product shot in a setting you cannot afford to travel to, or an angle you did not get on the shoot day.
Length is a quality cliff, not a slope
This is the least intuitive limitation and the most expensive one to ignore.
Two to three seconds is the reliable zone. Motion is coherent, subjects hold together, physics looks right. Four to six seconds is usable with review, and you should expect to reject some. Beyond that, error accumulation shows up as specific artefacts: subjects morph into adjacent objects, motion drifts into slow motion, limbs and props detach, and liquid or fabric that was convincing at second one behaves impossibly by second four.
The cost structure makes this worse, since generation is billed by the second across the market. With us it is 2 credits per generated second, so a 10 second attempt is 20 credits against 6 for a 3 second one, and the 10 second attempt is the one more likely to be unusable. Paying more for a higher failure rate is a bad trade in both directions.
What works instead is cutting rather than extending. Three separate 3 second generations, edited together with real footage between them, gives you nine seconds of screen time at 18 credits with three independent chances of a good clip. One 10 second generation gives you one chance at 20 credits. Short generations also sit better inside an ad, because peripheral shots under two seconds are where generated footage goes unnoticed anyway.
Disclosure: Genyad is our product, and it is footage-first. Generation is a gap-filler alongside assembly from your own library, not the main event. We do not offer AI avatars or synthetic presenters, and given the face failure mode above we are not planning to. We also have no static banner formats, no product-URL import, no product-feed or CSV rendering, no predicted performance scores and no direct publishing to Meta or TikTok.
Frequently asked questions
Can better prompting fix AI video artefacts?
No, for the six failure modes above. Prompting controls what the model attempts, not what it is capable of representing, so no phrasing gives a model character-level text or a correct finger count. Prompting does help with framing, pacing, lighting and mood, which is worth doing well.
Which prompt-to-video failures matter most for ads?
Product accuracy and on-screen text, because both can make an ad factually misleading rather than merely ugly. A generated version of your packaging misrepresents what a customer receives, and garbled text next to a price is a compliance problem. Hands and faces come next, since they break viewer trust in the first second.
Is image-to-video better than text-to-video for advertising?
For any shot containing your product, yes, and by a wide margin. Conditioning on a real photograph keeps the label, geometry and colour correct, which text prompting cannot do. Keep the camera move small and the clip under about three seconds and it holds up.
How long can a generated clip be before quality drops?
Two to three seconds is reliable, four to six needs review, and beyond that most models show morphing, drifting motion and physics errors. Since generation is billed per second, longer attempts cost more and fail more often, so cutting several short generations together is the better use of the budget.