One thirty-second vertical, specified from the brief to the upload, with every prompt written out and the bill at the end. Then the reference behind each decision: which model does what, the prompt formulas, the consistency mechanisms, the export specifications. All prices are the providers’ own published figures as of September 2026.
The economics of short-form as a business are on making video with AI. The short version, because it changes how much of this is worth doing: view revenue on Shorts is roughly $0.03 to $0.10 per thousand, so a million views is about $50. Short-form pays as client work and as a route into something you own. Not as ad revenue.
One clip, all the way through#
Everything below this section is reference material. This part is the actual job, so it comes first. The brief is real client work of the kind that pays: a thirty-second vertical for a small coffee roaster, to run as paid social, delivered with three hook variants.
Nothing here is a case study. These are the prompts you would send and the decisions you would make, with the failures you should expect named against each shot. Where a number appears it comes from the published prices further down.
The shot list, which is where the money is decided#
Five shots, thirty seconds. The third column is the one that matters, and every mark in it has a reason rather than a preference:
| Shot | Film or generate | Why | |
|---|---|---|---|
| 0–5s | The bag, label readable | Generate, composite the label | Legible type in a generated shot garbles, worst of all on anything moving |
| 5–11s | Hands tamping and pouring | Film | Fine motor action morphs; fingers merge and tools pass through objects |
| 11–17s | Inside the grinder, beans falling | Generate | Impossible to film. This is what generation is for |
| 17–24s | The roaster says one line | Film | A recurring real face. Generate it only if they cannot be on camera, and then as a bound element |
| 24–30s | Aerial over a hillside farm at dawn | Generate | Cannot travel there. Nothing is being replaced |
Two of five shots are generated. That is the normal ratio for commercial work and it is the single most useful thing on this page: the AI is doing the shots you could not otherwise have, and the day still goes on filming and cutting.
The two generated shots, in full#
Shot 3, the grinder interior. Settle it as a still first, because a still costs half a cent and the clip costs twenty-seven:
Extreme close-up inside a coffee grinder, dark roasted beans mid-fall against steel burrs, single hard top light, fine dust in the air, shallow depth of field, warm shadows, photographic.
Run a batch of five, keep one or two. Then animate the keeper:
Extreme close-up, dark roasted beans falling between steel burrs, tumbling and catching the light as they drop, single hard light from above, fine dust drifting through the beam, shallow depth of field, warm shadows. SFX: a low mechanical grind. Do not generate any background music.
Expect the beans to misbehave. Small rigid objects colliding is the weakest thing these models do, so some attempts will have beans passing through each other or hanging in the air. Run three, take the one where the fall reads, and use the middle of it rather than the whole clip.
One habit worth adopting on every prompt, from a maker who cut a three-minute short: build in a second of dead air at each end. He writes “a 1-second buffer to the start and end of every prompt… to give myself handles for clean cuts in post”, filled with something inert like a pause or a slow look. Generated clips drift at the edges anyway, and handles are what let you cut on the frame you want instead of the frame you were given.
Shot 5, the aerial. Note how the exclusion is phrased, because there is no negative-prompt parameter on every model and positive phrasing works on all of them:
Aerial view over a steep hillside coffee farm at dawn, rows of green shrubs following the contour, low mist in the valley, warm rim light on the ridge, photographic, an unbroken hillside with no buildings or roads.
Then the move:
Aerial tracking shot over a steep hillside coffee farm at dawn, camera moving slowly forward along the contour of the rows, low mist sitting in the valley below, warm rim light along the ridge, cold blue in the shadows. SFX: faint wind. Do not generate any background music.
Expect a road or a roof to appear anyway, and expect the rows to dissolve into texture at the far end of the frame. Both are fixed by cropping tighter or cutting earlier, not by re-prompting harder.
Where the cuts go#
Three joins, and each one is doing a job:
- Filmed hands into the generated grinder. Cut on the motion of the pour, so the angle changes at the same instant the material changes. A cut on a still frame invites the comparison.
- The roaster’s line into the aerial. Cut on the last word, not after it. The half-second of silence after a line is where attention leaves.
- One grade across the whole timeline, filmed and generated together, then light grain over everything. This makes them sit together. It does not repair a bad take, and a working colourist’s view is that these artefacts are “far too obvious to be fixed by simply adding film grain.”
Then the three variants#
The client bought hooks, not one film. Shots 2 to 5 stay as they are; only shot 1 changes, and the caption with it. That is the whole reason this format pays: the expensive work is done once and the variants are an afternoon.
The bill#
Two generated shots, a batch of five stills each, three video attempts each, plus one extra shot’s worth of material because some clips will only yield a second or two:
- Ten stills in Higgsfield Soul 2.0: 1.2 credits, about $0.05
- The same ten in Nano Banana Pro: 20 credits, about $0.78
- Nine video generations of six seconds on Kling 3.0 at 720p: 76 credits, about $2.96
- The identical job on Seedance 2.0 at 1080p: 486 credits, about $19.04
So three dollars to twenty for the whole advert, and the choice between those two numbers is which model handles falling beans and a receding hillside acceptably. Not a rounding error on a client invoice, and not a cost worth optimising either. What you are selling is the shot list and the judgement to bin the bad takes.
Before you upload#
One judgement worth making deliberately on this particular job. YouTube requires the altered-content label for realistic material that misrepresents a real place. A photorealistic aerial presented as this roaster’s own farm is exactly that, and needs the toggle. The same shot used as generic imagery of coffee growing probably does not. The distinction is what the advert claims, not how the footage was made, and it is the client’s exposure as much as yours, so put it in writing when you deliver.
Everything else on this page is the reference behind those decisions: which model to choose, the prompt formulas, the consistency mechanisms, and the export specifications.
Pick the model by what it gives you per generation#
This is the decision that shapes everything downstream, and the field moved during 2026. Clip length and reference handling differ enough between models that they are not interchangeable.
| Seedance 2.5 | Kling 3.0 Omni | Veo 3.1 | Sora 2 | |
|---|---|---|---|---|
| Per generation | 4–30 sec | up to 15 sec | 4, 6 or 8 sec | 16 or 20 sec |
| References | up to 50 inputs: images, video, audio | named elements, built from a 3–8 sec video | up to 3 images | one first frame; non-human subjects only |
| Several shots in one generation | — | yes, with a duration per shot | timestamp prompting | — |
| Voice bound to a character | — | yes | — | — |
| Negative prompt | — | yes, defaults to “blur, distort, and low quality” | not an API parameter | — |
| Aspect ratios | 21:9 to 9:16, seven options | set by the start image | 16:9 or 9:16 only | includes 1080×1920 |
| Extend | video reference edits and extends | — | yes, 720p only | 6 × 20 sec, 120 sec total |
| Price per second | $0.13–0.47 | $0.112–0.196 | $0.05–0.60 | $0.10–0.70 |
Published prices and documented schemas, September 2026: Google, OpenAI, and the fal.ai listings for Seedance and Kling. Seedance’s cheaper rate applies when you supply video references; Kling’s rises with audio and again with voice control.
How to read it. Kling 3.0 Omni is the one to reach for when a person or product has to recur, because its consistency mechanism is a named asset rather than a prompt you retype. Seedance 2.5 takes the most reference material of anything here and generates the longest single clip. Veo gives you short takes with first-and-last-frame control. Sora gives long takes and the only chained extension path.
There are more than these. Wan 3.0 Prime, Hailuo 2.3, LTX 2.5, Grok Video and Higgsfield’s own DoP models all sit in the same market at similar prices, and MiniMax H3 Max has a Director mode that streams a continuously directed session rather than producing clips. The four above are enough to choose between for short-form.
How the shot list changes with the model#
The walkthrough assumed six-second shots. That is a Veo-shaped plan. The same thirty seconds on Sora is two takes of 16 and 20 with one trimmed, and on Seedance it can be a single thirty-second take with a timeline in the prompt. Decide which before you spend anything, because the three routes cost differently and look different.
The film-or-generate column is where the money is decided, and the marks are not preferences. Legible text, hands performing the task, a recurring human face, water and fabric are documented failure modes, listed with their fixes in telling the difference from slop. Everything else is a candidate.
Writing a prompt#
All three providers publish a formula and they are the same formula with the slots in a slightly different order, so learn one and stop worrying about it. Google’s for Veo:
[Cinematography] + [Subject] + [Action] + [Context] + [Style and ambiance]
Google’s own example:
Medium shot, a tired corporate worker, rubbing his temples in exhaustion, in front of a bulky 1980s computer in a cluttered office late at night. The scene is lit by the harsh fluorescent overhead lights and the green glow of the monochrome monitor. Retro aesthetic, shot as if on 1980s color film, slightly grainy.
OpenAI’s for Sora is the same list in fewer words: shot type, subject, action, setting, lighting. Kling’s puts the camera later, as Subject + Action + Scene + Camera + Lighting and mood. Any of the three orderings works on any of the models, so pick one and stop reconciling them.
The rule underneath all three formulas is the same and worth stating on its own: describe what is visible, not what it means. Kling’s illustration is that “magic” gives the model nothing, while “swirling blue energy particles with an ethereal glow” gives it a target.
The formula is the skeleton. What makes a prompt work is a set of rules the makers publishing their own results agree on, and several of them cut against instinct.
Shorter beats longer. One Seedance user who tested this deliberately found that “a 70-word prompt consistently outperformed a structurally identical 200-word version”, and settled on “50–80 words, structured as subject + action in sentence 1, camera + style in sentence 2, constraints in sentence 3”. Google’s own example above is longer than that, which tells you the vendors demonstrate capability while users optimise for hit rate.
One camera move. “Slow dolly in, slightly handheld” works. “Dolly in while panning left” does not. Two moves in one instruction and you get neither.
Name the lens in millimetres. Writing 24mm, 35mm or 50mm makes the model commit to a focal length; leaving it out means it “defaults to a flat mid-range look across the whole sequence”.
Replace adverbs with physics. The word “fast” reportedly degrades output. What works is describing the mechanics: “feet striking hard, each stride full extension, arms pumping at 90 degrees”.
Drop “cinematic”. It means nothing to the model. A specific reference does: “Wes Anderson symmetry”, “Kubrick one-point perspective”, “Golden hour backlight, long shadows stretching forward”.
When you are animating a still, stop describing the scene. The reference already carries the look. Keep the prompt to “exactly two things — motion instructions and camera instructions”. This is the single most common waste in an image-to-video prompt.
Repeat the constraints at the end. Two makers working on long single takes arrived at this independently: one restates the hairstyle, wardrobe and no-cuts requirement in a final paragraph, the other closes with a block headed “KEY PHYSICAL RULES (MANDATORY)”. The model weights the end of the prompt more than the middle.
And where there is no negative-prompt parameter, write the negatives as positive states. Instead of a list like “jitter, bent limbs, flicker”, one user now writes: “Face stable, Limbs anatomically natural, Consistent lighting, no flicker, Body proportions consistent throughout.”
The vocabulary the models parse is ordinary film language. Movement: dolly, tracking, crane, aerial, slow pan, POV. Composition: wide, close-up, extreme close-up, low angle, two-shot. Optics: shallow depth of field, wide-angle, soft focus, macro, deep focus. One shot type, one movement and one lens per prompt is enough; stacking more of them produces mush.
Kling’s guide pairs each camera direction with the job it does, which is the useful way to choose one:
| Direction | Write it as | Use for |
|---|---|---|
| Close-up | Close-up on the speaker’s face | Dialogue, emotion, product detail |
| Wide shot | Wide shot showing the terrace and both characters | Setting, scale, scene coverage |
| Push-in | The camera slowly pushes toward the subject | Emphasis, tension, reveal |
| Pan or tilt | The camera pans across the room, or tilts up to the sign | Revealing space or vertical detail |
| Tracking shot | The camera follows beside the runner | Movement and continuity of action |
Pacing is prompted the same way, with a speed word attached to a reason: “a slow push-in as the character thinks” for calm or suspense, “a steady tracking shot following the subject” for readable action, “a quick cut to a close-up reaction” for energy. Where a shot needs an exact length, that is what the multi-shot syntax below is for.
The two prompts in the walkthrough above are written to this pattern. Read them against the slots and the structure is visible: framing first, then the subject and what it does, then the light, then the grade.
Sound is prompted separately, in three forms. Dialogue goes in quotation marks: A woman says, 'We have to leave now.' Effects are named: SFX: thunder cracks in the distance. Ambience is described: Ambient noise: the quiet hum of a starship bridge. Veo includes audio in the price by default, so you are paying for it whether you prompt it or not.
Three things about dialogue, from people who have tested them. The words have to be in quotation marks with an attribution: one maker’s rule is “put the spoken words/lyrics in quotes and put He/she/it sings or says”, and he reports that omitting them produces a dead or badly synced result “roughly 1 out of 5 times”. Distance kills lip sync — “the further the camera is from the character, the worse” it renders, and medium shot is the practical minimum. And ask the model not to produce music: baked-in music cannot be cut around later, so several makers prompt for effects and ambience only and lay the track themselves.
For things you do not want, whether there is a switch for it depends on the model. Kling exposes a real negative_prompt, pre-filled with “blur, distort, and low quality”, and a cfg_scale for how literally to follow the prompt. Veo has no documented negativePrompt parameter, so there you phrase it positively instead: Google’s guidance is “a desolate landscape with no buildings or roads” rather than “no man-made structures”. Check which you are using before assuming either.
Starting from a still#
Both providers support conditioning a generation on an input image, and it is the single biggest control you have. In Sora the input_reference image becomes the literal first frame, and it has to match your target resolution exactly. Veo takes an image-to-video input the same way.
The walkthrough does this at every generated shot, and the chart above is why: you are paying the video price for motion, not for discovering what the shot looks like.
The stronger version is two-frame keyframing, and the practitioner who wrote it up was explicit about what drove him to it: pure text-to-video “drifted every couple seconds. Character height changed. Background buildings shuffled. Robot color jumped between cuts.” His method is to generate the first and last frame in an image model, hand both to the video model, and describe the middle. Three details he says matter:
- The prompt has to ask for the middle explicitly. His template starts “Show what happens in between.” Leave that out and you get “a morph effect, like a Photoshop tween” instead of motion.
- Give the number of camera angles as an integer, not a vibe. He uses three for a five-second clip, four or five for ten seconds.
- Then take the last frame, regenerate a wider version of it in the image model, and use that as the next clip’s opening. That is how the sequence continues without restating everything.
One more from a heavy user of the same approach: prompt against the soundtrack. “Do not generate any background music. Only generate the corresponding sound effects.” Music baked into a clip is very hard to cut around later.
One constraint to know: in Veo’s image-input modes the personGeneration parameter is limited to allow_adult, where text-to-video permits allow_all.
Keeping the same face across shots#
Drifting faces and products are the most common thing that makes a sequence look wrong, and this is where the tools changed most during 2026. The old answer was to paste a description into every prompt and hope. The current answer is to build the character once as a reusable asset.
Kling’s elements. You create an element by uploading or recording a three-to-eight-second video of the subject, ideally showing it from several angles. The model extracts the appearance and, if it is a person, the voice. You give it a name, and from then on you write @Grace in a prompt and get the same person. Once a voice is bound to the element you do not restate it. Alternatively you supply multi-angle stills plus a voice recording of three seconds or more, which their guide says gives more accurate lip sync.
A prompt using two of them, from Kling’s own documentation, which also shows the multi-shot syntax:
Shot 1 (3s): Mid-shot, background @Image. @Grace sits on the sofa eating cookies as @Alan walks in holding @Samoyed. @Samoyed lunges for the cookie in @Grace’s hand. @Grace says, “Hey! Watch your dog!”
Shot 2 (2s): @Alan sits beside her, pulling the leash and lifting @Samoyed. Close-up, @Alan says, “He just likes cookies more than me.”
Shot 3 (3s): Close-up, @Grace smiles and says, “Well, he has good taste at least.”
Three shots, eight seconds, three named assets, dialogue inline, in one generation. That is the shape of current practice, and it is worth copying the format exactly: Shot N (duration): then framing, then action, then the line in quotes.
The blank version to copy. Keep the durations adding up to your model’s limit, which is 15 seconds on Kling 3.0 Omni:
Shot 1 (3s): [framing], background @Location.
@Person [action]. @Person says, "[line]".
Shot 2 (2s): [framing]. @Person [reacts].
Shot 3 (3s): Close-up. @Person says, "[line]".
Build @Person and @Location once, before you write any of that, and they stay usable across every clip in the series.
Seedance’s reference stack. Different approach, same goal: up to fifty inputs across images, video and audio in one generation. Their documentation splits the roles — images control appearance and composition, video controls motion style, audio controls rhythm and timing. Handing it a music bed as an audio reference to cut against is a real capability, not a workaround.
Veo’s reference images. Up to three images as content and style references. Fewer controls than the above, and reference mode forces the 8-second duration, so an attempt costs $0.80 at Fast 720p rather than $0.40.
First and last frame. Available on Kling as end_image_url, on Veo as a named feature, on Wan as an optional end image. Generate both frames in an image model, hand over both, describe the move between them. This is how you get a specific camera movement instead of whatever the model chooses.
What not to lean on: the seed. Google’s blog suggests reusing one for consistency, but the Gemini API documentation states that seed on Veo 3 models “doesn’t guarantee determinism”. A tendency, not a mechanism. Use elements or references instead.
Long single takes need production notes, not shot descriptions#
Once a generation runs to twenty or thirty seconds, a shot description stops working. Seedance’s guide is explicit that the prompt becomes production notes with a timeline, and the structure it recommends is worth following in order: duration and speed, then the job of each reference, then the opening state, then a timecoded timeline, then camera path, then continuity constraints, then audio, then the closing state, then what is forbidden.
Give every reference exactly one job, and say what to ignore. References are numbered by upload order — @Image1, @Video1, @Audio1 — and the pattern is:
@Image1 controls only the exact portable espresso maker: preserve its short cylindrical proportions, matte cobalt-blue shell, black rubber grip ring, circular copper button, and clear lower chamber. Do not copy @Image1’s studio background.
The rule underneath that: never give the same job to an image and a video reference. One carries the look, the other carries motion and timing. Overlap them and they fight.
Stage the action by timecode, so cause lands before effect.
0-4 seconds: the server places a full ceramic coffee cup beside an open sketchbook and walks away. 4-8 seconds: a busboy passes behind the seated customer; the edge of his tray lightly catches the cup handle.
State what must not change, and list the failure modes as prohibitions. This is the part that reads oddly and does the most work:
Keep the same messenger, jacket, helmet, bottle, phone, clerk, counter, cooler, and store layout from first frame to last. No cuts, no slow motion, no repeated entrance, no duplicated bottle or helmet, no object teleportation.
Where something leaves frame and returns, say how long it is hidden and what has to be identical when it reappears — same face, same coat, same walking speed, same direction. Occlusion is where continuity usually breaks.
The blank version, in the order the guide recommends:
[N]-second continuous single take, [location], [time of day],
all action at natural real-time speed.
@Image1 controls only [the thing that must not change].
Do not copy @Image1's [background / pose / lighting].
@Video1 controls only [motion style and pacing].
Do not copy @Video1's [subject / location].
@Audio1 sets [rhythm and timing].
Opening state: [what is in frame at 0 seconds].
0-[t] seconds: [action].
[t]-[t] seconds: [the reaction that action caused].
[t]-[N] seconds: [how it resolves].
Camera: [path and framing across the whole take].
Keep the same [list every person, garment and object] from first
frame to last. No cuts, no slow motion, no repeated entrance,
no duplicated [object], no object teleportation.
Audio: [dialogue, effects, ambience].
Ending state: [what is in frame at N seconds].
The prohibition list looks paranoid and is the part that earns its place. Every item on it is a failure these models actually produce.
Compare that with the Kling approach above. Kling wants several short shots with named elements and a duration on each; Seedance wants one long take with reference jobs and a timeline. Both work. They are different documents to write, and picking the model without knowing which document you are writing is how a day disappears.
How many attempts, and what they cost#
OpenAI’s documentation says it plainly: use shorter clips while you are iterating on prompt, motion or composition. The same logic applies to tiers. Veo’s are $0.05 per second on Lite, $0.10 on Fast, $0.40 on Standard, all at 720p.
How many attempts is the question everybody wants answered, and the widely circulated figures — one usable clip in four, three generations per shot — trace back to marketing blogs rather than to anyone who has done it. The most disciplined account we could find is from someone who made four short films in a month, and his working rule is batches rather than a ratio:
T2I: batch of 5 per shot. Usually 2 to 3 are trash, 1 to 2 are usable. I2V: batch of 3 per shot.
On top of that he deliberately overshoots the edit: “I intentionally overshoot by 20 to 50%… a lot of generations will be unusable or only good for 1 to 2 seconds.” For a three-minute piece needing 36 clips he generates 50 to 55. His analogy is a wedding photographer shooting a thousand frames to deliver fifty.
That is the optimistic end. Someone who spent ten days on a fantasy piece and then abandoned it reports the pessimistic end: “Generated 250+ video clips to get what you see here”, and that anything decent “often required 10-20 generation attempts per shot”. Both accounts are honest; the difference is subject matter. Falling beans and a hillside are forgiving. A dragon fighting a person is not, and he stopped because “this technology just isn’t ready for action sequences”. Price the shot you are actually asking for.
The other half of the discipline is choosing where to spend. Higgsfield resells most of these models on one credit system, which is the only place we found them priced comparably. At its Plus tier a credit works out around four cents, and one generation costs:
| Generation | Credits | Roughly |
|---|---|---|
| A still in Higgsfield Soul 2.0 | 0.12 | $0.005 |
| A still in Nano Banana Pro | 2 | $0.08 |
| Seedance 1.5, 5 sec at 720p | ~3 | $0.12 |
| Kling Omni 3 image reference, 5 sec at 720p | ~5 | $0.20 |
| Kling 3.0, 5 sec at 720p | ~7 | $0.27 |
| Wan 3.0, 5 sec at 720p | 8.75 | $0.34 |
| Sora 2, 4 sec at 720p | ~10 | $0.39 |
| Veo 3.1 Fast, 4 sec | ~11 | $0.43 |
| Seedance 2.0, 5 sec at 720p | ~22 | $0.86 |
| Seedance 2.0, 5 sec at 1080p | ~45 | $1.76 |
| Seedance 2.0, 5 sec at 4K | ~110 | $4.30 |
Higgsfield’s published credit costs, September 2026; the dollar column is our arithmetic at its Plus rate. Its tiers run $19, $47 and $99 a month billed annually for 270, 1,200 and 3,000 credits, and Seedance 2.5 is not available on the cheapest one.
Read the first two rows against the rest. A still costs half a cent to eight cents. Five seconds of video costs twenty cents to four dollars and thirty. That is a ratio between roughly ten and eight hundred, which turns “settle the shot in an image model first” from advice into arithmetic.
So for a 30-second vertical of five shots, overshot to seven shots’ worth of material with three video attempts each: twenty-one generations. On Kling 3.0 at 720p that is about $5.70 plus a couple of dollars of stills. On Seedance 2.0 at 1080p it is nearer $37. Same edit, same discipline, six times the bill, and the choice is which model’s handling of your subject you actually need.
Also worth knowing before you plan around any of this: at least one maker reports that the usable unit is not the clip but a fragment of it. “I had to go through each clip and find small usable sections.” Budget for trimming, not just for generating.
One caveat on that plan: iterating on a cheaper tier tells you whether the prompt and composition work, not what the expensive tier will produce, because they are different models. If exact matching matters, iterate within the tier you will ship on. For variant work, OpenAI’s batch tier is half price and you are not waiting on any individual render.
If you are on a subscription rather than API billing, do the arithmetic before you commit to a client. Runway publishes both halves: Gen-4.5 costs 60 credits per five seconds, and Standard at $12 a month includes 625 credits. That is about 52 seconds of generation a month, or roughly 13 seconds of keepers at a one-in-four rate.
The order of work, which matters more than any setting#
Two habits show up in the accounts of people who finish things, and both are about sequencing rather than craft.
Cut the whole thing as an animatic before you generate any video. One maker generated every still and the voiceover, cut a complete rough version, and only then generated video — at exactly the lengths the edit had asked for. That inverts the usual order, in which you generate clips and then discover the edit does not want them. Stills are between one-tenth and one eight-hundredth the price of clips, so the animatic is nearly free and it tells you which shots you do not need at all.
Do all of one stage, then all of the next. The maker of four shorts in a month is blunt about it: “Day 1 (night): do ALL text-to-image. Queue 100 to 150 and go to sleep. Do not babysit it. Do not tinker. Day 2 (night): do ALL image-to-video. One long queue. Let it run 10 to 14 hours if needed.” His reason is not throughput, it is attention: “If I do it in little chunks (some T2I, then some I2V, then back), I fragment my attention and the film loses coherence.”
And a note on where the disappointment usually comes from. A first-timer who produced 84 shots concluded: “The final cut didn’t come out as well as I’d hoped — mostly a story problem rather than a generation one. Too much narration, not enough actually happening on screen. Next one I’ll spend the time on the writing.” Nobody who has finished one of these says the models were the constraint.
Assembly#
You now have clips that do not quite match each other. What makes them read as one piece:
- Change the angle at every join. This is the most useful technique we found, and it comes from someone who hit the wall the obvious way. Chaining last-frame to first-frame for a continuous shot “degrades badly until it looks completely fried by the 4th iteration” — he compares it to re-copying a VHS tape. The fix: use the last frame as a reference and immediately cut to a different angle. “It will reconstruct the scene, and you can then come back to the original angle later.” Shot, reverse shot, back again. His summary is worth keeping: “Long seamless scenes are not a good idea with AI anyway.”
- Put the known failures on the seams. One maker structured a 28-second vertical so that the whip pans, a strobe section and a blackout landed exactly where the artefacts were, using them as cover. If a face drifts on a head turn, that is where the cut goes.
- Cut on movement. A cut landing while something crosses frame hides the mismatch between takes. A cut on a static frame exposes it.
- Trim the ends. Generated clips drift as they run. Take the middle and lose the last half-second.
- Find a prop that distorts everything equally. A nice trick from a short film: a megaphone “did most of the heavy lifting since it distorts every voice the same way.” Anything that applies the same treatment across every shot hides the variance between them.
- Match the motion. If your filmed shots are handheld and the generated ones glide, add movement to the generated ones or stabilise the filmed ones. Consistent camera behaviour reads as more coherent than higher image quality.
One correction to advice we gave earlier and to advice you will see everywhere: grain does not fix artefacts. A working colourist’s assessment of generated footage is that collapsed skin tones, waxy faces, smeared texture, banded gradients and clipped highlights “are far too obvious to be fixed by simply adding film grain.” A single grade across the timeline plus light grain does help generated and filmed material sit together. It does not repair a bad take. Re-generate the bad take.
And a warning about rhythm, from editors watching this material arrive: cutting every two seconds because that is what the model produces is visible. “It’s not an animated story, it’s a chain of 2-second gifs.” The traditional alternative is to cover a moment from several angles and cut between them, rather than generating two seconds of action at a time. That costs more generations and looks like film.
Captions: burn them in, because most viewing is silent. Keep them inside the safe area, below.
Export and upload#
All three platforms take the same master: 1080×1920, 9:16, MP4 or MOV, H.264, 30fps. One export covers them.
What differs is where each platform’s interface covers your frame, and this is far easier to see than to remember as numbers:
Export separately per platform rather than reposting a file with another platform’s watermark or letterboxing on it. That is what unoriginal-content detection is looking for.
Then the disclosure toggles, which differ:
- YouTube. Required only for realistic content that misrepresents a real person, place or event. Not required for unrealistic scenes, filters, effects, production assistance, or cloning your own voice. YouTube states the label does not reduce reach or earning eligibility. Setting it is free; not setting it where required escalates to a 90-day monetisation suspension and then removal.
- TikTok. Labelling required for synthetic humans, cloned voices and photorealistic generated product images, with some uploads auto-tagged. Barred whatever the label: imitating a news source, fabricated crisis footage, a public figure appearing to endorse something, a private person’s likeness without consent.
On TikTok specifically, its Creator Rewards rules exclude content with “minimal original input or edits”, and creators report AI-visual videos being flagged as unoriginal even when they wrote and edited them. We could not verify that against TikTok’s own documentation, so treat it as widely-reported experience rather than policy. Either way, if you are testing generated footage, test it on Shorts where the rules are written down.
Six people’s actual pipelines#
Published writeups, with their own numbers. Worth reading for the range as much as the method: the same year, the same models, and outcomes from seven hours to fifteen days to abandoned.
| What they made | Pipeline | Hardware | Time |
|---|---|---|---|
| 3-minute sci-fi short, all local | Local LLM for script → Flux stills → LTX 2.3 + Wan 2.2 VACE, with a JSON-structured prompt template | RTX 4090 | ~15 days, 5 on the script |
| Four 3-minute music shorts in a month | Music first → ChatGPT shot list → FLUX Fluxmania V → Wan 2.2 → CapCut | RTX 5090 | ~2 hours of editing each |
| 1-minute anime short | Own style LoRA in LTX 2.3, first frame decides everything, Nano Banana for shots the LoRA could not do | RTX 5090 | ~7 hours, deadline day |
| Vertical comedy sketch, two characters | Film screengrabs as image and voice references → MiniMax H3 in ComfyUI → DaVinci Resolve | RTX 3070, 8GB | 2 days, 4–5 hours of work |
| 15-second 5-shot anime sequence | Four reference images and one long multi-shot prompt, MiniMax H3 | 5070 Ti, 16GB | 136 seconds of render |
| Fantasy piece, abandoned | Nano Banana 3×3 storyboard grids upscaled with SeedVR → LTX-2 | RTX 6000, 96GB | 10 days, then stopped |
Three things to take from the spread. An eight-gigabyte card produced a finished piece in two days while a ninety-six-gigabyte card produced an abandoned one in ten, so the hardware is not the variable. Every one of them writes prompts with an assistant rather than by hand. And the two who finished fastest both started from a fixed reference set and refused to deviate from it.
One more technique worth stealing#
A maker of a horror series, blocked by one platform’s filters, uses a pencil storyboard as an intermediate step. He generates the shot images, then asks an image model, verbatim:
create a 5 panel story board using the images in sequence. pencil sketch only. No text. maintain character consistency
Then he uploads that single storyboard sheet to the video model as the reference, with a multi-shot prompt. His finding on length: for five shots, a ten-second prompt works best. The sheet does two jobs — it locks continuity across shots in one image, and a pencil sketch passes content filters that a photoreal still does not.
What one finished clip costs in time#
The money is small and the hours are not, which is the opposite of how this gets sold. Self-reported figures from people publishing finished work: 60 hours for a trailer on about $100 of credits, 70 hours for a two-minute piece, over 100 hours for a single sequence. That is 30 hours or more per finished minute at the ambitious end.
Short-form is far below that, but the shape holds: the generation is cheap and fast, and the selecting, trimming and matching is where the day goes.
The most useful scheduling advice we found is to separate the phases rather than interleave them. One maker’s routine for a three-minute piece: one night queueing 100 to 150 stills and going to bed, the next night running all the image-to-video jobs in a single queue for 10 to 14 hours, then two hours blocked out to cut it. Your attention is the scarce resource, and it is wasted watching renders.
Anyone quoting client work should price the shot list, the selection and the assembly. Not the render.
Where that work comes from, and what to charge: making video with AI for the pricing and rights, UGC for brands for the same buyer without the generation, and building an audience you own for what the views are actually for.
Documentation and price lists: Kling’s prompt guide and Kling VIDEO 3.0 Omni guide; Seedance 2.5 reference-to-video schema and its prompting guide and Kling v3 Pro schema on fal.ai; Higgsfield credit costs; Gemini API, Veo and pricing; Google Cloud prompting guide for Veo 3.1; OpenAI video generation guide and pricing; Runway pricing; ElevenLabs pricing; YouTube and YouTube Help on Shorts revenue and disclosure; TikTok Creator Academy. The practitioner methods, numbers and quoted strings come from public writeups on Reddit by u/Psi-Clone, u/Motor_Mix2389, u/crinklypaper, u/justin_wiggins, u/Time-Ad-7720, u/sktksm, u/Fresh-Resolution182, u/Zealousideal-Cry7806, u/Sniper_yoha, u/sarfarazh, u/Dohwar42, u/Ok-Wolverine-5020 and u/Legitimate_Bit2775. These are self-reported accounts of single projects, not measured results, and hardware and model versions vary between them. Safe-zone pixel values are community measurements. Shorts RPM ranges are third-party estimates with no published methodology. Cost worked examples are our arithmetic from the published prices, using the batch discipline described by the practitioner quoted above. Everything attributed to makers in this page is self-reported on public forums, not measured, and the widely circulated “one usable clip in four” figure traces to marketing blogs rather than to practitioners, which is why we have not used it. Checked 8 September 2026.