Long-Form Video With AI

Updated 2,389 words · about 11 min

Nothing generates a finished fifteen-minute video, though the limits are higher than people assume and moving fast. Seedance 2.5 does thirty seconds in one go, Kling 3.0 Omni fifteen, Sora chains six extensions into two minutes, and MiniMax H3 Max will stream a continuously directed session of up to fifteen minutes while holding characters and setting. What none of them does is produce a long piece you would publish: coherence degrades across a take, and a streamed session is raw material, not an edit.

So long-form is an ordinary production with generated material in named slots. This page is the pipeline: where AI goes, what it costs per finished video, and the structure that keeps people watching, since watch time is what all of it is paid on.

Why the length is worth the trouble#

Long-form keeps 55% of ad revenue against 45% for Shorts. Videos of eight minutes or longer can carry mid-roll ad breaks, which is the biggest single lever on RPM and does not exist below that length. And the Partner Programme door opens on watch time, which long videos accumulate far faster.

That last one is worth doing as arithmetic. The threshold is 8,000 qualified public watch hours in twelve months, which is 480,000 minutes:

Views needed to reach 8,000 watch hours, by video length To pass YouTube’s 8,000 qualified watch hours a 20-minute video needs about 60,000 views, a 12-minute video 89,000, an 8-minute video 120,000, a 3-minute video 291,000, and a 1-minute video 800,000. A twenty-minute video therefore needs about a thirteenth of the views a one-minute video does. Views needed to pass 8,000 watch hours Same threshold, same channel. Only the length of the video changes. 20 minutes 60,000 12 minutes 89,000 8 minutes mid-rolls become possible here 120,000 3 minutes 291,000 1 minute 13× the views of the 20-minute cut 800,000 Our arithmetic from YouTube’s published 8,000-hour threshold, which is 480,000 minutes, using assumed average view durations of 40% at 20 minutes rising to 60% at one minute.
The threshold is watch time, not views, so length does most of the work. Substitute your own average view duration from Studio: the shape holds unless your retention is wildly different at the two ends.
Video lengthAssumed average view durationViews needed for 8,000 hours
20 minutes40% (8 min)60,000
12 minutes45% (5.4 min)89,000
8 minutes50% (4 min)120,000
3 minutes55% (1.65 min)291,000
1 minute60% (0.6 min)800,000

The percentages are assumptions, not measured data. Substitute your own from Studio. The 480,000-minute figure is arithmetic from YouTube’s published threshold.

A twenty-minute video needs about a thirteenth of the views a one-minute video does to pass the same threshold. Platform ad revenue has the full rules and RPM by subject.

The script, and using a model without sounding like one#

The research slot is the highest-value use of these tools and the one that will publish a wrong number under your name if you let it.

What works is using the model as an editor and an opponent rather than a writer. Give it your outline and ask what a knowledgeable viewer would object to. Ask which of your claims has no primary source behind it. Ask what the strongest argument against your conclusion is. All three produce material you can use. Asking it to write the script produces the thing viewers recognise instantly.

Then a verification pass with a rule: every figure that reaches the script gets a source you have opened yourself. Not a source the model named — a page you loaded. This is not optional in a video about money, and it is where the time savings from the earlier steps go.

The disclosure position on all of this: script drafts, ideation and thumbnail generation are production assistance, which YouTube exempts from the altered-content label entirely.

Generated b-roll: the actual pipeline#

The workable ratio is a few seconds of insert per minute of video, cut into footage that is otherwise yours. Fifteen minutes carrying ninety seconds of generated material across fifteen shots is a normal shape.

How to produce those fifteen shots without wasting money:

Write them as a list first, with a duration each, and mark what is impossible to film versus merely inconvenient. Generated footage earns its place on the impossible ones: the interior of a machine, a historical scene, an abstract sequence for a segment about interest rates. On the merely inconvenient ones you will usually be happier filming.

Generate the still, then animate it. Both providers condition on an input image — Sora’s input_reference becomes the literal first frame and must match your target resolution; Veo takes an image-to-video input. Judging a still is free, judging a video costs a generation, so settle composition and light in an image model first and pay the video price only for motion.

Follow a published prompt formula. Google’s for Veo is [Cinematography] + [Subject] + [Action] + [Context] + [Style and ambiance]; Kling’s is Subject + Action + Scene + Camera + Lighting and mood. Both parse ordinary film vocabulary — dolly, tracking, crane, slow pan; wide, close-up, low angle; shallow depth of field, macro, anamorphic — and both reward describing what is visible over what it means. One shot type, one movement, one lens per prompt. Audio is prompted separately as dialogue in quotes, SFX: lines and Ambient noise: lines, and Veo bills audio whether you prompt it or not. Short-form video has the full prompt mechanics, since they are identical for an insert and a vertical.

Match the model to the insert. For a one-off abstract shot almost anything will do, and the cheap tiers are genuinely cheap: five seconds of Seedance 1.5 at 720p costs about twelve cents. Where an insert has to show the same object or person as an earlier shot, use a model with a real consistency mechanism — Kling’s named elements, built once from a three-to-eight-second video, or Seedance 2.5’s reference stack, which takes up to fifty image, video and audio inputs in one generation.

Match it to your filmed material in the edit. This is the step people skip and it is what makes inserts read as part of the video rather than as stock. One grade across the whole timeline and consistent camera behaviour: if your A-roll is handheld, do not cut to a perfectly gliding generated shot without adding some movement to it.

Two cautions from people who do this professionally. Grain unifies material but does not repair it — a working colourist’s verdict on generated footage is that collapsed skin tones, waxy faces, banded gradients and clipped highlights “are far too obvious to be fixed by simply adding film grain”. And grade it as though it were 8-bit footage, because it behaves like it. If a take is bad, replace the take.

On upscaling, expectations are worth lowering. The order people use is upscale first, then grade. But necessity is contested and the throughput is punishing: reported speeds for Topaz’s Starlight Mini run from 0.9 frames per second down to 0.1 depending on machine and source, and one user on a 5090 reports “only 10 seconds of video need 4 hours export time”. Another found it made almost no difference to generated footage: “the before and after was nearly identical.” The rule people settle on is that the source decides the ceiling — “it won’t add new leaves to a tree.” Generate at the resolution you need rather than planning to rescue it later.

What to keep out of generated shots, because these are documented failure modes rather than taste: legible text, hands performing the task, a recurring human face, water, fabric. Telling the difference from slop has the full list with fixes.

Dubbing, with the mechanics#

This is the clearest paid margin in the category and the part with the least guesswork, because both halves are published.

ElevenLabs charges 3,000 credits per minute for automatic dubbing without a watermark, and the Creator tier is $22 a month for 121,000 credits. So a fifteen-minute video costs 45,000 credits to dub, about $8 at that tier, and the tier covers roughly two and a half such videos a month before you need more. Voiceover from a script is cheaper still: fifteen minutes of speech is around 12,000 characters at one credit each, so about $2.

Two things worth knowing before you plan around it. Commercial use requires at least the $6 Starter tier — there is no commercial licence on the free plan. And instant voice cloning starts at Starter while professional cloning needs Creator.

Then the delivery mechanism, which is the part people get wrong. YouTube’s multi-language audio lets you upload your own dubbed tracks to new and already published videos, from Studio on desktop. The file is audio-only and roughly the length of the video. Access needs Advanced features and is still rolling out to a subset of those channels. The gotcha: this is separate from YouTube’s own automatic dubbing, and if an automatic dub already exists for a language you have to delete it before uploading yours.

Doing it this way keeps one channel with its watch time intact rather than splitting an audience across language channels. AI-assisted services covers selling this to other people, which is the better business than doing it for yourself.

What a fifteen-minute video costs#

Ninety seconds of inserts is fifteen shots of six seconds. The realistic attempt rate is not the “three generations per shot” figure that circulates — that traces to marketing blogs. The most disciplined practitioner account we found works in batches instead: five stills per shot, of which “2 to 3 are trash, 1 to 2 are usable”, then three video attempts per shot, and a deliberate overshoot of 20 to 50% on the total shot count because some clips turn out “only good for 1 to 2 seconds”.

So budget for about twenty shots’ worth of material, three video attempts each: sixty generations of six seconds, or 360 seconds of billed output.

  • On Veo 3.1 Lite: 360 seconds at $0.05 = $18
  • On Veo 3.1 Fast 720p: 360 seconds at $0.10 = $36
  • On Veo 3.1 Standard: 360 seconds at $0.40 = $144
  • Stills for those shots in an image model: a few dollars
  • Voiceover if you are not recording yourself: about $2
  • Each additional language, dubbed: about $8

So $20 to $150 of generation for a video that might carry mid-rolls for years, and the tier you shoot on is the whole spread. Iterating on a cheaper tier tells you whether the prompt and composition work but not what the expensive tier will produce, since they are different models; if exact matching matters, iterate on the tier you will ship.

The hours are the real cost. People publishing finished AI-assisted work self-report 60 hours for a trailer, 70 for a two-minute piece, over 100 for a single sequence. The most useful scheduling habit we found is to keep the phases apart rather than interleave them: one night queueing 100 to 150 stills, the next running all the animation jobs in a single 10-to-14-hour queue, then a couple of hours blocked out to cut. Attention is the scarce input, and watching renders wastes it.

Structure, since retention is the whole economy#

State the payoff in the first thirty seconds, then earn it. Not a teaser, the actual claim. “This costs £40 a month and I am going to show you the invoice” gives someone a reason to stay for eleven minutes.

Write in segments that could each stand alone. Retention graphs fall at transitions. A segment opening with its own small question survives the join; one opening with “so, moving on” does not.

Put something concrete every ninety seconds. A number, a document on screen, a demonstration, a named example. Abstract stretches are where people leave, and they are also where generated b-roll is most tempting and least useful.

Plan the mid-rolls. Above eight minutes you place them yourself. Put them after a payoff rather than before one, and treat the automatic placement as a starting point.

Do not pad to reach eight minutes. Padding shows as a retention cliff, and a five-minute video that holds beats a nine-minute one that loses half its audience at minute three, mid-rolls or not.

What gets the channel demonetised#

YouTube renamed Repetitious Content to Inauthentic Content in July 2026 and named three categories that cannot earn: content following a template with little variation and reproduced at scale, content built on manipulative or shock formulas, and AI personas giving advice on finance, law, health or medicine.

The test is a question about your next video: if it could be produced by swapping the topic and changing nothing else, that is a template. True whether a person or a model filled it in, which is why the policy does not mention AI.

Two formats are finished. Narration read verbatim from an external source over stock footage, and story-aggregation channels reworking forum posts. In January 2026 YouTube deleted or hid the videos of at least eighteen channels publishing this material; one had over 1.2 billion views. What gets you demonetised has the rules and a self-audit.

There is also a production trap with numbers on it. Creators who adopted AI video tools early tripled their output on average, and only 34% held or improved their engagement rates. Three times the work for the same income is a worse job than the one you started with.

Who this suits#

Anyone who knows something specific and can talk about it for fifteen minutes without repeating themselves. That is the qualification, and no tool supplies it.

People already making long video who want the overhead down: research, inserts, dubbing, metadata.

Anyone selling video to businesses, where explainers and internal video are ordinary paid work with a budget attached.

Who it does not#

Anyone who wants output volume to be the strategy. The policy blocks it and the engagement data says it fails before the policy gets to you.

Anyone who needs money inside six months. Selling a skill and UGC pay in weeks; ad revenue does not.

Anyone unwilling to be the author. Every part of this that survives the rules depends on there being a person with a view.

Documentation and price lists: Kling’s prompt guide; Seedance 2.5 schema and MiniMax H3 Max Director on fal.ai; Higgsfield credit costs; Gemini API, Veo and pricing; Google Cloud prompting guide for Veo 3.1; OpenAI video generation guide; ElevenLabs pricing; YouTube Help on multi-language audio; YouTube Help on mid-roll breaks; YouTube on Partner Programme changes; YouTube Help on disclosure. RPM ranges are third-party estimates with no published methodology. The view counts and the cost worked examples are our arithmetic from published figures, using stated assumptions about view duration and the batch discipline described above. Hours, keeper rates, grading and upscaling observations are self-reported by makers on public forums, not measured; the circulated “one usable clip in four” figure traces to marketing blogs and is not used here. Checked 8 September 2026.