⋮⋮
hirenpateldev — blog/posts/ai-short-film-pipeline.md
⎇ main
📁
🔍
📦
Explorer
HIRENPATELDEV
📁.config
📂about
📄README.md
📄uses.md
📂blog
posts.ts
#tags.ts
📂projects
repos.ts
📂off-keyboard
📜reading.log
🚗trips.log
🏋workout.log
.git-log
{}skills.json
$contact.sh
📄README.md×
posts.ts×
ai-short-film-pipeline.md×
← back to postsblog/posts/ai-short-film-pipeline.md
--- frontmatter ---
title: "Making a Short Film With AI, End to End"
date: "Thu Sep 10 2026 00:00:00 GMT+0000 (Coordinated Universal Time)"
tags: ["ai", "video", "filmmaking"]
readMins: 12

# Making a Short Film With AI, End to End

The pipeline I use to get from a one line idea to an exported 60 second film, the settings that matter, and the prompt that does the real work at each stage.

The tools in this post will be renamed or replaced inside a year. The order you do things in will not, and that order is the part worth writing down. So this is a pipeline rather than a review.

Start with the economics, because they decide the shape of everything else. Inside Flow, generating a still image costs nothing, you are capped by a daily limit rather than credits. Generating video costs 20 credits for one output and 80 if you ask for four. Every decision you can push into the image stage is one you get to make for free, over and over, until it is right. Every decision you postpone to the video stage costs real money and gets a fraction of the retries.

The other constraint is the film. Keep it between 30 and 90 seconds, one location, one character and at most one other. That is not modesty, it is what the medium supports. At this length there is no room to develop a second character anyway.

The tool chain

  • Claude for everything made of words. Loglines, the script, the shot list, the revision passes.
  • Flow, Google Labs' filmmaking tool, and the hub for all of this. Nano Banana generates stills inside it, Veo generates video from those stills. One month free, then about $20 a month.
  • Suno for music.
  • ElevenLabs for sound effects, and for voice when a character speaks across more than one shot.
  • DaVinci Resolve for the edit and the export. Free tier, and its colour tools are the reason to pick it over CapCut.

You can reach Nano Banana through Gemini directly, but keeping stills in Flow means references, keyframes and video all live in one project.

text
STAGE        TOOL              WHAT COMES OUT           COST
concept      Claude            10 loglines, keep 1      free
script       Claude            four beats, 30-90s       free
style        Flow / Nano B.    the look, as images      free
character    Flow / Nano B.    front, profile, rear     free
shot list    Claude            scene / size / movement  free
keyframes    Flow / Nano B.    3x3 grid, then singles   free
video        Flow / Veo 3.1    one clip per prompt      20 credits
music        Suno              one cue, spliced         cheap
voice + sfx  ElevenLabs        lines, foley, ambience   cheap
edit         DaVinci Resolve   the film                 free

The prompt spine

Every visual prompt takes the same shape, six parts in order:

  1. Role. "You are an award-winning cinematographer."
  2. Subject. Who, doing what, in what posture. Posture carries more than expression at this size.
  3. Environment. The place, the time of day, what is in frame that nobody is touching.
  4. Light. Direction, quality, colour. "Golden hour through a window on the left, long shadows across the table."
  5. Shot. Size, angle, lens, depth of field, colour grade.
  6. Mood word. Two or three words at the end to give the model something to aim at. Optional, and it does more than it looks like it should.

Underneath it sits one habit: decide what you want the audience to feel before you decide what they see, then work backwards to the objects that carry the feeling. Loneliness is not a prompt. An empty chair pulled out from a table set for one is a prompt.

Stage 1, concept

Open Claude and ask it for film ideas and you get ten competent, weightless loglines that could have come from anyone, because you gave it nothing of yours to work with. Answer these on paper first, and do not tidy them up:

  • What set this off? A moment, an image, something that annoyed you.
  • What question is the story asking?
  • Who is it about, in four adjectives?
  • Which one or two genres does it live in?
  • How should the audience feel walking away?
  • Which images do you keep coming back to?
  • Where does it happen?
  • What are the constraints? Length, cast size, dialogue or silence.

A logline answers three questions and nothing else. Who is this about, what do they want, what is in the way.

text
You are a short film director known for small, quiet stories.
Give me 10 loglines for a 45 second film.

The raw material:
  spark      [what set the idea off]
  themes     [the question the story is asking]
  character  [three or four adjectives]
  genre      [one or two, no more]
  mood       [how the audience should feel leaving it]
  location   [one place]
  images     [the shots already stuck in my head]

Every logline must:
  rest on one character, carried by expression and body language
  have a setup, a turn, and a landing
  arrive somewhere true rather than somewhere moral
  be shootable in one location

Do not judge the first batch. Notice which ones pull at you, then iterate, because the first ten are almost never the answer. They are how you find out what you are circling.

Aim small. A person deciding whether to send a text, or a student wondering whether to put their hand up. Micro-moments beat macro-stories at this length.

Stage 2, the script

Four beats, and the structure is not a suggestion:

  1. Setup. The character wants something.
  2. Catalyst. Something interrupts, and the obstacle arrives.
  3. Struggle. They try things. This is where the emotional arc lives.
  4. Twist or resolution. They get it or they do not, and either way something true lands.

Go and watch Pixar's For the Birds. Three minutes, and you can point at the frame where each beat starts.

text
Take logline 4 and open it into a short film outline using this
structure: setup, catalyst, struggle, twist or resolution.
Keep it small, intimate and visually expressive.
It has to play in 45 seconds, so budget about 8 shots and write
to that budget.
Tell it through what the camera sees.

Budgeting the shot count inside the prompt matters. Left alone, the model writes a lovely two minute story and lets you find out three stages later, once you have paid to generate half of it.

Then edit, which is the half of Claude that gets wasted. Run these as separate passes rather than one big "make it better":

  • Point out anything confusing or unclear.
  • Here is everything I know about this character. Strengthen them with it.
  • Shorten the action lines by 20% and keep the visuals strong.
  • Replace exposition with visual cues.
  • Make the dialogue less on the nose, more subtextual.

Stage 3, visual development

If you do not define the look of your film, the model invents one, and it is rarely the one you wanted.

Find the style first. Throw everything you know about the character into Flow and generate freely. Anime, Pixar, photoreal, noir. This costs nothing, so do a lot of it, and expect to change your mind late. Landing on photorealism, reaching a later scene and realising the film wanted to be animated is completely normal.

Then pull a clean character plate.

text
Take this image and make it a full body studio photograph of the
character, standing in a neutral pose against a clean neutral
backdrop.
Keep the art style exactly as it is in the source image.

Generate front, profile and rear, one folder per character. Locations get the same treatment and have to match the characters stylistically. Animated character, animated room. Mixing them is instantly visible.

Shot list next.

text
Create a shot list for my short film, roughly 45 seconds.
Columns: scene, shot number, shot size, camera movement, notes.
Prioritise emotional storytelling, simple staging, micro-moments.
Keep it feasible for AI generation.

[paste the full script]

Then the contact sheet, which is what turns a shot list into something you can generate from. Hand Flow one reference image and ask for a nine panel grid from the same scene. Because they come out of a single generation, they share a character, a location and a colour grade for free.

text
Act as a storyboard artist. Take the attached reference image and
expand it into a 3x3 contact sheet of keyframes for a 15 second
sequence about [one line synopsis].

Hold constant across all nine panels:
  the same character, same face, same wardrobe
  the same location, same time of day
  one colour grade across the whole sheet

Change only framing, angle, blocking and expression.

Cover a range: wide, medium, over the shoulder, insert,
extreme close up, and one high angle.
Depth of field follows shot size, deep on the wides and shallow
on the close ups.
Introduce nothing that is not in the reference. If the scene
needs tension, put it off screen as a shadow, a reflection or a
gaze out of frame.

Output one single grid image. No text, no labels, no captions.

Then extract panels one at a time with "output the panel at row 2, column 3 on its own, at full size." One gotcha: these grids come out horizontal, so if your film is vertical, set the aspect ratio to portrait before you extract.

Refining a still. Starting over is the wrong move. These adjustments do the work instead:

  • Zoom. 2x in or 2x out on an existing image.
  • Time shift. "Move forward one second in this action." Useful for building a matched pair of frames.
  • Add or remove elements. Put the scarf on the chair, clear the noise out of the sky.
  • Perspective shift. Turn a wide into an over the shoulder, a low angle, a three-quarter profile.
  • Lighting. Ask for cinematic lighting or a shallower depth of field.

First and final frames. When a shot has a specific destination, build both ends as stills before generating any video. Combine the character plate and the location plate, dress the set, and that is your first frame. Build the last frame the same way, changing only what the action changes. Now the video model interpolates between two images you already approved instead of guessing, and this is where the free retries pay for themselves.

Stage 4, generating the video

Four settings, and getting them wrong is expensive.

  • Aspect ratio. Match the stills you already generated, decide once, carry it through. Mismatched frames are obvious in the edit.
  • Outputs per prompt. One. Four costs 80 credits and you will burn a month of budget in an afternoon.
  • Model. Latest available, currently Veo 3.1, which generates synchronised sound with the picture. It comes in fast and quality, and quality can cost five times as much. Fast is often better, which is not what you would expect and is worth knowing before you spend the difference.
  • Mode. Flow offers text to video, frames to video, ingredients to video, and create image. Use frames to video and ignore the rest, because you have already done the work the other modes are for. Within it, first frame alone is the everyday choice. Add a last frame to lock a specific emotional beat.

Build the scene like an editor. Wide, then closer, then intimate. Establishing shot for the geography, mediums and over-the-shoulders for the performance, close-ups where the meaning lands. For a film this length that means one establishing shot, one or two character shots, two or three struggle shots, one twist shot, and one final image. That is the whole film, and every shot has to earn its place.

Speak in cinematic language. "Cinematic" is a word the model has seen attached to a million different looks, so it means nothing on its own. These mean something:

  • Size: extreme wide, wide, medium, medium close, close up, extreme close up, insert.
  • Angle: eye level, low, high, bird's eye, over the shoulder, dutch.
  • Movement: locked off, slow push in, pull back, pan left, tilt down, tracking alongside.
  • Lens: 24mm for a room, 50mm for something natural, 85mm when you want the background to fall away.
  • Light: soft window light, hard side light, a practical lamp in frame, backlit through haze, golden hour.
  • Grade: warm amber, cold teal, desaturated, high contrast with crushed blacks.

Be willing to name things wrongly. My favourite example from the course: a dog character kept dropping onto four legs whenever the prompt said "dog", and asking for "anthropomorphic" turned it into an actual sumo wrestler. What worked was calling it "the brown character" and letting the reference image carry the species. If the model keeps fixating on a word, take the word away.

When it goes wrong, and it will:

  • Be precise, not poetic. The model wants clarity, not atmosphere in the instruction itself.
  • One character, one action, one dominant emotion per clip. Complexity multiplies failure.
  • Reapply the reference image when the style drifts. It snaps back more reliably than any amount of rephrasing.
  • Adjust the lighting cue. It has a disproportionate effect on realism and continuity.
  • Do not settle for the first output. Three to five variations, changing angle, lighting or emotional tone.

Continuity is the honest weak spot, and it is why most AI films read as trailers. Traditional editing solves it with coverage, shooting the same action from several angles with overlap, and you cannot do that when you are generating isolated shots. Flow's Scene Builder is the partial answer. Jump to cuts to the same moment from a different perspective. Extend continues a shot forward in time, which is how you fake a single uninterrupted take. Save frame as asset lets you start either one from a frame you choose rather than the end of the clip.

Stage 5, sound

Sound is at least half the emotional experience and costs almost nothing, which is why skipping it is the worst trade in the pipeline.

Music. Prompt Suno and expect to be wrong twice before you are right. In the edit you will splice sections of the track rather than laying it flat, so the beats land where the cuts do.

Voice. Veo 3.1 generates voices along with the picture, and for a 30 to 60 second film that is often enough. You need ElevenLabs when a character speaks across several shots and has to sound like the same character each time.

Cloning is the reliable way to build one. Generate a clip in Veo where the voice is already in the right register, pull about 15 seconds of clean audio, and upload it as a voice clone. If you do not have 15 clean seconds, loop what you have in Resolve and export that. You can also describe a voice from scratch in text, though note the guardrails around child voices: describe energy instead of age, so "young male, energetic, feisty, a little scrappy" rather than naming a child.

Four settings control the output:

  • Speed 1.0. Move it for a character with a distinctive rhythm.
  • Stability 50%. Higher is more consistent but flatter, lower has more life but wanders.
  • Similarity 75%. Higher keeps it recognisably your clone, but push it too far and you get artefacts.
  • Style exaggeration 0%. Counterintuitive, but turning this up makes the model over-emote and invent emotional beats that do not fit. Expressiveness should come from your writing.

ElevenLabs v3 lets you put performance notes in square brackets inline, like giving an actor a direction. That covers ordinary shifts. For a genuinely different register, build a second voice for that mood and switch to it.

Effects. Same tool, sound effects tab. Generate each one separately so you can place and level them independently.

Layer them. A real sound bed runs to about four tracks:

  • Ambience, fading in and out between clips rather than cutting hard.
  • More of the same on a second track, overlapping the first so the joins disappear.
  • Hits, stingers and specific foley. A stinger on the extreme close up does a lot of work.
  • Music on its own track, spliced so the beats hit the cuts.

Ambience matters more than it sounds like it should. A room with birds outside reads as a room. The same room in silence reads as a rendering.

Stage 6, the edit

Organise before you import. Folders on disk for video, music and sound effects, everything in place first. Resolve links to file paths, so moving a file after importing takes the clip offline.

Assemble before you trim. Lay the shots out in shot list order and watch it through. It is meant to be too long and slightly awful, and that pass tells you what is missing.

Then cut. Two things worth knowing beyond ordinary trimming: Transform in the Inspector will keyframe a zoom across a clip, and a push from 1.2 to 1.8 adds drama the generated clip did not have. And if an over-the-shoulder has the shoulder on the wrong side for where the character is looking, flipping the clip horizontally fixes the eyeline for free.

Colour is mostly one trick. Flow's output is good, but clips drift in temperature between generations and it shows side by side. On the Color page, select the clip you want to change, shift-select the clip you want it to look like, right click, shot match. That does most of what colour correction needs on a film this size. Everything else is a deep rabbit hole.

The poster

Two minutes of work, and the film needs one image to travel with. Feed your own best frames back in as references.

text
A film poster for a short called [title], about [one line].
Match the visual tone of the attached references.
Cinematic composition, heavy negative space, almost no text.

Swap "poster" for "thumbnail" to get the other one.

The run sheet

When I come back to this, this is the part I am actually reading.

  1. Answer the concept questions on paper. No AI yet.
  2. Ten loglines out of Claude. Notice which ones pull, then iterate.
  3. Expand to four beats, budgeting the shot count inside the prompt.
  4. Run the five revision passes separately.
  5. Settle the visual style with free stills before anything else.
  6. Character plate, front, profile and rear, style locked. Location plates to match.
  7. Shot list out of Claude from the full script.
  8. 3x3 contact sheet, then extract panels one at a time. Portrait before extracting if the film is vertical.
  9. Build first and final frames as stills for every shot with a destination.
  10. Flow settings: matching aspect ratio, one output per prompt, Veo 3.1 fast, frames to video.
  11. Budget the shots: one establishing, one or two character, two or three struggle, one twist, one final image.
  12. Speak in shot sizes, lenses and lighting. If a word keeps derailing the model, take the word away.
  13. Scene Builder's jump to and extend for anything that needs to feel continuous.
  14. Music from Suno, spliced to the cuts. Voice cloned in ElevenLabs at stability 50, similarity 75, style 0.
  15. Four audio tracks: ambience, more ambience, hits and foley, music.
  16. Folders on disk before importing to Resolve. Assemble untrimmed, then cut.
  17. Shot match every clip against your favourite one.
  18. Export, then generate the poster from your own frames.

The pipeline here follows Graham Tallman's AI filmmaking course on Udemy, 37 short lectures and about an hour and a quarter end to end. The prompts above are my own rewrites of the ones he supplies.

thanks for reading. say hi →12 min read
⎇ main● 0 ⚠ 0UTF-8LF⌘K⌘F>_Ln 1, Col 1Spaces: 2Hiren · Ahmedabad