← All transcripts

EXACTLY How to Make an AI Short Film (Full Workflow) Transcript, AI Summary & Key Points

Youri van Hofwegen · Jun 01, 2026 · Howto & Style · 16:50 · EN-US

Watch on YouTube

AI Summary

A cinematic AI short film can be created in 16 minutes by building reference images before generating video, then linking clips with video references and assembling them in CapCut. The workflow uses character, location, and prop references to reduce the AI's freedom and preserve consistency. Higgsfield Cinema Studio generates cinematic images and video, supports character emotions, dialogue, audio, multi-shot prompts, and video references that carry lighting, mood, and tension from one clip into the next.

Key Points

  • AI created the short film in 16 minutes.
  • Building reference images before video generation gives the AI specific character, location, and prop information instead of making it imagine everything from a text prompt.
  • The workflow creates three types of references: characters, locations, and props.
  • Ryder is the main character, and Vance is an experienced sniper who helps take down the bad guys.
  • Ryder's appearance uses a stock-style reference image, while Vance is created with the AI cast feature.
  • Vance's AI cast settings use the action genre, a $200 million budget, the 2020s era, the sage archetype, a white male in his 40s, an athletic and tall build, facial hair, and a military outfit.
  • The location workflow creates both a clean bridge and a destroyed version of the same bridge so the destruction remains consistent.
  • The damaged bridge reference preserves the original bridge while adding broken steel beams, burn marks, and smoke.

AI in practice

Used for

Business ideas

Create short action films by first building consistent references for characters, locations, and props, then generating sequential video clips that carry visual continuity, mood, dialogue, and sound from one scene to the next.

For
People who want to produce cinematic AI short films without separately recording dialogue, managing lip sync, or performing extensive manual editing.
Solves
Directly asking a video generator to invent the characters, locations, props, action, and continuity at once often produces inconsistent characters, missing details, changing environments, and disconnected scenes.
  • The Ryder and Vance action short film: Ryder and Vance prepare for an armored convoy, destroy a bridge, fight across the bombed bridge, and finish with the bridge clear; the creator claims the characters remain consistent and the result feels like a real cinematic action film.
🔒  Build steps and tools for 1 idea. Unlock

Tools & resources

4 items

CNo. 0049
AIAINotes.us AI product

CapCut

capcut.com

CapCut is a cross-platform video editing application and web service owned by ByteDance that provides tools for creating and editing videos. It offers AI-powered editing features, smart templates, effects, and export options for desktop, mobile, and web for social media content.

Mentioned in
4 videos
Kind
AI
CNo. 0020
AIAINotes.us AI product

Claude

claude.com/product/overview

Claude is an AI assistant developed by Anthropic, positioned as 'The AI for Problem Solvers'. It is a general-purpose AI system used for tasks such as generating code prompts, refining requirements, creating advertising strategy and copy, and processing creative content like storyboarding and video prompts.

Mentioned in
86 videos
Kind
AI
HNo. 0028
AIAINotes.us AI product

Higgsfield AI

higgsfield.ai

Higgsfield AI is an AI-native creative platform for generating and editing images and videos from text prompts and reference images, with tools for cinematic video creation, visual effects, and AI-assisted content production. It also supports animating photos, creating shorts, generating voiceovers, cloning voices, upscaling, reframing, removing backgrounds, and using an AI agent to automate creative workflows. The platform is available on the web and mobile.

Mentioned in
20 videos
Kind
AI
HNo. 3367
AIAINotes.us AI product

Higgsfield Cinema Studio

higgsfield.ai

An AI filmmaking tool from Higgsfield for creating cinematic scenes and short films. It builds reusable character, location, and prop references, then generates connected video clips using those references to maintain continuity. The workflow supports emotional control, dialogue, background sound, video references, and editing of the resulting scenes.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of EXACTLY How to Make an AI Short Film (Full Workflow) — Youri van Hofwegen (16:50). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Youri van Hofwegen. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 This entire short film was created using AI in just 16 minutes. >> Everyone in position. >> Eyes on the lead vehicle. >> Now. >> Two down. >> Three more on the south side. >> [sighs and gasps] [panting] >> Bridge is clear. >> Previously, I used to think it was impossible, but after finding this workflow, I actually pulled it off. So, in today's video, I'm going to show you every step inside from avoiding rookie mistakes to keeping your characters consistent so you can make your own AI film from scratch.

01:08 Now, the whole workflow comes down to one principle, and I see a lot of beginners completely missing it. So, we need to make sure you understand this before moving on. So, what most people do is take their movie idea straight into a video generator and spend hours trying to create the perfect prompt. And after dozens of iterations, the result still looks pretty bad.

01:26 The details are missing, the characters change between scenes, and the model forgets to add things that were clearly mentioned in the prompt. So, they end up blaming the AI for these problems. But, the reality is AI models nowadays are 10 times better than what they were just a few months ago, and there are people creating insanely good-looking movies with them.

01:44 So, the model is clearly not the problem. The problem is that you're giving the AI way too much freedom. Even if your prompt is highly detailed, the model still has to imagine the character, the location, and every small detail on its own. And because it has figure out all of that at once, it starts taking shortcuts and changing [music] things up. So, what you want to do instead is stop giving it that freedom and take full control.

02:04 And that's exactly why the best AI creators never start with video generation. They spend the majority of their time building reference images first. Because once you have strong references, the AI doesn't need to guess anymore. It already knows what the character looks like, what the environment looks like, and what objects need to be included. This is one of the biggest secrets in AI filmmaking.

02:25 And if you do it properly, you'll get 80% of the work done in the first part of the workflow. So, for today's short film, we need to create three types of references. One for the characters, one for the location, and another one for the props. So, let's build the first one. For this, we need two characters. One is going to be Ryder as the main character, and another one is Vance, who's going to help me take down the bad guys in the movie.

02:44 For this, I'm going to be using Higgs Field as it has all the tools and AI features I need for the whole production. So, if you want to follow along, I'll leave a link in the description below. Once we're inside, let's head over to Cinema Studio. Now, the reason I'll be building everything here is because Cinema Studio has been trained on real movies only.

03:03 So, this means that no matter how simple or complex your prompts are, the results are always going to look cinematic. To create the first character, we need to make sure we're inside the image section. Here, I'm going to select auto instead of a specific AI model because it understands itself what I'm trying to build and chooses the best image generator for that specific job.

03:20 So, instead of spending time researching things like what a Nano Banana Pro is better than GPT Image 2 for this specific character, I just leave it on auto and Cinema Studio handles the rest. All I need to do now is paste in the prompt. But before that, I'll set the aspect ratio to 9 by 16 and set the resolution to 4K. Now, because I already know the exact look I want for the main character, I'm going to start with a stock style image.

03:43 So, I'm giving the AI that image as a reference to keep Ryder's look consistent. Not only does this give Ryder the exact look I wanted, but it also feels super real. Then, I paste in the prompt. Now, let's press generate, and here is the result we get back. Honestly, this is really solid, and that's because of the simple, yet intentional prompt I wrote for it.

04:01 Now, let's go ahead and generate the second character for the movie. However, for this one, I won't use the same method as before because I don't have a reference image of him yet. I just have a clear idea in my head of how I want him to look. But, when you're in this situation, one of the best ways to build your character is by using the AI cast feature.

04:19 So, let's go ahead and change from auto to AI cast inside the image section. And once you do that, you'll see the build your cast panel on the right side. Click on it, and you'll get all these pre-made options to work with. Now, all we need to do is go through each one and pick what fits our character best. The first one is genre, which sets the overall visual aesthetic.

04:38 So, things like action, drama, and even horror. For Vance, I'm going with action. Then, we have budget in millions, which tells the model what kind of production scale we're aiming for. I'm picking $200 million because I want it to feel like Hollywood production. Next is era, which sets the time period the character comes from. And this changes the wardrobe, the grooming, and the overall styling.

04:58 That's why I'm going with a 2020s as I want to keep my character modern. Then, we have the archetype, which is basically the character's role in the story. You have a lot of options you can choose from. It can be the hero, the innocent, or many more. But, for Vance, I'm picking sage because he's an experienced sniper. After that comes the identity. This is where you choose the gender, race, and age.

05:15 So, I'm going with a white male in his 40s. Then, we have physical appearance, which includes things like body type, height, and a lot of details you can set up. But, for this character, I'm just picking an athletic body, tall for the height, and some facial hair. Then, we move to details, where you can choose smaller features like tattoos and even facial scars.

05:35 However, I'm going to leave it empty for this example. And finally, we have outfit. Now, because Vance is going to be a sniper, the only reasonable option from here is the military outfit. But, if you don't find one that suits your character, you can always use the prompt box from the left side to add a custom outfit. Now, I can press generate and honestly, the result looks like he was pulled straight out of an action movie.

05:55 The face has that weight you'd expect from a real Hollywood casting choice. The beard has a nice texture, the outfit fits perfectly, and all of this without me writing a single line of prompt. Now, both characters are saved inside the Higgs Field library, but we still need to create two more references [music] before moving on to the video generation.

06:11 So, the next one we're going to build is the location, but for this one we're going to do something a bit different than we did with the characters. Instead of generating just one image, we're actually going to make two versions of the same location, but I'll explain why that matters in a second. For now, let's go back to the image section and change the mode.

06:26 We're going to switch it from AI cast to cinematic locations, and just like everything else inside Cinema Studio, this feature is built specifically for cinematic shots. [music] So, because of that, we don't really need to worry too much about adding complicated details for the lighting, the shadows, or the colors. All we really need here is a simple prompt.

06:44 For this, I'm using Claude to generate the prompt for me. So, now I just walk it through the location [music] I want, and here's the prompt it gives me back. Once I have that, I set the resolution to 4K, and then I press generate. And honestly, this is exactly what I had in mind. The lighting feels like that harsh noon sun [music] you'd actually expect to see in the desert.

07:02 And on top of that, the canyon walls in the background give the whole shot real depth. Now, this is the location where our action takes place, but later in the short film, the bridge gets blown up. Now, I could just let the AI video generator imagine what a destroyed bridge should look like, but that's where the problem starts because the AI doesn't really know exactly what I have in mind.

07:21 And this is where the principle [music] from the beginning of the video comes back again. We need to stop giving the AI so much freedom. Right now, if I leave it up to the model, it can just invent new things and decide on its own how everything should look. But what we want to do instead is take full control over the exact result. So, that's why we're going to force the AI to use the same bridge just in a damaged state.

07:40 To do that, I build a second version of the location, but this time it's the destroyed bridge. That way, I give the AI no room to add anything I didn't ask for. So, I switch back to auto because that's the mode that lets me upload an image and edit it. Then, I drop in my clean bridge as the reference image. The most important line in the prompt tells the AI not to change anything else.

07:59 So, I press generate and look at what comes back. It's the exact same bridge, but now it has broken steel beams, burn marks across the deck, and smoke coming out. But, before we can finally move on, there's one more thing we need to build, and that's the prop. For this short film, the bad guys arrive in armored cars. So, I need a clean image of one that I can reuse inside every clip.

08:16 For this, I'm switching the model one more time, and I'm going with Soul Cinema 2K because it's designed for clean hyperrealistic shots. And here's the prompt I'll paste in. Then, I press generate, and honestly, the convoy car looks like it came straight out of a real action film. And just like that, you know how to create any reference image. So, let's take these into the video section.

08:36 And as you can see, there's a lot going on here. It might seem overwhelming at first, but once you understand what each part does, the panel is actually really simple to use. So, let me quickly walk you through the director panel on the top. First, we have genre, where you can choose from multiple options like action, drama, and more. Then, right next to it, we have camera movement, which controls how the camera behaves during the shot.

08:56 And nearby is the speed ramp. This is very useful if you want full control over the speed of one specific scene. I usually use this to create slow-mo shots. After that, you have the basic settings like duration, resolution, aspect ratio, and whether the audio is on or off. The last setting is shot control, and here you have two options. You can go with smart, where the AI chooses the camera angles and timing for you, or you can go with custom multi-shot, where you take full control over how each shot is set up.

09:24 I've tested both, and honestly, smart gives me way better results than what I could create on my own. I don't know what they did to it, but it works super well. So, that's the one we're going with. So, let me show you how I use these to create the first scene. And on the left side, we have the reference media panel. This is where we drop in the characters, the location, and the prop we just created.

09:41 That way, the AI knows exactly which assets to use when it generates the video. So, I click on elements, and here you can see all the assets I'm going to use. Then I add them one by one, starting with the characters, then the location, and finally the car. But before I write the prompt, there's one really useful feature I want to show you. Next to each character we added, there's a small smile icon.

10:01 And if you click on it, you can actually control the emotion of each character. So, just like in real life, you can tell your cast to be either happy, angry, or act out any other emotion you want. And on top of that, you can also control its intensity, which makes the performance feel more natural or more dramatic, depending on what you need. So, for this opening scene, both characters need to feel serious because they're preparing for the enemies that are about to arrive.

10:23 That's why I set both of them to vigilance. After that, I go back into the director panel. For the genre, I choose action because the entire short film is built around an action sequence. I leave both the camera movement and speed ramp on auto, as Cinema Studio usually does a great job choosing these. The only settings I want to change are the duration, which I set to 12 seconds and 16 by 9 for the aspect ratio.

10:43 And I also keep the audio turned on, so the AI can generate the dialogue and background sound directly inside the clip. Now, when it comes to the prompt, this part is actually really simple because most of the hard work was already done with the image references. But you still need to be careful here because if you overcomplicate the prompt, you can actually make the result worse.

11:02 That's why I always use a framework I call the multi-shot framework. This helps me write the prompt faster while still giving the AI a clear structure to follow. And here's how it works. [music] Instead of putting everything into one huge paragraph, you split the action into separate shots. So, like shot one, shot two, shot three, and so on. Each shot covers one specific part of the action in the exact order it should happen.

11:24 So, the AI doesn't have to guess the timing on its own. And the second part of the framework is the dialogue. If a character says something inside a shop, >> 10 seconds out. >> Everyone in position. >> Eyes on the lead [music] vehicle. >> I just write the line in quotation marks. Cinema Studio reads that as a voice line for that character. So, I don't need to record audio separately or fix the lip sync later.

11:44 So, here's the prompt I write using the multi-shot framework. Then I press generate. And honestly, the result looks really clean. The convoys feel heavy and slow, almost like real armored vehicles moving into position. And both characters keep that vigilant look you would expect right before an ambush starts. Now, for the second clip, we're going to use a new feature.

12:02 And this is one of the most important parts if you want all the clips to feel like one continuous scene. The feature is called video reference. Now, the normal way most people handle continuity in AI video is with a start and end frame. You take the last frame from your first clip, use it as the first frame for the next one. And this helps keep the visuals consistent.

12:18 But, there's one big problem with that. The mood resets every single time, and the tension you built in the previous scene doesn't really carry over. But, the video reference feature handles this in a much better way. Instead of only giving it one frame, you can drop the entire previous clip, and the model will be able to study everything inside. So, this means it keeps the lighting, the mood, and even the tension from the previous scene.

12:40 Then, it uses all of that to continue the next clip in the exact [music] same direction. So, for the second scene, I take video one and drop it into the video reference field on the left. Now, for the prompt, I'm going to use the same multi-shot framework. I'll update Ryder's character's emotion from vigilance to rage when he's destroying the bridge.

12:56 However, Vance stays on vigilance as he's still waiting for the right moment to shoot. And I'll also add in the damaged bridge image to the references so that the AI knows how it'll look after the explosion. So, I press generate, and honestly, this is exactly what I wanted. When you watch the first clip go into the second one, everything carries over perfectly.

13:13 The lighting stays the same, the mood is still there, and it doesn't feel like the scene restarted. It just feels like one continuous moment. So, now, we can move on to the third clip where the action gets even more intense. This is where the most action happens in our short film. The convoy gets hit and the entire bridge turns into a battle zone. And this is also where I tweak one setting I don't usually touch.

13:34 I'm talking about the number of generations. You see, when you have a scene this dense, the AI has a lot more variables to handle. So, one single generation is way more likely to miss something or get one of the beats wrong. That's why for action clips, I always set the number of generations to three instead of one. This gives me three different takes of the same clip and then I just pick the best one.

13:52 Now, you can either set the number to three from the start like I'm doing here or you can generate one take at a time and only reroll [music] when something doesn't land the way you want. Either way, the process is the same. You watch all of them, compare the results, and choose the best one. Outside of that, I drop video two into the video reference field.

14:10 The location stays on bombed bridge since the explosion already happened and I push the duration to 15 seconds because there's more action to fit inside this scene. Now, for the prompt, I'm using the same framework and here's what I wrote. I press generate. And once all three takes are done, I quickly run through them. Honestly, all of them look good but I really like this one the best.

14:29 So, I save it for the edit. Now, for the final clip, the process is pretty much the same as before. I drop video three into the video reference field, then I bring the number of generations back down to one and set the duration to 13 seconds. For the emotions, I keep Ryder's character on rage because he's still finishing things up but Vance shifts from vigilance to trust because the threat is over and he can finally relax.

14:49 Now, for the prompt, here's what I wrote. So, I press generate. And honestly, this lands exactly how I wanted. >> Bridge is clear. >> So, now all that's left is to put everything together in the edit. And for this, I'm using CapCut. I drop all four clips in order, one after another on the timeline. But the thing is, we don't even need to do much editing here because thanks to the video reference feature, all the clips are already linked together really smoothly.

15:21 You can't even tell they're separate generations. The audio is also already in place since Cinema Studio put the dialogue and the sounds directly inside the videos. So, the only thing you really need to do here is cut out any parts you don't like. And after that, you can export the final video. So, here's what I got. >> 10 seconds out. >> Everyone in position.

15:44 >> Eyes on the lead vehicle. >> Now. >> [screaming] >> Two down. >> Three more on the south side. >> [sighs] [panting] >> Bridge is clear. >> Honestly, this came out really cinematic. The characters stay consistent across every clip, and the whole thing feels like a real action film. So, if you want to build a cinematic AI short film like the one you just saw, go sign up to HeyGen Field using the link in the description. Thanks for watching, and I'll see you in the next one.