C-Dance 2.0 is described in the videos as an AI video-generation system that creates scenes using character sheets, image frames, environment references, and voice-reference videos to maintain visual and voice consistency across scenes.
An AI image-generation model attributed to OpenAI that creates storyboard images, character reference images, and revised thumbnails from detailed prompts. The described outputs can include story titles, multi-panel layouts, character details, props, and other narrative elements.
Kling 3.0 is an AI video-generation product or model described as generating comparison video scenes from environment and character reference images.
Searchable transcript of 7 Updated Tips to Master Consistency in AI Video (Characters & Backgrounds) — Tao Prompts (18:03). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Tao Prompts. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 To tell a real story, you need perfect consistency in character and environment. Without it, you may find yourself looking a bit off. I've packed thousands of hours of testing into seven updated tips to keep your AI actors and sets identical in every shot. The foundation of consistency in AI video starts with the characters. If you generate the same character twice and they look like two different people, it breaks the immersion of the story immediately.
00:28 Now, it's pretty straight forward to keep the character looking the same in a close-up shot without a lot of motion going on. But if you're trying to generate a longer AI video clip where the characters moving around and if you see them from different angles, it's much harder to preserve the consistency. Luckily, character consistency is a pretty simple fix now if you do it the right way.
00:47 Before you start making the AI video, start by creating character sheets like this one. We'll use these as references for the AI video models so that they stay consistent in the videos. Here's an updated tip though. Typically, when you make character sheets of people, you have close-up shots on their faces and also full body shots that show how they look head to toe.
01:06 The problem here is that if you're zooming on the face of the character in the full body shot, it's usually a little bit blurry and hard to see all the details. And then when you use this as a reference for the AI video models, it doesn't know if it should use the person's face from the close-up shot or the full body shot. Then this can lead to lower quality results from the AI video models.
01:26 So instead, generate character sheets like this where you have a close-up shot on the character's face, but for the full body shot that shows their entire body, remove their face from the picture. That way the AI model knows exactly how the character's face should look with all its details. Now, I'm going to put all the prompts that I used inside this tutorial in the description.
01:46 So make sure to go and download that. To create these character sheets, I'm going to use the AI image generator inside of this tool called OpenArt. The AI image model I'm going to use is the GPT image 2. And I'm actually going to generate an AI character of myself. So, what I've done is I've uploaded a set of photos of me, and then I'm going to put in this character reference sheet prompt, which I'll leave in the description, and it'll generate this reference sheet of me that has a close-up shot of me and also a full
02:15 body shot without my face on it. Once I have my character sheets, I can start creating AI scenes for my story with my characters inside of them. To do so, I'm going to use my character sheet as a reference along with some shots I had of the environment. To use the character sheet as a visual reference, I just have to drag it into the prompt box. And I'm also going to use this other image, which has the visual style I'm trying to generate.
02:37 Then I just have to put in the prompt that combines everything together. So, I've generated a shot of my character inside this new environment. Now, we want to hold on to those character sheets because they're going to come in handy when we turn this shot into an AI video as well. So, I'm going to go and create a video. So, when I create the AI video, I'm going to put in this image frame that I just generated as what the first frame of the video is going to look like, but I'll also include character reference sheets as
03:06 well to tell the AI video generator exactly how my character should look throughout the entire video. The AI video model I'm going to use is called C-Dance 2.0. I'd recommend you using this if you want to maintain the maximum amount of character consistency across different shots. And then inside the prompt, I'm going to describe a scene with multiple characters inside of it of both the man and the robot.
03:26 And here's what the resulting video looks like. >> Relax. Nothing's going to sneak up on us out here. We'd hear it coming. >> So, the first shot is pretty straightforward. It's just a front view of the character. It's hard to screw that up. But when the AI video cuts to the next scene, scene from behind the shoulder of the man looking at the robot, this is where the character reference sheets come in, both for the appearance of how the robot should look inside the scene, but also how the man should look when we're
03:57 looking at him from behind. When it comes to consistent characters, keeping their appearance looking the same is only half the battle. The real test of consistency comes from how they sound when they talk. Your characters can look the same across different scenes, but nothing breaks consistency faster than if as soon as they talk, their voices sound different.
04:17 Now, typically when AI models generate scenes of your characters, it's going to generate a slightly different sounding voice each time. So, how do we keep the voices from the character sounding consistent? The first thing I do is I need to create a demo of the voice that I want to use for each character. So, what I've done is I've created a screen test of them talking on top of the character sheets.
04:37 >> same thing I felt the night before we lost the last team. >> When you're designing the characters' voices, you can customize them a bit in terms of the emotion and tonality as well. At this point, we don't have to worry too much about how the animations look. We just want to make sure we're locking on how the voices sound. So, what I've done is I downloaded those AI video files of my screen test, and I've extracted these MP3 audio files.
04:58 Now, what you could do is use these MP3 files as a reference for how the characters' voices sound, but for the specific AI video model I'm going to use, which is C Dance 2.0, and I think that's the best one when it comes to creating dialogue and voices for your characters, the voice reference actually works better if you use this blank black MP4 video recording.
05:24 So, I'll show you an example of what it looks like. >> Something's tracking us. It's not loud enough to be an animal, but it's not quiet enough to be nothing either. >> It's literally just a video file with a black screen and the character's voice on top of it. And we're going to use this as a voice reference to direct how our character sound. So, now when I write a prompt for the AI video model, not only am I going to tell it specifically what my characters look like through the image references, I'm also going to
05:53 tell it the voice that the characters speak with and add a reference to the specific voice file. So, for this specific dialogue scene, I'll tell the AI video model to start with a close-up shot of the woman from her reference character sheet image number three. She's glancing up warily and she quietly says, "I think something's tracking us. It's been following us for hours."
06:14 And the woman speaks with the voice from video one. Uh where video one is actually the character voice reference. By the way, to use your references inside of OpenArt, what you're going to do is type at to tag it and it will let you select from the set of references which you uploaded into the video model. Now, when we generate the scene, the characters will speak with our pre-made voices.
06:40 >> I think something's tracking us. It's been following us for hours. >> You've said that on every expedition. It's all in your head. >> It's been pacing us since the ridge, matching our speed exactly. >> All right. Let's just keep moving and stay close. We'll be extracted soon. >> So, that's how we make our characters look and sound the same across different scenes.
06:59 But, what about the larger environment? Character consistency is an important first step, but it doesn't matter if your characters look the same inside every scene if the environment's constantly changing when you change the camera angle or create a new scene. We need the world around the characters to stay consistent as well so that it feels like a seamless story instead of random AI shots.
07:20 There's a couple different methods to achieve environment consistency and there's pros and cons to each of them. The first method is to use what's called the elements feature or technique where you're uploading a reference shot of your environment along with the characters and using all of them as ingredients, mixing them together, and generating an AI scene.
07:41 So, let's say I want to create a scene of these explorers stumbling upon this abandoned encampment with some radioactive material. I could use a shot of that environment as a reference along with the images of my characters and also how their voices sound as well. And then inside the prompt, I'm just going to describe the relationship of my characters to the environment throughout the scene.
08:04 So, for example, in this line, I'll say that the man crouches behind a radioactive crate while tagging my image of the environment. When it comes to consistency in the AI video models, the actual AI video model you use matters a lot. There's a couple different AI video models right now that are worth using. One is Sega 2.0, which is the one we've used so far.
08:26 There's the Kling 3.0 video model and also Google's Omni video model as well. But, let's take a look at how the results from the different AI video models look. So, first up, let's see how Google Omni does for this specific scene I'm trying to make. >> Don't touch that. Just look. >> This camp's been abandoned a long time. Whoever left it left fast.
08:49 >> Fast enough to leave this behind? >> Right off the bat, I noticed that the color of the scene has completely changed from the initial reference I used. This happens in Google Omni a lot. It likes to create these much more desaturated, more bland-looking colors in my opinion. Now, the character movements, dialogue, and emotions all look pretty good.
09:12 However, it's not the most cinematic shot. Also, Google Omni is limited to only a 10-second video, whereas the other video generators can give you up to 15 seconds. Now, let's take a look at how the Kling 3.0 AI video model performs. >> Don't touch that. Just look. >> This camp's been abandoned a long time. Whoever left it left fast. >> Fast enough to leave this behind?
09:37 >> So, for clean through but not the colors look a lot more accurate. However, the acting and the dialogue is not very good. If you listen to the dialogue, my voice actually changes between the different lines I say. Also, the woman doesn't really speak her lines correctly. It's also not the most interesting looking scene. Now, let's see what happens when we use C Dance 2.0.
10:01 >> Don't touch that. Just look. >> This camp's been abandoned a long time. Whoever left it left fast. >> Fast enough to leave this behind? >> Out of all three, I'd say C Dance did the best. It preserves the colors from the reference image. It shows the parts of the environment relevant to the scene. And also the It shows the characters interacting with the environments.
10:23 And the dialogue also sounds pretty good. Now, it can be pretty frustrating when you're trying to generate that same scene over and and over and over again in your AI video model, and you keep getting different results or results that look inconsistent. You want the environment to match each other across your different scenes, but they don't and you end up regenerating the videos over and over again, wasting tons of credits.
10:47 There is a faster way to get multiple camera angles of the environment that uses just a single prompt. Using this method, we're able to generate an AI video that has a bunch of rapid fire clips showing the environment from different camera angles. And from what I've seen, this method is one of the best ways to get multiple shots of your environments that look consistent when seen from different directions.
11:10 This is called burst mode, where you start with a image frame of your scene with whatever character references that you want to put inside of it. And then inside your prompt, you're going to tell the AI to generate rapid fire shots of the scene and putting a list of all the different camera angles you're trying to recreate. And so as on my prompt, I've attached the shots of all my characters, as well as 20 different camera angles that I want the AI video model to generate inside of this scene.
11:40 And I'm going to generate all 20 of those camera angles inside just a 10-second video clip. Now, not every single one of these camera angles is going to look correct. Sometimes AI is going to make mistakes, but if I go and extract the individual frames from that rapid-fire video generation, what I end up with is a collection of shots that show what my environment and the scene looks like from all the different camera angles that I want.
12:08 Then, what I can do is use these image frames that I extracted as references to my AI video generator once again to create a longer scene. So, you know how to generate consistent shots of your characters, as well as the environment. Now, if instead of just generating a collection of shots and camera angles, you want to add in more of a story component where the scenes flow between the different frames, sometimes it can be a good idea to create a storyboard first before generating the AI videos.
12:36 Now, here's a couple examples of 12-panel storyboards that I've made, and each one with a story that flows between all of the panels. What these AI video models can actually do is go and animate multiple panels of your storyboard all simultaneously at the same time inside the same video clip. And when you generate a storyboard that has all of these panels inside of it at the same time, the AI models are going to do a better job of preserving the consistency of your characters, as well as the visual style across the
13:07 different scenes. This way, I can keep my characters in the world consistent, but actually tell a story inside of my AI video. Here's a tip, though. Don't try to animate all 12 panels at the same time. That's way too much information to put inside of an AI video generator in one go. I'd recommend splitting the full storyboard into rows of four panels or less and then use your AI video generator to animate these rows of four panels.
13:35 This way I can fit four different shots into that same video generation. Now, I'm going to put a link to this platform I used to create the AI video models down in the description, so you can go and create your own AI stories. So, one thing that AI video really struggles with is creating big motions and movements. One tip that I always give is that you want to create scenes with a small amount of motion and movement inside of them if you want to preserve the maximum amount of consistency.
14:03 It just leaves less room for error. But, what if for that specific scene you want to create you need a big camera motion. You need the characters to move around a lot. How do you keep the environment and the characters looking consistent throughout that entire scene without them drifting? If that's the case, it's not enough to just use a single image frame as a reference.
14:23 Instead, you need an image frame that defines how the beginning of the video looks, but also how the last shot of the video looks as well. This is called start and end frames. We're reset two image frames for the start and end of the video and tell the AI model to fill in what's in between those shots. This way you can create a lot of camera movements without the world completely losing consistency.
14:45 To do this, go to this start and end frames feature and inside of here you have the option to drop in a frame for the start of the video and also the last frame of your video as well. For the specific shot, I want to start with this image of the woman who's trapped inside of a spider web and I want to generate the shot with a camera rotates all the way behind her to show the skeleton that's also trapped in the web.
15:11 When you're using this method to generate big movements and motions, one thing I would recommend is putting this phrase that says single continuous shot into your prompts. This just makes sure that the AI video model is actually able to generate that full continuous movement without splitting the video generation into a bunch of different friends. The bigger the movement you're trying to create, the more I would recommend using the start and frames feature.
15:38 It's just a great way of locking in that exact visual style you want to maintain through the entire video clip, as well as how the environment and the characters look. All the techniques I've shown so far for keeping your AI work consistent will come in handy at some point in your AI video journey. But, there is a final fail-safe technique. It's one of the most fundamental methods you can fall back on, even if all the other methods fail, and that's chaining together image references to create a consistent world.
16:05 So, we start with an image generated of a scene, and then to generate the next shot, we take that image from the first scene we created, and use it as a reference. So, if I start with this close-up shot of me as a hiker, I can use that shot as a reference for the next image frame I want to generate, which is going to be this over-the-shoulder shot of me looking at a bear.
16:29 Then I can use those first two shots as image references again to create a third shot, this time of me facing off against the bear directly, looking him in the eye. By chaining the different image references together, you can create a complete sequence of shots. And then each of the individual shots can be animated one by one using an AI video model.
16:53 This method does feel tedious compared to some of the quicker ways I've shown in this tutorial. However, if you want to get the exact results you want with the camera angles, the way the characters are arranged inside of a scene, generating the frames one by one like this is the way to go. Here's another example where I'm starting with this front view image reference taken from in front of the woman, and inside the prompt I'm going to tell it to rotate the camera to an over the shoulder shot behind the woman to show an
17:25 old skeleton. I'm also going to use this other image reference right here as a style reference and tell the AI video generator to use the same visual style as that image. And that's how I'm able to generate multiple consistent scenes of this world. If you want 10 more tips for how to create hyperrealistic AI videos, go watch this tutorial right here.