Blender is a free and open-source 3D creation suite developed by the Blender Foundation. It provides tools for modeling, rigging, animation, simulation, rendering, compositing, motion tracking, video editing and game creation. Blender is distributed under the GNU General Public License.
fal is a generative media platform for developers that provides a unified API and SDKs for running image, video, audio, 3D, and code-generation models, including open models, user-provided LoRAs, and custom model endpoints. It offers a gallery of production-ready models and runs inference on a globally distributed serverless GPU infrastructure that can scale deployments from zero to large numbers of GPUs. fal also provides on-demand GPUs, compute for training and fine-tuning, private model endpoints, dedicated clusters for custom workloads, and observability tools for monitoring deployments and usage. The platform develops and post-trains open-weight models such as MiniMax H3 to improve generation speed, cost, and control for real-time video applications.
H3 Max is fal’s post-trained video-generation model, based on MiniMax’s open-weight video model. It combines model post-training with systems and hardware optimization to reduce generation time while maintaining video quality, supporting fast generation, interactive streams, and professional video workflows. Its continuous-video experiments can retain previous scene context and respond to new directions while a stream is running, with development focused on controllable camera movement, lighting, characters, motion, and lip sync.
H3 Max Director is a public AI model for action-controlled continuous video. It maintains memory of previous scenes and accepts live prompt-based direction while video is running, enabling control over elements such as camera movement, lighting, characters, motion, and lip sync. It is described as fal's post-trained version of MiniMax's open-weight video model, combined with systems and hardware optimization for real-time generation.
H3 Max Turbo is a faster variant of fal's H3 Max generative video model, which fal post-trained from MiniMax's open-weight video model. It is designed to generate a five-second video in approximately 1.5 seconds, with a small quality tradeoff compared with H3 Max. The model's speed supports experiments with real-time and continuous video generation.
Twitch is an interactive livestreaming service for content including gaming, entertainment, sports, music, and continuous AI-generated video streams.
Wispr Flow is an AI voice-input and dictation tool that uses speech recognition and contextual editing to convert spoken commands into formatted text for writing, messaging, and interaction with desktop applications.
Searchable transcript of How Real-Time AI Video Is Changing How Creators Work — a16z (39:04). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by a16z. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Generative media is along with coding agent market what we call is token market fit. Everyone's waiting for a large consumer moment in AI. I believe H3 Max makes it possible. >> Were you surprised by the speed up and the gain you could get from post training this model? >> We have a version called HDMax Turbo that's public that can generate like a 5-second video in like 1.5 seconds.
00:21 From a cost standpoint, it's also like 2x less. People starting creating these beautiful scenes using an LLM model GPT Astra in Blender and all of a sudden it unlocked a whole new workflow for Hollywood and professional people. >> We have been very very focused towards speed, performance, quality and now we have a really good base model. The next month or two is going to be fully focused on Welcome Gork Banan to our podcast again.
00:51 Uh we did the last one last year. This is long overdue and we have such an exciting model to talk about which is false um H3 Max. Um the day when it came out I was calling it it's really in the league of its own. Like it's so funny to see the benchmarks where you have like you know the dot of this model on the far left or far right and then everything else is like on the other half >> and that graph is actually log scale.
01:12 So it's actually further but we had to to to fit it in we had to do log scale. >> That is hilarious. Um and and it you know the >> time the time portion the quality is not yeah >> for sure the internet noticed for sure there are so many viral tweets about it like people really played around with this model. Maybe just give us the backstory of what inspired you to uh post train this openweight model from Miniax and how did you get the quality and speed to where it is?
01:45 >> First of all the Miniax H3 model is is the first truly opensource very capable like latest generation video model out there. So even though we work with some of the other model labs to run inference for them, we never had the had had this this capability like had the right to add this capability on top of it. So when Minimax came up with their open very capable open source model that is is truly last generation can take references like very familiar architecture to to any other video model.
02:19 We thought this is a great opportunity to go all in and and see what we can do. And again, we did like many different things that we are going to talk about like that that combined gave the results that that you show on the graphs. But yeah, the the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open source.
02:43 What is the idea like given like you know FA has been known to be like a generative media inference serving platform like what is the idea to get into post- training openweight model like you know you talk quite a bit about uh about it in the in the blog of like combining the system work with you know the the model uh itself like maybe talk more about um the work behind that >> like generative media is is I would say along with the coding agent market What we call is is token market fit and the way we define it is as
03:18 can a single person productively spend a lot of tokens and and the the amount is like 10k a month something like that. So there is incredible amount of demand in the market to to generate video to to generate many things at the same time and a person who is doing this for their daily job they spend in front of a computer and do this all day long and they spend thousands of dollars lots of tokens um and since around April the whole industry and file itself we've been compute constraint we we are growing as much as we
03:56 are adding compute Like there are things we do here and there but the whole industry has been compute constraint and we we've always been looking for efficiencies where we can relieve that a little bit so people can can use this more. So that has been the the idea behind everything we've been doing since April and this just came at the right time because this makes everything maybe an order of magnitude more efficient.
04:24 So gives more compute for for other other models or even like more tokens can be generated using using HDMax. I think would would agree on on that like >> yeah like just from a systemwide optimizations which is what we have been doing for the past three four years you can maybe make the model 2x 3x faster while producing the same quality right because it's at the end of the day same model same architecture you have the same constraints you're just trying to optimize what uh what what you can get out of an what you can
04:56 get out of the chip itself and there there is a roof line there and like you know we we have been approaching that roof line more and more uh especially like you know lately because our entire team has been focusing on how do we get out more video pixels from a single chip uh as much as possible and this this new set of like post-training related optimizations with like you know system/model code design enables us to go beyond that roof line by an order of magnitude and like we we just like felt the pressure we have
05:29 been working on it on top of open source image models For the video models, we did one version with ideoggram. We did one version with flux. So we have been we have been like experimenting with how do how can we build post training infrastructure to take an existing model build kernels and systems design around it to run it very very fast for a specialized version that can beat anything else that we would get just by running the model itself.
05:54 And you know combination of that plus you know just like getting a frontier video model on our hands and all this expertise we were able to you know go by like an order of magnitude in terms of speed. >> Incredible. Let's let's dig into that. Um I may get some of number numbers wrong but um uh the >> there's like efficiency numbers there's cost numbers there's speed up numbers like not not everything means efficiency but it all adds up to be very efficient.
06:21 Yeah. >> Yeah. I guess what what is stunning to me is like there is like a magnitude lower cost and also like much faster. I think it was like 35x speed up >> compared to the original minimax and point. >> Well, at ELO score you didn't really sacrifice quality. So yeah, just maybe reveal a little bit more of the the the secret sauce behind of like is this more of like um the the the type of system work you have done that like did you have to do like model architecture change like is it some system work that really
06:52 brought down the the um uh the cost and and latency and like >> how about like the the next generation of chips like uh GB200 fits into the the the whole story >> it it it's just like compounding effect of like multiple multiple different optimization variables that we have been targeting. The first one is obviously okay you go from you go from like the base model to a model that's like post train to be like more efficient uh for for diffusion models this is just essentially how do you go from like running 50 steps to
07:23 running something like 20 steps right like you're just like trying to optimize that pipeline but as as soon as you go from 50 steps to 20 steps you lose quality so you you need to target in the optimization scene okay I want to improve the quality and then I want to apply the optimization so we have like checkpoints of this that are significantly higher better quality but obviously slower.
07:41 So what we initially did was okay let's run our post training and RL pipelines so that we can improve the model's quality and then apply optimization stack on top of it so that the end result gets you to the same quality or like even like higher quality than the original model but at the same time you're like an order of magnitude faster. So most of the gains come from like you know post- training this model uh to be like you know compatible that you can run this on like less amount of steps but on top of that you add
08:11 like all the kernels and systems engineering work that you do that brings your like hardware utilization from like 30 40% which is like standard in like many inference workloads to like 70 80%. Like you're essentially trying to and 70 80% on theoretical MFU which is like impossible to reach. So you're essentially at the roof line of what you can get out.
08:28 And then these models are not just like a single oh you just give a prompt and you get a video back. There's like they're actually pipelines underneath. You need to get take a prompt. You need to you know like use a run a LLM like a very large LLM to go expand that prompt to a format that the model was initially trained at uh run the generate the video in the latence space and then decode those latency back into pixels and then depending on the workload there might be an upscaling component involved.
08:54 So there's like multiple components and every single component by default is unoptimized. There's still like lots to be gained there and like we just like looked at it from a perspective of we are going to get the maximum out of every single component. This made us go around LLMs at super high speeds right there. There there is that component but for a different workload.
09:13 This is not like something like an agent decoding LM workload where you have very high cache rates where you have higher sessions. It's a single shot. You take you give a prompt you get a prompt back and there's no caching. you're operating at low batch sizes. So there's a completely different set of optimizations on on the prompt expression side, completely different set of optimizations on the diffusion model, completely different set of optimizations on the VAE that you take from latence to pixels and you just
09:36 combine all of these to to have an effect that compounds. From a hardware standpoint, going from something like hoppers to black walls, you see something like 2 to 3x improvement by by itself. But from from from a cost standpoint, it's like pretty comparable because the cost is also like in in that league. So it I would say it only reduces your like wall clock time but like not just the efficiency itself.
09:58 But it it obviously helps if you want to go significantly beyond real time like if you want to generate you know 5 seconds of video in less than you know 3 2 seconds then you need like some of these like latest generation hardware uh today uh to to unlock that possibility. Maybe this is a detail of a question like is the model being served on single GPU or like it's >> it is like majority of the video models today run in a single node configuration which is 8 GPUs because once you start scaling beyond 8 GPUs the
10:26 efficiency gets less and less because of the communication overhead uh and com existing both like the existing minimax H3 endpoints as well as like other video models are probably getting served at like you know single node configuration same with this it's like running on in parallel across eight GPUs And do you think there will be more efficiency gains in there that you can either you know optimize more of the steps in between by sacrificing maybe some of the um like narrow down the user experiences let's say like
10:56 the different type of inputs and outputs or like as you're think of parallelism like is there more juice to squeeze maybe that's the >> we released a turbo version of H3 Max. So initial idea was calling this H3 Turbo and like we're like we don't want to call this Turbo because the quality is like better than the original one, right? Like that's like this this needs to signify how good of an achievement it is.
11:17 So we released H3 Max but like a week later we had like you know like our team was like we can run the 2X faster at like 97th percentile of quality. Like we run they're like almost the same, right? like there's still like there's like a notice like there's a small noticeable uh loss in quality, but we we have a version called HDMax Turbo that's public that can generate like a 5-second video in like 1.5 seconds, which is like insane.
11:39 And that's also like 2x the 2x like from a cost standpoint, it's also like 2x less. So there like depends on like how how okay you are with like losing quality you know you can go down and like today these models are so cheap and so fast that I don't think people need any faster or like any cheaper like it's already like at at a point where uh the the from a cost standpoint compared to the frontier itself it's an order of magnitude cheaper compared from like a speed perspective it's more than an order of magnitude
12:09 faster and like you just like enable all the experiences I think we would need to Like I think we would need to see what else levers that people would need but I my bet today is we just need to improve quality more than like the speed at these speeds right like let's fix the speed and let's try push for quality and controllability of these models which is like you know what we have been pushing in the past two 3 weeks I think controllability is is key like when we first did it we did um text to video and then image to
12:37 video and then references came later which which adds a ton of controllability ility and it's basically the default mode how people use these models these days references and then we are now adding different uh Laura's fine tunes of of the base model as well we are working on like a lip syncing version we are working on a different different camera angle Laura different style Lauras so again open source adds a whole ecosystem around the model and it really really helps >> were you surprised by the speed up and the gain
13:14 you could get from post-training this model like like I I I saw it as a little bit of a surprise that like one day I think it was a Saturday I'll launch the model and the Sunday people put it on Twitch it become a real-time model like I think that's the interesting part of like when you >> we we did like we spent a ton of money doing on our own I don't know like tens of thousands of dollars even and like the results were unbelievable and and then like the plan was to just release the model without doing external emails
13:45 and then okay we decided let's let's hold off let's not let's not tell people that this is like so much faster and so much better before we have some external validation so so we waited like 3 4 days to uh all these other like eval platforms to actually run the eval so we we matched the results that we have externally as well and that's that's how we launched it because as you said the results were a little too good to be true and And it was >> I guess were you taken by surprise the the real-time uh use case that came
14:17 out of it or what are some examples that you think this model has like unlocked of the experiences the prior models could >> this happens at fal once in every couple of months where like the whole company gets gets hold of something and and the creativity just explodes and everyone is just working on a a new little app or or a a different optimization Laura whatever it might like the whole company gathered around this this model and like some some front- end engineers started working on like interesting applications uh
14:56 we can talk about our like world model accelerator team which is which is brand new they they started working on the the live uh experience >> RTC >> the web RTC live experience so like there were five six different parallel little projects within the company and like I I think we broke a record on Slack that day how many messages were were sent in the company because like and and and like we have a distributed team.
15:25 We have we have people all around the world like mostly in San Francisco but it's it's like incredible when you see like the 24-hour development like when people like work 16 17 hours and then someone else wakes up and picks up that like and that went on for like three four days and that's that's when we released all these projects. >> Take me into that.
15:47 It's so interesting cuz um like you imagine like a model or product launch being like planned out like having again like having all these like eval vendors being ready lined up and you know ship something out and then like you let the the world um or the external users take it and then experiment and like build experiences put online. Yeah, it seems like you know people internally who are very creative just like took this drop everything they were doing like launched a experience that got really popular on Twitter.
16:16 Do you want to tell us about that one? >> Yeah, of course. One of our engineers, Rahan, um just just by himself completely um started streaming a live stream of continuous generations of H3 Max from his laptop like he was his computer. >> His computer. Exactly. He was doing some like prompt tricks trying to keep like a coherent story and then like he started live streaming that on on Twitch.
16:41 In parallel, Level Zio, famous Twitter influencer at this point, had a had a similar idea and he reached out to us that he has a website ready. He wants to like host the streaming himself and and have a website that does infinite streaming internally. Also, we had another team who was working on a continuous version of HDMax. So, HDMax is like the the Rahans version and and levels.io version were independent clips.
17:14 It's still very fast, but uh the clip starts, it ends, and then you take the last frame of the of the clip, try to put it in the next one, and try to create a contin the second clip doesn't really remember anything from the first clip other than uh the last frame. But internally, we were the the ML team was working on a version where it's the the transition is more seamless.
17:40 there's like two minutes of memory. So like you're in a scene and when when you direct the the model or someone else enters the room, it actually like everyone looks at that person entering and the scene is continuous. So internally we were working on that and then another team was working on an experience we called file live for the continuous version.
18:00 So we had three parallel efforts going on that were all independently going viral on Twitter. >> And these were all like spontaneous like you didn't like >> you didn't plan for it for any of them. Exactly. And they just became products and experiences in the next in the following days. >> But let's talk about the how how we made the model more continuous.
18:22 That was very surprising to me because I've I've never seen uh that actually work on on a video model before. So yeah we so going back we have been like very very focused towards role models and essentially like action controlled or like you know action-driven real-time continuous streams of video and the problem till like you know something like HDMX was quality was not good enough at all.
18:47 It was just like you know you it degraded a lot. It didn't remember the past before but we built the infrastructure. We built the infrastructure that we can like go stream video, have people control it in real time, being able to like multiplex it to multiple people, very low latency. And at the same time, our ML team was essentially trying to take every single video model and try to apply these set of like optimizations and tricks to okay, how can we make this gen instead of a 5-second video, 15 second video, 30
19:14 second video. uh but you were always like below the real-time factor uh where you know you you were always like you know you like you never could generate like 5 seconds under 5 seconds once HDMax unlocked it the ML team was like this is insane which are like separate teams internally we have a research team we have an inference team we have an ML team they're like they they saw this and like this is insane we can apply all these like set of learnings that we had in previous models where we attempted to do this where
19:41 instead of trying to generate a 5-second chunk that's trying to generate you know like a 50 like 10second video uh and then the 5 seconds from previous one is still attended. We still remember it. And like as the video goes up, we can like extend that memory up to 2 minutes. And you need to do extremely clever optimizations because attending to a 2-minute video is just extremely extremely comput intensive.
20:02 It just like it goes up exponentially uh uh from from like a comput standoint. So like we we we did like lots of optimizations there. But at none of the state we were able to okay we can remember back to two minutes which is like generally good enough from a memory perspective and then obviously with like prompt tracks you can still like continuously uh remember more finer like grain details above the 2-minute mark and you can essentially stream infinitely.
20:26 be capted at an hour uh from that perspective and then that team just like released that model under H3 Max director which is public for people to use and I think it's the only model that can generate like you know up to 60 minutes continuous videos that is action control you can like you know start with a prompt say uh like there's like an office setting and someone is like you know working and then like 30 seconds later it just imagines by itself 30 seconds later you can like say a woman walks in through the door
20:52 like it it can take the prompt and reflect it immediately which is the most fun part. >> And the office is still the same office. The camera can pan back to the original person and the original person is still there in the same state. Yeah. >> So you know we we we released that and it it got like we we we did this like five website to just like demonstrate it because it's like people need to see how cool this is, right?
21:15 This is a new technology. I I don't think people are like really aware and it it got also like very viral immediately because we also let people vote on what the next action is. It was like you know uh like a form of crowd source to like the chat was controlling what whatever was happening which which is fun but obviously you know we we limited on like the options and then they could pick oh like a banana enters the office instead of a moon and it's like more fun and like people start like you know having this and we
21:39 start adding more channels and like every channel had a concept. There's like a channel where it's like full chaos. There's a channel where it's like cartoons from like ' 80s and like the model is like extremely capable and it just like remembers like so many different concepts and there's like it has like a big big memory uh from from like a styles and like you know concepts perspective.
21:57 So it it just became like a very fun experience underneath. >> Again like there's so many really incredible experiences coming out of this like H3 Max director was just another huge surprise to me. It's like I I I found it interesting in the in the gem media market that you it's not like you know like language model you have like this linear graph of like just continuously compounding on like you know intelligence capability and so on like feels like in the field you're operating in it's always like a few months of
22:28 like sort of quiet time but like a lot of things are bubbling but like in a very short period of time like everything bursts like all these things in com in combin combination come together of like the the um uh base model being good enough like you can get the latency down to the point where you can like >> references. Yeah. >> Yeah. um like get the real time experience but also like apply controllability on top of that real-time exper like this just opens so many you know opportunities of like live experiences where
23:01 like end user can control what's happening on the screen which is incredible like uh we have imagined a lot of these experiences but never been able to like really play around with it. Maybe just like tell us more about what you're seeing from the market of like how are people using like um uh the director uh capability like what are you uh seeing creators are creating that you haven't seen before uh and what do you think that unlocks as far as you know what >> people can do with the this this medium?
23:31 Yeah, it's been like almost three weeks since we released H3 Max and already it is the most popular video model on the platform on on the file platform by like double almost like little more than double >> um in terms of like volume. Um so in in a lot of other platforms it's also becoming uh the default model that people interact with because it's so fast so cheap it just makes sense if if you if you come to a platform this is the experience that that you want to see.
24:04 Um so in terms of like popularity and volume it's it's taking over at least from our uh vantage point. And for for Max Director again there has been I don't know tens of different versions of these live streams. Some of them are uh still going on and like becoming more and more popular. We are trying to work with some like AI IP holders people who have like AI shows on Instagram and Tik Tok and and do trainer Laura on on their style and do a live version of their show.
24:39 So uh we have couple lined up already. So that's going to be very exciting. Um and like the way people like if you talk to uh a creative technologist prompting with voice has already become uh something that like they use all the time is like using whisper flow or the chatt voice mode. Yep. and and now like you can keep talking to the model and it it like almost as if it's a real director in a real movie set uh directing like the camera directing people where to go.
25:11 You you can do that and like our creative engineers started using these models like that. So we'll see like a lot of interesting experiences are built as we speak. Very interesting. As in like the video is playing. >> Video is playing you are like talking to the video and uh what what's what's being displayed changes accordingly. Yeah, >> that's incredible.
25:33 Um, and talking about like how the the memory piece holds now like are again this may be a technical detail like the capability of remembering what happened in the last scene or in the last couple minutes of scene like are you remembering that through like the the the frames the images or is it like through text? It it essentially um no it's it's essentially the like it remembers the raw video obviously very very compressed because you can't attend to the full video but it essentially knows like most of the like the
26:02 details happened in the past 2 minutes from its own generations and above the 2-minute mark it has like think of it as like a it has a evolving system prompt on top of the 2-minut mark from 2 to 60 minutes where it knows like the overall structural detail. So it remembers like the last few scenes. If you think a scene is like 15 30 seconds, then it remembers like the last four to eight scenes and then on top of that like there's like a continuously evolving gradually evolving system prompt that like keeps remembering
26:31 the like the overall uh coherence of the of the world. Like everyone's waiting for a a large consumer moment in AI. Now it's like good enough and and cheap enough that like a truly novel social AI experience can be built on top of it. Maybe let's talk more about the e economic side of this like what is the um uh I guess one just like talking about serving cost for like same uh minutes of video uh with it's it's three max um and how has it changed your thinking around like your footprint of like inventory of chips like
27:13 how do you want to have um like different uh steps of experiences serving to the end user >> but I mentioned this a little bit Like everyone talks about how uh complex the next generation LLMs are, but video models are actually very complex as well because the pipeline has different components and sometimes they require different hardware configuration for efficiency things like that.
27:37 So the if you were to do this even more efficient, let's let's go maybe even cheaper, uh we would probably run different parts of the pipeline in in different types of hardware. Um, another interesting thing would be to um run it on consumer hardware for people to run it in their in their own machines at home like optimizations don't translate 100% but translate somewhat somewhat close to that and then we can do extra work to translate more of it.
28:11 Um so doing these optimizations in different types of hardware and combining the pipeline in a way that it's even more efficient. Uh I think that's what we are going to do in the in the next coming weeks. >> Amazing. So you will have uh people like Rohan that can stream partially of the experience from from his computer but also having like the director and the the control plane more living on the >> Exactly.
28:37 Yeah. >> On the cloud makes sense. So we talk about all the consumer experiences um this model could unlock. Um and it seems like Botwan is happy with all the efficiency uh like squeeze out of the GPUs. Now we're talking more about how do we like improve quality and controllability of these models so that like you know the high end of the market the Hollywood creators uh directors can can take this to the next level.
29:00 Um I saw some demos uh coincidentally like you know this model came out of uh came out the same week or week prior to Astra. People were combining the Blender experience with uh H3 uh Max from from FA like um talk about how it's going to impact the the the Hollywood world. >> You using Blender with one one of these AI models together is a extremely popular workflow for for professional work.
29:29 Basically, you you you render a low resolution of your scene, what you want to do using using Blender previous like nonAI technology, and then once you add that video as a reference to an AI model, you basically get close to 100% controllability. Um, and this this is an in incredibly popular um workflow for BFX artists, people who are doing this professionally because they want exact they they want to get exactly what they put in into into the model and as as you mentioned uh a week after we we launched H3 Max uh
30:06 people starting creating generating uh these beautiful scenes using an LLM model GPT Astra in in Blender and all of a sudden it unlocked a whole new pipeline using LM to create a Blender scene and then passing that to the H3 Max model or or any video model. But it's it works very well with H3 Max because it's extremely fast and you can like try many things all at once in parallel and that that unlocked the whole new workflow for Hollywood and professional people and it gets you to like close to 100% controllability.
30:40 As as I said, we have been very very focused towards speed, performance, quality and now we have a really good base model. I think the next month or two is going to be fully focused on okay how much controllability we can add to these models so that professionals at studios, professionals who want to actually produce like produce content that fits their use cases perfectly can leverage these models.
31:06 Uh the team has been working on an amazing you know like a lip synchronization model where you know you can just supply the audio you can supply your your uh you can supply like a video or an image reference and then it can like synchronize the lips perfectly. Same with like motion controls. Uh you can just take a motion of someone dancing and apply it to like your uh AI generated character and it fits perfectly.
31:29 And this like you can like get these results with like basic prompting and you're going to get like 80% 90% reliability. What we are targeting is like 99.9% reliability in the outputs so that you can actually trust the model every single uh aspect of this generation perfectly and that's like what we've been pushing. Uh one big launch that we had last week was the camera controls uh which is essentially you can direct where the camera is going within the video perfectly to the to to the degree and >> and this is by like
32:01 describing in the prompt or like generating the >> you just essentially like underneath you give a JSON of like I want camera at like 0000 at t0 I want camera at like 90 degrees angle at t1 like you essentially supply a a structured uh structured description of where your camera needs to be at any point in time and then the model is like perfectly conditioned to to regard it as like the only source of truth and it doesn't like hallucinate uh on like where the camera should go and it just like you can essentially
32:33 reconstruct 3D scenes from a single input like because the model itself is like very good video model but at the same time you know it's like perfectly adheres to to the camera itself >> and this is because uh the base model itself already has the understanding of the camera angle that you It doesn't respect it. It just under like you need to tune them all.
32:50 You need to tune the model to a significant degree. And this is like what enables like at large scale you know post training infrastructure. We now have the infrastructure to take H3 max add any capability to it. Same applies for any new model right if there's a new video model. We we we we essentially spend most of the time building it as an infrastructure than just like one-off training runs so that we get we can build like services around this for not just like you know open source models but for like frontier close
33:16 source models as well because we see in the market this is like the biggest gap is just how controllable these models are like first we start with text to video where you put a prompt you get a video back was good but like you never could describe the perfect character for you and we had image to video where you know use an image editing model and then generated like you know the first scene and then the model was like obviously much more fitting but you still couldn't like say oh I want this new character appear at
33:41 like second three you need to put it to your first frame or like you can't like you can prompt it but this was never perfect and then we added reference to video where you know uh you can provide like an initial starting frame and you can also provide I want these characters with these voices like you know that's also like a big big unlock where you can essentially say this is the this is the voice for this character and now like you know we are adding oh within the scene I want camera to look at this degree at like t0
34:04 on camera to look at this degree at like t3 and then we're adding lighting controls where you essentially say where the light is coming from. These are all all compounding on top of each other and like we just have the unified infrastructure to just apply this to any model at this point. >> That's incredible. Hollywood is our fastest growing segment and um there there's a lot of noise about how AI might distract Hollywood but Hollywood usage was non-existent a year ago and in in the past year it grew and now it's it's
34:33 the fastest growing segment like um Amazon MJM studios in in their conference they they released their NAR uh tool it's it's mostly backed by uh file infrastructure behind the scenes and we are seeing an incredible incredible pool coming from Hollywood and it's exactly what they need these like small point solutions rather than generating everything from scratch.
34:56 They want to be able to extend the video a little bit. They want to be able to change the the camera controls. They want to change the lighting and someone has to build these solutions for them. What what Hollywood needs and what the creators actually need and what the research labs are working on. there's a little bit of a disconnect there and we believe we can come in and do these little post training projects to to close that dep gap because we work with all the Hollywood studios and we hear from them what what they
35:28 need and these are exactly the things they need these small point solutions that actually make them more efficient push out more more video and uh AI can actually close that gap very nicely >> maybe say in a little bit different way like we have been staring at this problem for the last three years as well. Like we see like companies trying to like build a you know uh movie director like video model like by uh either pre-train or post-train on on the video side but what I'm hearing is like different people expressing
36:02 the way they want the output to come out very differently. consumers talk about it and then like write the prompt and generate the results very differently from a Hollywood director which is obvious right like professionals want to talk about like you know these camera angles they want to talk about the lighting like you sort of have built a library or like a a collection of >> of um post-training like I would call it data and toolkits that can apply these on any model that you can like grab the weights on so that
36:37 they're adapted to like a different audience where they can express their creativity in a bit different uh fashion to control the model when it unlocks a lot of the capability underneath >> and half the problem was capabilities of the of these models we are solving that the other half of the problem was legal and data residency things like that we we we made a ton of progress there as well uh we now have have a system of people can apply with with their own IP and we unlock their own IP in in the models and we are
37:11 going to grow that and that's going to be a very powerful thing we do with Hollywood studios also we we now have seance US hosted as well we already had previously other Chinese models seance was the missing part every Hollywood studio wanted us to have it US hosted now that's available and so there are no no obstacles because in front of these Hollywood studios now everything is ready and we we believe they're going to 10x 100x their AI usage in the coming months.
37:45 >> It's such a exciting world for uh movie lovers, consumers, people who consume a lot of video and and creative content >> and uh we have our conference generative media conference next next week. Uh this is our second time we are doing it. Last year it was mostly consumer AI. There were like maybe couple Hollywood executives here and there just curious about it.
38:08 And now it's dominated by studios, new AI studios who are like offshoots of the bigger studios trying to do like only AI AI shows but also like the the biggest of the Hollywood studios are also there because now they have big plans integrating AI into their workflows into their existing systems. So you can see the the change in the attendance of the conference as well.
38:35 >> That's awesome. Well, for the audience, check out uh the content coming out of the the Gem Media conference. It's going to be very very exciting. And thank you so much for Van coming on to our show. It's super exciting time for Gem Media. Thank you.