← All transcripts

Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon Transcript, AI Summary & Key Points

No Priors: AI, Machine Learning, Tech, & Startups · 13 days ago · Science & Technology · 38:14 · EN

Watch on YouTube

Answer

Diffusion may win AI inference because parallel token generation maps better to GPUs than autoregressive sequential generation, improving speed, hardware utilization, and intelligence per watt or dollar. The broader architecture contest remains unresolved.

AI Summary

Diffusion models generate multiple tokens in parallel through iterative denoising, giving them a potential inference advantage over autoregressive models that generate tokens sequentially. Inception’s 2024 research matched autoregressive model quality and perplexity at the GPT-2 scale while generating text 10x faster. Stefano Ermon argues that diffusion models use GPUs more effectively, reduce inference costs, improve the economics of reasoning and reinforcement-learning post-training, and may eventually offer better controllability and data efficiency. Inception’s Mercury models serve production customers, use an OpenAI-compatible text-in/text-out interface, and compete with speed-optimized models from frontier labs. The broader architecture question remains open, and adoption is limited by immature serving infrastructure, kernels, open-source models, and deployment tooling.

Key Points

  • Stefano Ermon began working on generative models at Stanford in 2014, when models often produced grainy MNIST digit images and generative modeling was not yet a major research area.
  • Ermon initially viewed generative models as a way to learn structure from unlabelled data and develop world models, rather than anticipating current large language model capabilities.
  • 2019 — Ermon’s lab developed score-based generative models that became part of the foundation of modern diffusion models.
  • Diffusion models generate images by starting with noise and iteratively refining it, rather than generating pixels one at a time; the approach now underlies leading image and video models and some music and protein-generation systems.
  • 2024 — Ermon’s research matched an autoregressive model’s quality and perplexity at the GPT-2 scale with a diffusion transformer trained on the same data and parameter count.
  • 2024 — The diffusion model generated text 10x faster because it generated many tokens at the same time.
  • Autoregressive inference is sequential and memory-bound: each token depends on the preceding tokens, causing substantial weight movement across the memory hierarchy and relatively little arithmetic.
  • Diffusion inference processes many tokens in parallel, creating a workload closer to training and better suited to GPU computation.

AI in practice

Agents

  • Operate voice conversations for customers. 2 held 17:37

Business ideas

Build transformer-based diffusion language models that generate multiple tokens in parallel instead of generating tokens sequentially. Serve them through a production stack designed for latency-sensitive applications, while maintaining compatibility with existing model interfaces and workflows.

For
Companies building latency-sensitive AI applications, especially voice-agent providers and other applications that need high-quality model responses within a fixed latency budget.
Solves
Autoregressive language models generate one token at a time, creating a sequential, memory-bound inference workload that uses GPUs inefficiently and increases latency and inference cost.
  • Inception's 2024 research result matched an autoregressive model's quality and perplexity at the GPT-2 scale while generating text about 10 times faster.
🔒  Build steps and tools for 1 idea. Unlock

Tools & resources

4 items

CNo. 0714
AIAINotes.us AI product

Cerebras Systems

cerebras.ai

Cerebras Systems (branded Cerebras) is a company that designs AI‑optimized compute hardware and software for large‑scale model training and inference. Its products include the Wafer‑Scale Engine (WSE) and CS‑class systems, which provide high‑performance, large‑memory accelerators for deep learning workloads.

Mentioned in
3 videos
Kind
AI
MNo. 4124
AIAINotes.us AI product

Mercury

In the AINotes directory

Mercury is Inception’s family of diffusion-based language models for generating text and code. Unlike autoregressive language models that generate tokens sequentially, Mercury generates discrete tokens in parallel to improve inference scaling, hardware utilization on standard GPUs, and latency. The models target production serving and latency-sensitive applications such as voice agents, with quality comparable to speed-optimized models from frontier laboratories while generating outputs significantly faster.

Mentioned in
2 videos
Kind
AI
NNo. 0712
AIAINotes.us Tool

NVIDIA GPUs

nvidia.com

NVIDIA GPUs are a family of graphics processing units produced by NVIDIA Corporation, intended for graphics rendering, high-performance computing, and hardware acceleration of workloads such as machine learning and scientific computing. Major product lines include GeForce for consumer gaming, RTX A (formerly Quadro) for professional visualization, and data-center GPUs (e.g., A100, H100) for AI training and inference.

Mentioned in
3 videos
Kind
Other
ONo. 0368
AIAINotes.us AI product

OpenRouter

openrouter.ai

OpenRouter is a platform that provides a unified interface and API for accessing and comparing multiple AI models and their pricing. It aggregates model endpoints so developers can route prompts to different providers from a single place.

Mentioned in
17 videos
Kind
AI
🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Why Diffusion Will Win AI Inference with Inception Co-Founder and CEO Stefano Ermon — No Priors: AI, Machine Learning, Tech, & Startups (38:14). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by No Priors: AI, Machine Learning, Tech, & Startups. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:05 Hi listeners, welcome back to No Priors. Today I'm here with Stephenau Man who is a longtime Stanford professor and now co-founder and CEO of Inception. Stephenau has a extraordinarily broad body of work around generative modeling, but is especially well known as one of the fathers of diffusion. We talk about his company challenging the large labs and why speed and efficiency are going to be the name of the game in AI over the next few years.

00:31 Welcome Stepheno. Stepheno, thanks so much for being here. >> Great to be here. >> I would love for uh us to just start with a little bit of your research background and how you ended up starting your company. >> For sure. Yeah. I've been doing research in generative models for like uh basically my entire career. I started at Stanford in 2014 as an assistant professor and I was working on building generative models.

00:55 Back then the the research area was not particularly hot. Uh you know the the models were not quite working well. We were still like building little generative models over emnest and it was like a big success if you could generate these grainy images of digits and you know it was even hard to publish papers back then in on on on that topic and you had to kind of like justify training a generative model as a way to learn features from unlabelled data that then could maybe help you do better at supervised learning

01:24 because that was the thing that everybody cared about. But then you know things took over of course and so it was like a I was at the right place at the right time working on the right thing and so I've been doing research in in that space since uh since the beginning basically. >> Did you have like a a um besides a curiosity in the area a personal hope for what the models would do back in 2014 and 15?

01:48 >> Yeah. I mean I I always felt like that was going to be the that was the right way to think about uh kind of like learning from unlabelled data that like building a generative model is really the right way to make sure you understand the structure in the data. That was kind of like the way I was getting at. I I I was not even dreaming about the kind of capabilities that this LLMs that we that we have today uh could could uh could do that.

02:09 Uh but I was thinking more from a I think world models perspective like I was I was working a lot on images and so thinking about okay like I have a world model I can imagine what's going to happen if I were to stand up and walk out the door like I can kind of like picture that in my mind and that's important uh to make decisions and kind of like model predictive control when having this kind of model of the world requires some generative capabilities and so I always felt like okay that's the right direction to work on

02:37 I felt like this is going to be very hard as a problem is going to keep me busy for my whole career and it's a good problem to work on and then of course I was very wrong and things [laughter] evolved much faster than than I was expecting. >> Yeah, I think that's kind of universally true though. Um and sort of walk me through the the state of your research and how that led you to start the company.

02:59 Yeah. So I was working on generative models of images initially working on autogressive models which were very slow and kind of like very blurry and then vaes and then GANs took over. >> Yes. >> And uh back then we were very unhappy with the state of yeah generative models for images like the GANs were they they they worked but they were very unstable to train very hard to reproduce results.

03:23 And so we were trying to see is there a way to build something that is as good as I can but it's more principled. And so we started working on scorebased generative models which are basically what eventually became uh diffusion models back in 2019 with with my PhD student. And so we kind of like came up with this idea of let's train a neural network to den noiseise images.

03:43 And if you can den noiseise an image then you you really are understanding enough about the structure of the image that it should be possible to build like a generative procedure based on these deninoisers. And that basically became the the kind of like underlying technology of modern diffusion models where instead of generating images you know left to right one pixel at a time you kind of like start from pure noise and then you gradually refine the object until you get like a clean picture at the end.

04:12 And that started yeah back in 2019 in my lab and then it kind of like took over the space and and even today the best models for image generation, video generation, music to some extent a lot of the protein stuff they are based on diffusion and so my group has worked a lot on various kinds of diffusion models technique for accelerating them to to generate samples very quickly to improve the quality of these models.

04:35 And so uh since we were able to get them to work on images, I started thinking about how do we get diffusion models to work on text or or code generation and can discrete objects like is there a way to move beyond autoagressive models to uh something that it's more parallel more [snorts] with built-in error correction and so I've been doing a bunch of research at Stanford on getting distri models to work on text and code generation.

05:02 Um we had a breakthrough in 2024. We published a paper um basically showing that for the first time it was it was possible to match the quality of an autogressive model at the GP2 scale. So less than a billion parameters still fairly academic but we were able to train basically still a transformer model as a diffusion model on the same data. We were able to match the quality like the same perplexity and you were fitting the data just as well as an auto reagressive model with the same number of parameters.

05:32 But the diffusion model was significantly faster because it's diffusion because you're outputting many tokens at the same time. We were able to generate text like 10x faster compared to the autogressive model. And so that felt very very exciting. And I really wanted to see what happens if you scale up if you train bigger models. And so I started inception uh a company to basically scale up the technology and and try to build um commercial scale diffusion based language models.

06:00 >> Everyone has now seen um the outputs of diffusion models in in in particular images. Yeah. And uh I I would argue that it's like increasingly a dominant form of like generated short form video from diffusion models is like a dominant form of entertainment in other parts of the world and it will likely become so here. Um it's kind of unbelievable at least to me even having followed the field for you know the last decade plus uh the the quality that is possible today.

06:28 Um so I think that is kind of obvious right and it's such a big use case in images and and video generation that um folks are even creating you know hardware to support better uh um better performance here. Um it's not intuitive that would work for other fields or that you know this is a interesting competitive direction to uh the um current you know full transformer focused like AGI labs um can you offer some intuition on that?

07:04 >> Yeah so it's it's a very interesting kind of like state of the world right now from a researcher perspective because like there is like two main paradigms two ways of building generative models. There's auto reagive where you you kind of like have a model that predicts the next tok the next token or the next pixel and then you generate left to right one token at a time.

07:20 And then there's diffusion which is a course to find generation um kind of like iterative den noising kind of generation [snorts] and as you said like we have continuous modalities where diffusion dominate. There is discrete modalities text and code where primarily all the big labs are kind of like betting on the same architecture autogressive models and as we move towards more and more like multimodal models and kind of like there is this idea that maybe we'll have a model that can handle all modalities and we'll know

07:51 everything about the world what architecture will that be like will it be an autogressive model will it be a diffusion model uh nobody knows I think that I think the jury is still out there uh at inception betting on diffusion models because we believe that uh what matters eventually will be inference time scaling and there are fundamental reasons for why diffusion models are better than auto reagentive models at inference time.

08:16 So even if you think about the story of autogressive models, there was an inflection point in 2017 when people switched from RNNs to transformers, right? And and why why was that? The the the problem was that RNNs had to essentially process tokens sequentially, one at a time and training was very slow and so people came up with this idea of let's have an architecture that allows you to process many tokens at the same time in parallel.

08:41 And that was a transformer and that that was the thing that scaled better for training and that enabled a lot of the successes behind LLMs. >> But if you think about inference now not training inference generation uh auto reggressive models are still sequential uh the computation is one left to right one token at a time. You cannot generate the 10th token until you've generated everything that comes before it.

09:06 That kind of workload is um does not map well to GPUs. that kind of workload is extremely memory bound. Uh you're spending most of your time moving around weights across the memory hierarchy and you're doing very little arithmetic and that's a fundamental problem of auto reggressive models. And so what's the equivalent if you think RNN's transformers auto reggressive models the equivalent at inference time is a diffusion based is a diffusion model because a diffusion model is built to have at inference time a workload

09:38 where you process many tokens at the same time. And so the the the the workload that we have at inference time in a diffusion model, it's basically very very similar to the workload you have for training where you're processing many tokens at the same time in parallel. >> And so it's built to essentially have an inference workload that maps really really well to map moles, but the kind the kind of things GPUs do really really well.

10:05 And so we bet on trying to build the architecture and trying to build the kind of models that will scale best at inference time because um you know economics are dominated by you know the kind of intelligence per watt the intelligence per dollar that you're able to get from the models. If you think about a lot of the advances with reasoning models, a lot of it is scaling test and compute, right?

10:28 And so being able to scale better along that axis will also matter. And even if you think about RL post training, a lot of the bottleneck is generating rollouts >> like letting the model explore, you know, and and then scoring the trajectories and then improving the model based on the kind of things it finds. And so inference is again a key bottleneck for our outpost training.

10:50 And so if you have a model that scales better at inference time, then automatically you're going to get better scaling during our outpost training. And so that's why we decided to bet on a diffusionbased LLM because it's inherently more parallel and the bitter lesson is that the more parallel solution is the one that is eventually going to win. How did you think about um applicability or what experiments did you run in terms of cracking the nut on discrete versus continuous modalities?

11:20 Because I think people have also uh shaped the existing you know dominant paradigm through new tokenization uh efforts or or uh methods to make video and voice work for example. Um uh it's you know this is you're you're not in the same tokenoriented um paradigm. How do you make it work here? >> Yeah. So, there was a lot of research that uh that went into figuring out how to apply a technology that was inherently very tied to kind of like continuous structure in the data.

11:53 >> So, if you think about a diffusion model, it's learning how to den noiseise images and it kind of like makes sense for continuous data because if you think about even two pixel colors, you can kind of like interpolate between them and it will still make sense. But if you think about two words, there is not necessarily something in between them, right?

12:08 It's all discreet and so it required a lot of R&D and a new science that had to be developed to figure out how to how to extend those kind of ideas to discrete spaces. >> What can you claim about how well it works today? >> We think it works really well. So we've been able to to train uh diffusion based LLMs that are comparable in quality with the the speed optimized models for Frontier Labs.

12:32 So our Mercury models are on par with the haiku models, flash models, mini nano models from OpenAI if you look at benchmarks uh while being significantly faster. So we've uh crossed I think the the the we went from you know pure research prototypes to things that are actually used like we are serving these models in production today. We did all the work of figuring out how to even just like build a serving engine, right?

13:00 you cannot run this diffusion based LLMs on VLM or SG lang like you have we had to build our own serving engine and we can handle a lot of the complexity of like real production workloads u and uh we we've we've solved all these challenges and we can deliver this kind of like new experience end to end to to real customers today >> actually a great time to just talk about where inception is as a company like how many people what are you guys actually serving um sort of state of research yeah >> yeah so it's We are about

13:30 two years old uh around 50 people uh spending a lot of time still on R&D kind of like figuring out what's the right way to train these models um how to accelerate inference like it's not obvious how you even if you think about an auto aggressive model it's pretty clear there's not a lot of things you can do there in terms of like okay you generate one token at a time and that's it in a diffusion based model we know that there is a lot of different possibilities for trading compute for quality at inference time like

13:58 even And if you think about image diffusion models or video diffusion models, there's a lot of techniques that you can use to kind of like accelerate sampling of distillation or like fancy uh differential equation solving techniques that allow you to sample very very quickly from these models. And so there is a lot of research on the training on the inference and then engineering like just like thinking about data mixes evolves um RL post training infrastructure like there is a lot of work that uh that needs to happen

14:28 to figure out how to how to build recipes that work for this new model and we try to leverage existing things as much as possible. For example, it's still a transformer-based model. So you don't have to throw away uh a lot of the work that has been done on on good architectures. uh we still use attention uh we still use uh a lot of the public data sets that people have created and evils and benchmarks.

14:51 So you know we're a start we're a startup we try to be scrappy we try to use existing things as much as possible and and kind of like focus on the on the the aspects where we can have the highest impact and then where we can be the most differentiated and right now it's speed in the future who knows like it's possible that a diffusion based language model will be maybe significantly more intelligent than than an autogressive one like nobody knows that that's why I think this is very exciting because we're developing

15:18 these really powerful AI systems but it's all very fresh. It's all very new. Uh I doubt we've discovered the best way of building these systems. There's got to be alternatives. There's got to be other ways of building these models and and eventually yeah efficiency will be very important. Like if you think about the AI factory like how is that going to work?

15:40 I think nobody really knows and just being able to play in that space and thinking about alternative ways of creating intelligence. I think it's exciting. >> Absolutely. And I I also think that in a uh increasingly like fundamentally compute and supply constrained environment uh the you know for for some I I wouldn't say that there wasn't very focused research on um efficiency of models but it was a second order uh uh sort of consideration for many of the industrial research efforts versus like pure capability scaling

16:16 on what we've thought, right? Um and and new methods in that. But, uh I I just think if you if you fundamentally believe that we can use all of the power we have and there will be economic reasons to do so, then I think the focus on efficiency is going to go dramatically up, right? Um you know, I I I think many people now argue like one of the most important decisions for um a company in the AI space is like how do we use the power we have between, you know, training and the most valuable use cases for any any watt.

16:44 on that topic like here in you know uh September of 26 uh where does speed win like where do people care enough about this already? >> Yeah, it is basically applications where latency matters. I mean generally speaking I think everyone cares about speed in the sense that if you can give me the same uh quality but faster people will always pick the the faster solution.

17:09 And then we're seeing it with this like faster versions of even the the models from Frontier Labs. People are willing to pay more to get access to to faster models, right? And and I think once you get used to a fast model, it's hard to go back. It's kind of like broadband, right? And then get faster and faster and if you if you were able try, you know, people cannot go back once you once you get the fast model.

17:32 >> Uh are there customers that you can talk about publicly that um you know care about this? Today >> there are a few that that we can mention uh like in the voice space for example um open call is an example um they they're building like voice agents um they you know speed of course the pipeline is like you have an ASR model you have an LLM that it's kind of like doing all the tool calls and figuring out what to say next uh it has to be a reasoning LLM typically to to have the highest level of quality and then there

18:00 is a texttospech uh component at the end um speed matters a lot to um they were previously using uh serving their L&Ms on on on Cerebras. So they were using custom chips to get to the kind of speed that they need to to deliver the best experience to their customers. And then they switched over to to our diffusion based LLMs because they can essentially get the same speed as what you would get if you were to run an auto reggressive model on custom hardware.

18:29 If you have a diffusion based LLM that it's built to be parallel, it's accelerating at the software level, then you can get the same speed on Nvidia GPUs, uh, which means much more, um, availability. I mean, GPUs are scarce, but there's more of that than than custom chips, uh, and lower cost, higher quality. So, that's an example in the in the voice space.

18:50 Yeah, I was actually going to ask you how you think this um interacts with the uh hardware landscape as well given we've seen enough now demand from use cases who are like yes I want a big expensive chip with a lot of SRAMM and people will pay for the outputs of that. Yeah, in in coding and other use cases. >> Yeah, for sure. For sure. And I think hardware is one way to to to to accelerate things and the good thing >> and software might be better if we can use the existing hardware.

19:16 >> Exactly. Exactly. And especially they are complimentary. That's the exciting piece is that to some extent uh the gains that you get from the software they are multiplicative with the gains you get from the hardware >> and maybe someday people will develop uh system uh you know hardware that fits even better the models that you're building. >> For sure.

19:33 For sure. Yeah. Yeah. If we just like zoom out to um the you know inception in the broader industry uh I I think there is a vein of concern and correct me if I'm wrong uh that uh it's very hard to invest in new architectures today because some you know if there are advances in architecture or methods it will simply be absorbed by uh large players with the resources to scale compute.

20:01 um talk to me about how you think about going and competing as David in this situation. >> Yeah, it's a very valid point and something that is also like top of mind for us. Um I think initially for us for sure like the mode is sort of like the IP the trade secrets like the the ideas that we and our researchers have to to to build these models and make them better.

20:22 um as we mature as a company and and one of the reasons we are not just doing pure research but we're also like you know we've developed a product and we have real customers and we are getting feedback on on the models from from the real world is that by doing that we are also like developing components that are also very important to to deploy these models for example like a serving engine uh you know if you don't have the serving engine you can't really serve these models in production and So by forcing ourselves

20:52 from the very beginning to go out and and deploy something end to end, we're learning a lot about how to serve these models and how to build uh software that it's kind of like needed to to to run these models and and that again becomes IP like even if you train the diffusion based LLM, if you don't have the serving engine, if you don't have the VLM equivalent to serve it, you're still stuck and you still cannot use it.

21:18 on the same line like along the same lines. We are working with real customers and we're getting feedback on the models. We figure out what works, what doesn't. We collect data sometimes from them. We create evolves based on what they're seeing. And so that again becomes part of the the technical mode because uh you know, of course, those things are a little bit harder to to replicate.

21:40 You can tell me if this doesn't make sense as a a question to ask um technically, but one of the things that diffusion models benefit from structurally in images or video generation is you know you're replicating something where there should be some consistent structure in the world. Uh voice as well, right? It is whatever is really possible and most likely.

22:04 Um there are like some of the fields where uh AI has been most valuable to date. Uh I I'd say like you know a lot of the input data you use to train like code data for example. Um uh it's it's very messy right and you know one could argue that a lot of it doesn't actually have the like correct real structure you're looking for. How do you think about that when it's like human generated input data versus you know images, video, voice?

22:32 Yeah, it's a good question and fundamentally if you think about whenever you train a generative model what you're doing is whether it's an autogressive model or to some extent even a diffusion model is you are trying to uh identify structure in the data to by essentially building a compression scheme. That might not be obvious but whenever you train these models you're effectively trying to identify common structure by trying to find an efficient way of compressing the data.

22:58 And so the more you can compress the data, the more structure, the more patterns you're identifying. And that's how these models work, which is the the the amazing thing is just like by predicting the next word, uh you are learning something about the structure of the data. And that's the same whether you're using a diffusion model or you're using an autogressive model.

23:17 Both methods are essentially trying to learn a compression scheme. And when I mentioned the the original 2024 paper when we showed that we are achieving parity with auto reggressive models the metric that we're using is basically perplexity which is a notion of how much structure have you identified in the data >> and so even though it might not seem obvious we were actually able to identify at the GP2 scale >> the same amount of structure as an autogressive model.

23:50 >> Yes. I I think like that that empirical result is there but the intuition would be like well this you know code is not uh the the data set that you are working on is not like grounded in physics right there's a lot of noise in there and it sounds like that is you believe that's a manageable problem. Yeah, there is noise in in everything. And so to the extent, you know, the numbers don't lie, to the extent that you you're able to drive the perplexity down, then it means that you can actually build like a a

24:15 compression scheme uh and that that will get you that sort of like level of compression. And so the the structure must be there and the model must have been able to uncover it. And then it's more a question of an inductive bias like you know is is a transformer a better way of identifying those patterns or something else. uh is next token prediction the right modeling framework or is it more like den noising and that's that's very much an empirical question that I think at the moment we don't have a tools even to

24:46 understand >> can I ask a question just because you used um a voice customer as the example here uh one of the uh benefits that some people building these AI products have identified of having an LLM in the middle of this voice pipeline is they understand how to um do alignment a little bit better there or controllability. I imagine that has to look different for a diffusion based model.

25:11 Can you talk about that? >> Yes, that's the that's a key value proposition and one of the things that they they always look into is you know yeah to what extent uh a lot of the value they provide is like the harness and making sure that the models indeed are doing the right thing. And the interesting thing about a diffusionbased LLM is that we've built everything to be backwards compatible.

25:28 So it's still like the API is the same. It's still OpenAI compatible, text in, text out. And it so happens that the models we've trained are good at following instructions. They're good at outputting, you know, if you're using JSONs, like structure outputs. They can handle all of those things. >> And uh it was good enough. It was better in fact than the models they were using before.

25:50 And so they are still able to provide the the kind of like level of service to their customers by using Mercury. >> Well, very simple if the interfaces are the same and you can just use your same stack. Yeah. But it could be that that I think that's actually a very interesting point is that we know that diffusion models are typically easier to control compared to autogressive models.

26:09 And the reason is that if you think about an autogressive model, you kind of like have to wait until you've generated the full object to know whether or not it satisfies let's say a constraint or whether or not it's aligned or whether or not what whatever you know at some brand whatever it is that the the objective function that you care about. Maybe you're generating a molecule and you care about solubility and then and then you kind of like have to wait until you have the full object to be able to score it with some

26:34 reward function. But a diffusion model it's more course to find generation. So from the very >> you can progressively [clears throat] do it. Yeah. >> From the very beginning you know kind of like is this object the kind of thing I want or not and you can steer the generation in the direction provided by an external reward function or some set of constraints.

26:52 And so there is a lot of evidence in the academic literature at least that diffusion models are easier to control and there are different ways of steering them that are just not possible without aggressive models. So that would be a different interface >> for the model that maybe might not be even available for autogressive models. I think that's something that uh we've been thinking a lot about like what would be the right what's the right product experience that we can build around new capabilities that are just not

27:24 provided by autogressive models. >> Are there capabilities that you imagine um in inceptions models uh having at scale that today's models don't have beyond performance? >> Yeah, that's the thing we don't uh we don't know right and that's the it's emergent. Yes. Yes. Like right now the wedge speed uh we know they are much faster. That was the initial bet because that was easier to test.

27:46 It's also like easy to you know to to measure and uh it it's obviously valuable right but as we that's why I find it so exciting is that as we learn more about how to train these models you know we don't know what we're gonna find. And uh there is a decent amount of evidence in the academic literature for example that diffusion based models are more data efficient compared to autogressive models.

28:14 And the intuition is just like if you think about training a diffusion model um you're learning by denoising. You start with an image you add noise and then you learn how to remove the noise. And so it's effectively doing data augmentation in the sense that the same image is augmented by many noisy views. >> Okay. >> Yeah. Uh and so they tend to be a little bit more data efficient and so if that's you know holds up at scale and then you believe then maybe we'll get into >> tasks where we have less data.

28:47 >> Yeah. Where data becomes the bottleneck then it becomes more interesting. Right. >> And so we'll we'll see. That's why it's so exciting because things are changing and and and then you know this technology is so important and so valuable that having something differentiated I think will create value >> if we uh project out you know five years um uh that's actually way too long in AI world if we project out two years like uh what do you think is the workload split between diffusion and uh traditional um models?

29:18 I think we're still uh we're not at the frontier level of of intelligence and I think a lot of the workloads do require frontier level intelligence but in my estimates like even if you just go to you know open router has this very nice way of kind of like looking at all the different use cases and you can kind of see you know the research and conversational and coding and software engineering and log processing log like they have like a nice hard taxonomy basically of tasks >> and I was doing some estimates and I think

29:48 there is like between 20 and 30% where latency is really really important and so at the very least as a lower bound I think it could be >> address >> addressable by models that are >> within a given latency budget they will give you the highest possible quality >> and then just you know all technology approaches have trade-offs what are what are the uh challenges of working with diffusion models >> yeah it's it's a different stack and so one of the challenges that we had to build a lot of things in house and uh there

30:20 is not a mature sort of like ecosystem of uh uh if you think about the serving engine or like kernels like uh a lot of the things had to be developed in house >> and so there is not really anything open source or there are some open source models but they're not particularly good and so that makes it a little bit more difficult to you know deploy to get customers to try things they they're not used to it so that's that's been uh one of the challenges >> I imagine that also reflects externally right you know in a

30:52 landscape where folks uh at some sophistication where they would care about cost and performance and though you might uh for for certain use cases you will care about cost and performance from for the beginning uh there's an increasing amount of interest in post training right and so I imagine in a uh new architecture that's even more challenging >> so we had to build our own stack for doing SFD for doing RLHF doing RL uh I mean that becomes IP to some extent.

31:22 So it it's it's it's one of the reasons we decided not to open source everything was really to to keep the IP a little bit closer to us and not opening it. But then there are downsides like there is less less an opportunity for the community to contribute. It's harder to to adopt. It's hard to do on prem kind of like deployments and so there are pros and cons with the two choices.

31:44 Can you talk about you know scale of your own training uh and then um like current or aspirational uh and then uh at 50 people I'm sure you're continuing to hire like why researchers or engineers or others should consider um you know investing in this direction or working at >> yeah so we're not able to share much about the the the the training uh the size of the models or the flops or all of that it's kind of like a trade secret but we are continuing to push the frontier and inception is a great place to be if you

32:17 want to be have an opportunity to shape the the direction of the field. Like it's still a relatively small field. There is a lot to be invented and so a lot of the people that decide to to come to inception instead of joining one of the other labs is really that they want to have ownership and uh they they like to invent new things. They like to they they like to be in a space where there's more of a green field and more opportunities to um try things.

32:41 there's less that it's known or available out there. It's a little bit more open-ended and so we we tend to attract those kind of people. >> I think one thing that is both exciting and causes some despair amongst research friends is um uh you know the ability to use models for recursive self-improvement in the research field itself. What given you're working on like a very different direction, what is your view on this?

33:08 Yeah. I mean, it's uh it's something that we >> I mean, explicitly it sounds like you still feel there's work for you and your team to do. >> Oh, yeah. Yeah. I think we're not there yet. Uh maybe we don't have access to the models that other folks have, but I feel like there is still uh uh we use models a lot of course and it has accelerated the the the speed at which we can iterate, try ideas and um you know, we use models from Frontier Labs and and uh uh it's been it's been great.

33:38 it has accelerated our development process a lot. Um at the same time I think uh at least that right now I don't know what what it's going to be in six months or a year but right now uh the the the human ingenuity is still like super important and the ability to come up with the the right ideas and kind of like prune the space and and kind of like identify directions that are more promising has been really important to us.

34:04 50 people is not that many people for let's say like a full stack you know uh research serving product company um or however you would think about describing it uh how you know how do you organize and then how do you think about how you allocate your resources here >> it's a small team but everyone is very talented and and uh they they work very hard and and we have access to agents that are making us a lot more productive and so I think the numbers are uh you know are sufficient to do to do a lot and and in fact often

34:38 the I feel like the the the bottleneck is more compute than than than people. Uh but yeah, the team is organized like there is a there is a a product team effectively that is handling the platform and working with customers to to to make them successful with our models and then so there is basically a team that is serving the the current best version of the model and then there is a team that is building the next version of the model and that includes training RL inference and so that's more research.

35:11 Stephano, one last question for you. Uh, you know, the 24 paper was a super interesting result, made a big splash. You've been working in this field for a long time. A lot of folks would say that, uh, would claim that, you know, academic AI research is very challenged in this era of, you know, being able to scale resource a great deal. Like this is certainly true to some degree given you you started a commercial company around it as well.

35:36 But how did you get confidence in the directions that you were working in having impact or being promising before you really had th those 24 results? >> Yeah. And and I think it was like a um a collection of results that uh that I I had been that I've been working on in in my lab. Not necessarily like of course there's like the early diffusion work that we did in the lab.

35:58 We showed that the at the kind of scale of models that we could train on uh academically. we were able to kind of like beat Gans and so and then while being much more stable and then the whole thing took over and then became stable diffusion mid journey sort of all of that started from ideas that were developed in academia in my lab but that's not the only one like I was co-advisor so I worked on flash attention for example right that's another thing that came out from academia that then eventually had a huge impact in

36:29 industry right or another example is DPO that was another project started out as a rotation a project in my group. Uh it's an algorithm that is used to align, you know, LLMs and diffusion models and everywhere, right? And and that's again something that that was developed entirely in academia and it was just like based on a on a clever insight like some some interesting mathematical structure that you have in that problem that allows you to come up with a very different and more efficient way of post-raining and

36:59 aligning these models, right? So uh there are gems there are uh lots of opportunities for finding uh new and better ways of solving important problems. Uh one of the nice things about academia is that it allows you to take these contrarian bets. Um as you said I mean there is the challenge that maybe we don't have enough resources and there's never enough resources and if we had more compute we could be more efficient.

37:25 uh but the you know you have access to amazing students and everyone is kind of like trying to develop the new thing people are not scared about taking bets and uh that's that's why academia has been so impactful I think or even in the AI space a lot of the important ideas have roots or even were created in academia. >> Awesome. Super inspirational.

37:49 Thanks so much for being here Stephano. >> Thank you so much for having me. Find us on Twitter at no [music] prior pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every [music] week. And sign up for emails or find transcripts for every episode at no-briers.com.