Searchable transcript of The Best AI Companies Have Unique Data Acquisition Strategies | Simile Co-founder & CEO — 20VC with Harry Stebbings (01:05:03). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by 20VC with Harry Stebbings. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 I think there's a world in which in about 2 to three years we're running a single simulation session that people will pay $100 million for it. This is Jun Park, founder and CEO at Similey. They predict the future. They're a simulation market that try and predict future human behavior. It is incredible. >> My fundamental thesis here is for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible.
00:24 This was one of the best AI technical conversations we've had in a long time and it was incredible to have June on the show. People live through different stages in their life and they have different careers, different jobs. At each stage of their life, were they the reason why that thing was successful? If you squint, were they the common denominator?
00:42 Ready to go. June, I'm so excited for this dude. When Shardell told me uh that I had to meet you, I'm going to be honest. Shardell does not tell me often that I have to meet someone. So, I was like, "Wow, this is I feel honored. Thank you, Shardell." Um, and then we met when I was on holiday with my family. And I remember my my grandparents were like asleep upstairs and so I was whispering to you and I remember being like so excited by what you were building but then also having to be incredibly respectful of the
01:21 sleeping elderly people next door. But thank you so much for joining me dude. >> Thank you for having me. Excited to be here. Now, when I spoke to a lot of your investors and friends before, they all said that I had to start on the very unique background you have of specifically kind of became very well known for a particular project and it centers around Valentine's Day and a simulation that happened as a result.
01:43 C can you explain what happened and how that potentially led to the early days of simile? For sure. So, this was 2023. uh we were we had this idea that large linkage models are often used for simpler tasks like classification, simple generation but we thought that these models actually had a lot more potential. One of the early observations that we made was that these models are trained on so much of human behavioral data, sentiment data that were expressed on the web.
02:11 So if you poke at them sort of at the right angle, you could actually extract a lot of realistic human behaviors out of them. I thought that was really interesting and it was also particularly interesting in that it was domain agnostic. So if you look at the literature in computer science for many decades we've always had the vision of creating agents that are meant to be generalizable that are meant to really be able to act like human in any environment.
02:37 And my mind went to well maybe we have that opportunity here. So what we ended up doing was well if we were to fast forward many years into doing this what would be the most ambitious vision that we might have and that was creating entire lived experience of a town. So the idea here was we would make a game town and we would populate it with 25 NPCs.
02:56 So non-playable characters except these characters would actually wake up in the morning, do their routines, go to work, have relationships and do all that. They would actually remember their interactions. They would actually plan their days. And some of the surprising things you end up seeing was the simulation itself was set the day before Valentine's Day and you actually see these agents come together, have parties like self-organized.
03:21 So they would actually plan parties, they would decorate the cafe and so forth. We thought that was really interesting. Now two fundamental contribution from that work. One was it was one of the earliest example of creating agents. So this particular uh set of agents were paired with back in back in the day GPT3.5 text. So we didn't quite have chat GPT back then.
03:43 Uh and then was paired with memory planning and reflection really the first times that those concepts came out to be uh an explicit part of the architecture in quote unquote agentic workflows. The reason why we actually got that inspiration was if you had more than one agent side by side you want them to remember each other. Back in the day, lynching mortals didn't really have the concept of memory.
04:05 So I thought, okay, you have to give them me the memory so that they don't say, "Hey, nice meeting you every time they meet their roommate." So we had them give have this concept of memory and planning and reflection to make sense of very long-term landscape. >> How do you solve that memory problem? Cuz everyone says, "Oh, we have a memory problem today."
04:25 How do you solve the memory problem of agents to prevent that from happening? So back in the day was actually the initial idea was fairly simple which was that these language models are actually quite good at processing natural language. So we put everything in markdown text file that was it that sort of worked. Now the issue there however is because the language models have context window and because and even today even if the context window is getting larger the kind of experiences that these agents can have in the
04:55 small game town is immense and imagine now if we were to bring this to real life in world like the one we live in the amount of memory that we accumulate is huge. So the the problem becomes how do you make sense of this large quantity of memory. So imagine you went to get omelette five times in a day. You want to make sense of that aside from oh I went to get omelette five times you know five times throughout the week or something like that.
05:19 So we had this concept of reflection which basically was every certain interval. It's like a shower thought you have you ask agent explicitly to get bunch of their memory pieces and basically make sense of them. Why did you get omelette so often this week? Were you busy? Do you like omelette? Why are you studying for this test so hard? Like you are in library every single day.
05:45 Like does this matter to you? And they will actually start formulating ideas that are more higher level than what happens on the ground truth. So gradually they start to realize oh this particular research topic I'm actually quite invested in it. This might have actually have something to do with my childhood or my fundamental memory. This actually shapes who they are as a person.
06:04 So that ends up becoming a very useful function in creating these agents that have personality that actually has a point of view in the world that can actually make sense of a lot of this data. So that's how we did it back in the day. And so when we think about a simulation model today, for those that don't know, a simulation model is essentially that it's the creation of agents that then produce a set of activities or actions that then will show us what a simulated future world might look like.
06:32 Is that correct? That's right. >> Got you. Okay. When we think about then building a simulation model company, would you say Sim is a simulation model company? Yeah. We are a company that is creating foundation model of human behavior that can then be used to create simulations of individuals, simulation of subopuls and down the line the simulation of the entire ecosystem and even the market.
06:56 Do you sit on top of core foundation models? How do you think about the relationship for those listening between an open AI anthropic frontier model provider and you? Yeah, so this is a great question. So the way we see it is if you look at large link model companies today, fundamentally the task they have at hand is to create super rational intelligent machines that are good at coding, that are good at natural sciences and mathematics.
07:21 Simile doesn't really care about any of those. What we care about is if we have a person make a mistake in this context, we want our models to make the same kind of mistake. We want our models to be biased in the same way humans are. In a way, we want to be a representation of people's values, preferences, and taste, sort of their subjective half of their brain.
07:44 That's what we care about. I love that. A lot of what people say is different to a lot of what people do. How do you think about the chasm of what people say and what people do and how that impacts your models? For sure. So say do is real and you know if you look at the web data it is fundamentally data of what people have said not what they have done and obviously large language models today are trained prelim uh preliminary um mainly on this web data.
08:16 For us we actually do collect a lot of behavior data. We collect uh transaction data. We collect observational data. We also partner with our um customers uh our vendors to collect some of this data. But my personal hot take here is a lot of observational behavior data. What they're amazing at is actually helping you create a correlation of the observation and what could happen in the future.
08:42 Good for prediction task. But my take here after interacting with so many of our customers and also being in research, no one really cares about prediction. no one really cares about what's going to happen in the future unless you're trying to predict the stock market. What people actually care about is they want to shape the future. They want to know imagine you're a Starbucks doesn't really help them to know that your fraction of sales is going to tank in two quarters.
09:09 They'll hear that and they'll be like what do we do about them? That's terrible. What they want to know is how can we prevent it? What do we need to do now to change the future? And there what you really need is causal mechanism. You need a model that can actually reason about causal mechanisms and counterfactuals. So the kind of data that we care deeply about is a lot of randomized control trials.
09:32 We actually run a lot of AB testing. We show the models. Imagine people have done this versus that. This is how their behaviors would actually change. That becomes a core part of our training asset. So this is actually the data collection that goes beyond observational data that similarly collects. Is data collection acquisition. The hardest element of building simulation models for you like if we think about the kind of core pillars for traditional models it might be compute algorithms and data is is data the biggest
09:59 challenge for you data is an important piece of simil for sure u my fundamental thesis here is for AI companies of this generation you need to have an interesting data strategy that's going to be defensible and for us really the data collection challenge comes from two angles one is actually sourcing people sourcing people here is a little bit different than what other language model companies might consider to be their people or their population.
10:22 We don't go after these expert programmers or expert scientists. We go after people like us like everyday people living their everyday life. Um that's but what we care about is are they representative? Do we actually have the same representation of people as we do in the world that we live in? And then actually asking the right questions to these people.
10:46 What are the experiments? What are the questions that actually get at the fundamental core nature of who they are? Some of the questions we actually ask at the start of our data collection at times is actually saying something like tell us the story of your life. Where did you grow up? What did you experience? What were some of the hardest problems that you had to tackle or decisions you had to make?
11:05 Tell us a lot about these people. And that's what we try to do >> in terms of people don't want to predict the future. Just so we can drill down on that. I thought they do like Starbucks if they can predict that Frappuccino sales will be down in two quarters, they can amend bluntly their buying cycle, they can change how much they purchase. Isn't that valuable?
11:26 And what am I missing? But that's the thing. The reason why they want to know is so they can change their strategy. So certainly talking about oh how much u resources they actually need to actually serve this market that is a kind of changing in behavior but fundamentally it is about counterfactuals. So well we have this market that we want to serve we want to maximize our value as a company.
11:51 What do we need to do to make sure that we react to this dip in the market whatever it may be. Now fundamentally though the work that we do is about people we do uh we try to simulate people and represent people's perspectives. So the value that we provide is counterfactual in terms of what your consumers what your population would do. When you look at what can be done for some of the biggest brands you mentioned like a CVS there incredibly valuable for surveys for customer feedback for determining what customers
12:21 really want moving forwards. I don't know how to say this. I'm being rude. How do you do you want to just be a next generation qualrix and how do you prevent that being the end goal? Yeah. So the way we see it is again fundamentally the core primitive of what we're trying to build is very straightforward. You tell us what population you're interested in and we're going to model them.
12:43 And really so far the layer of innovation has lived in the more tooling layer. How can we create better survey tool? How can we create a better interview tool? Simulation is fundamentally about something different which is how can you create the most generalizable model of people so that we can represent people's viewpoints at scale that goes beyond simply running surveys or interviews down the line I actually see simulation as a field moving into context where hey can we actually create simulations of many people
13:18 interacting with each other so that you can understand all the downstream implications of of your decision-m making or if imagine you have a new product you're about to launch can you actually simulate the entire launch and how the audience might actually react how the market might shift and this also goes into the scientist part of me also get quite excited by the vision where simulation I do think can also be a cure for many of what we call quote unquote wicked problems a good example uh here might be things like
13:49 climate change requires collective action across many stakeholders who have different incentives. One of the reasons why such problems are so difficult is actually finding the right equilibria state where all all different parties come together to make decision for global good is very difficult. Can we actually simulate those decision-m processes? Can we actually simulate even in things like in what conditions does a democracy fail?
14:15 Can we actually predict that? These are the kind of questions that simulation ultimately can answer. Can I ask you in thinking about like democracies failing elections for a government? This would be an incredibly useful tool. How do you think about who you can and should work with versus who you shouldn't? Yeah, this is for us where the principles matters so much.
14:41 The way I see it, simulation as a piece of technology is one of the twin pillars of technology. I'm a fan of science fiction. You read any advanced science fictions, there's always two pillars. One is some form of AGI that always shows up. The other is simulation. And like with any powerful technology, the misuse, the potential for misuse is quite real.
15:03 And the way we see it, simulation at its best ought to be representation at scale. People have different viewpoints, different perspectives, different taste. Many of their viewpoints are not considered in rooms where important decisions for them are made. We want to always say we listen to our people, we listen to our customers, we listen to our stakeholders.
15:25 In practice very difficult. This is a way for us to ensure that in every decision making we actually listen to people at scale. That's the north star. How much data do you need to feel confident in an accurate prediction outcome to be displayed? Like is it a 100 people? Is it a thousand people? Is it a million people? You want to have more people represented so that you can segment down to specific subop.
15:52 If you look at any social scientific literature if you have a very narrow population of interest, you would usually get statistical significance in the study that you want to run by the time you have thousand people. However, often times the kind of ways that people query our system is they want to come in and say, "Hey, filter down to X XYZ population."
16:13 Those filters are often created on the fly for us to then be able to simulate people's responses across all those filters. Thus being we want to represent the entire population. So that's the journey that we're on. >> Is it self-fulfilling? Do you get better and better at predicting over time? >> You know, I think that certainly is the case because right there is the data flywheel.
16:38 There is the learning that occurs as we get more and more simulated results and see what happens in the ground truth. That absolutely yes. And this is obviously one of the core value proposition for our early partners because they know in their business context is similarly getting better and better and better and do they have that compounding advantage?
16:57 It's kind of like Alph Go, you know, they just beat the out of the model and played it a thousand times. Every day of activities and outcomes in the world is another game of Alph Go where you can correct the model on what was wrong and what you missed and what didn't happen. And 10,000 days in, you should almost be better than the model at the model.
17:14 Do you know what I mean? So, this is actually quite interesting. Um, you might think, let's actually think about a different example. So how does data flywheel work in simulation and why would it work? If I take a brief detour and talk about coding, the reason why coding agents has had such massive improvement over the years was because their learning their reward function was extremely clear.
17:40 If you make a suggestion and your user says accept, fantastic. If they say reject, also very useful. You very quickly know what is good and what is bad. That actually was one of the core learning mechanism for these models. And it might be easy to look at simulation as a field and say well where are you going to get the reward? Because fundamentally all the things that you're trying to predict is happening in the future.
18:03 It's going to be hard to validate. It is true. At the same time I actually think simulation has even better mechanism which is the world is our ground truth. We live in the ground truth world. So what we can do is every single day we can be generating tens of thousands of hypothesis. Each hypothesis is mapped onto an end statement. If this happens, we know whether we can validate the simulation to be right or wrong.
18:33 And we're basically watching the world every day seeing which of those hypotheses are answerable at what time. And we can basically say a month goes by, we generated a million hypothesis, x percentage of them came true. This is the best way to learn about the world. Does it take a huge amount of compute to run these simulated environments at scale? And well, compute is an important piece of simulation.
18:57 Of course, a lot of the work that we do is to make our simulation be more efficient. So, a lot of our computitionally actually goes in to create the initial breakthroughs in technology. So it is actually exploring different ways to train is exploring different kind of data set. Once we have a point of view we can very quickly make it efficient. So some of the things that I've seen uh within simile as we built this company over the year is right now we have a model that's been in production.
19:30 This model used to cost about 100 times more to run than it does now. And some of it does happen because we actually found different ways to model but with the same reward model with the same philosophy but just in a way that's much more efficient at inference time. So there are these kind of tricks that we can play and these kind of scientific advancement we can make to make things cheaper.
19:53 A lot of the investment however does go to find that initial point of view. Can I ask you when you look at like serviceable market or total addressable market TAM and venture speak you know you obviously have your CVS's and your huge enterprises who would absolutely want to work with you it can also be consumers >> like regular consumers wanting to see what happens if and running their own environments.
20:21 Is this a play for everyone? Is this a play for the biggest companies in the world? How do you think about the TAM for something like Simile? >> So the start of my career really came from research obviously and the job of a researcher is to serve the humanity that we do our research obviously for our own enjoyment as well. We love the process of finding new things in the world.
20:41 But fundamentally it is a service. It is a belief that if we are able to make scientific breakthroughs, this is going to down the world down the down the line serve everyone in our society. That is how I see simulation as a field as well. So right now we do serve enterprise customers for a couple of reasons. One obviously I'll be frank there is the budget that there is a clear product market fit that we see today.
21:09 uh that does excite us and at the same time it is an amazing way to validate the technology. It is very important to us that we get the feedback loop to be as tight as possible. So we know when our simulation is right, when our simulation's wrong and we're improving it every single day. And obviously there's this side you know part here that's just as important which is I have a colleague uh when I was at Stanford uh my office next to mine uh was Pat Han who was one of the founders of Tableau.
21:45 He's a graphics professor also tur uh he won touring awards a very well-known uh person in this in this landscape. [clears throat] An advice he actually gave me and some of my colleagues was the best way to get feedback is to actually ask people to pay you. That was their core philosophy at Tableau and I also see this here. So getting the best kind of feedback matters a lot.
22:07 So enterprise market market research right now is an wedge that we found that actually have significant budget that have immediate product market fit. But down the line I do want this technology to be used by rest of our society because fundamentally what we are trying to serve is help people make better decisions. Before we move on to expansions that it could be used for when did you know you had product market fit?
22:32 You said you felt that pull. When were you like ah we we got product market fit here. Many of the Fortune 500 uh board members and their seuitees reached out in part uh they do come to Stanford to see some of the demos that are happening in the lab and they all saw the smallville demo after it got released and everyone thought oh my god if we can simulate a market like this this is going to change the way we operate.
22:57 So you could immediately sense the product market fit and really this was the forcing function for us to then say okay this is actually quite interesting. We're actually going to show and validate that our simulation cannot just be an interesting demo but it's going to be accurate. So we spent about a year actually demonstrating that we can create models of people that are actually amazing and validated at predicting people's behaviors across surveys, behavior experiments, real environment.
23:28 And we show that we can actually predict people's behaviors and attitudes 85% as accurately as people replicate their own. We put that work out at the end of 2024. And that's really what started the field around synthetic panels simulations and that's the market that we're seeing today. >> Will synthetic panels be larger than human panels in 3 years time?
23:48 The way I see it, synthetic panels will be larger than our what we know to be the current human panel market. In part because this can really raise the ceiling of the kind of questions we can answer. Well, what I see today in the market is actually quite broken. We have so many questions we want to ask about our market. If we were to release this product, if we were to have this particular strategy, this particular policy, you're a scientist, then you want to run this study or you want to try like this macro scale
24:22 experiments. You are looking at maybe 5% of those ideas get answered. The rest of the 95% we never bother experimenting with because we either don't have the ability to do them especially if it's something at the emergence scale like we literally don't have a way to run those like emerging scale experiments at the same time we don't have the budget and time for it.
24:47 So a lot of the decisions that we make as a society we base on our gut instinct. Sometimes they're good but sometimes they're very biased based on our own narrow experience. So what simulation will do is unlock that limitation and us to actually test every single hypothesis that we have about the world before we have to launch into the world. How do you balance the pursuit of the next dollar and serving customers who pay a lot of money I'm sure versus research prioritization and maybe focusing dollars there over
25:21 building out a customer success team and an FD team. How do you balance the profit maximization with the research purity? So, Simile is interesting as a company. So, Similey is a company that has a real product and engineering team. But at the same time, we are a research company. The co-founders, the four of the three co-founders are researchers. So, we have as co-founders myself, Michael Bernstein, Percy Young, Laney Ellen.
25:51 Myself and Michael Percy were all researchers at Stanford. So I led research around agents simulations. Michael was one of the co-authors of the image that really kickstarted the AI revolution and he's been a leader in humans and AI. Percy was the person who literally coined the term foundation model. And the vision for this particular area is the vision that we can actually create the next paradigm shift in AI and in the way we view technology and the impact of the technology in the form of simulation.
26:21 The reason why we are able to however operate as a research lab but also have an amazing product and engineering function and go to market function that's led by my counterpart Laney is the alignment between what the technology can do the promise of technology is so close to what our market actually requires. The better the model gets in representing people, the better simulation we can create, it immediately means better experience for our users because they'll have much more grounded, much more accurate simulation.
26:56 It is very difficult to maintain both a lab and a product company if there's not that alignment. But when there is, it can be quite magical and that's what we're seeing at Simil. >> Can I ask you when we think about the pursuit of some of the largest companies on earth? We mentioned I'm not sure which customers were able to say versus not say, but you mentioned CVS.
27:13 People always think it's like multi-year incredibly long sales cycles. Was that something that you experienced or was it a different experience for you getting and working with some of the biggest companies on the planet? >> What's been fascinating to me coming into the field of simulation especially in this market was last year when I started the company and when I so I left Stanford uh June of 2025.
27:37 So it's been exactly one year. I actually thought our field will actually the market will take about a year or two before they warm up to the idea of simulation. So we'll basically find build the right foundation for this company and for this market and we'll go aggressive based maybe towards the end of 2026 was what I had in mind and that's not what we experienced.
28:01 What we experienced was our customers were moving extremely fast also in ways that like truly made me change my perspective on corporate America. our uh leaders uh that we work with for instance CVS I've been working with this uh particular uh leader uh Shri uh who is their VP of insights extremely forward-looking extremely ambitious extremely hardworking and amazing counterpart to vision like simile but what I've also found was the pain they were feeling in their day-to-day work was so real it was way more acute than
28:38 I could have imagined that when they realize that there is or there could be an answer in this market for addressing some of those pains of very slow experimentation, budget and so forth, they are ready to drop everything and try us out. So we actually saw some of the largest customers in the world move at a lightning speed for enterprise where we saw them close deals within 3 months.
29:03 >> 3 months. Wow. Okay. That's very different to what people traditionally think. What matters more to them? speed of output. In other words, being able to get results very quickly on their simulations or accuracy of simulations. >> Yeah. Uh it is both. There are so many questions that they are truly relying on their gut decision today that if they can get some form of evidence to at least directionally guide them in the right path, then they're ready to try it.
29:31 And then they very quickly realize that oh, this is actually an amazing way to interact with a lot of data. This is amazing way to gain evidence that is actually quite accurate and they actually one of the ways we actually got some of our first customers was in the first call they actually had a finding from you know large consulting companies and they basically queried our system.
29:52 Hey if we were to rerun this what would the system say and we predicted the outcome of studies that took three to six months but just within 2 minutes that's very powerful. It must be so compelling in a customer conversation to be able to say like you did this campaign. If you had done this campaign, it would have been 12% more effective. Do you want to buy our product?
30:15 [laughter] Like it is such a good sell. I'm a seller. To be able to have that data is unbelievable. How do you think about value extraction efficiently? And what I mean by that is that if you work with like a CVS or you name any of your big companies that you work with, these are massive companies where if you're able to do your job efficiently, you can move the needle to the tune of hundreds of millions for them.
30:36 In some cases, billions of revenues. Charging like a million bucks feels like a large chasm between value generated and value extracted. How do you think about closing that chasm to be more fair? Yeah, it's a great question and I I see the market moving in this direction. One of the core premise and one of the ways that our customers are actually finding value in simile is actually avoiding really damaging decisions that could have costed them hundreds of millions of dollars.
31:07 >> So it's prevention not optimization. >> It's both. Uh but certainly prevention is a huge it's an obvious value case, right? That oh wow that could have been a total disaster had we run that. That would have costed us half a billion dollar. We ran simulation and that prevented it. That's nobrainer. And this is a this is a true painkiller in their case.
31:26 >> If you do your job efficiently, can calcian poly markets still exist for a lot of their markets? >> So it's an interesting question. I do certainly think there is an overlap here in that we are companies that are fundamentally interested in the future and helping people at least get a glimpse of what the future might be. Where I see simile come in is we are a company that is not just interested in what's going to happen but more on how it's going to happen and why.
31:57 So in that way and this is also the value proposition that our customers are most inspired by. It's one thing to simply predict, but can we actually show here are all the steps that your ecosystem is going to take to get to that particular outcome and this is the way you can prevent that or you can encourage that. That is ultimately the power when it comes to team building.
32:20 You said about the craft of team building and I think it's really interesting because it's an ongoing challenge building the best team. What have been your biggest lessons coming out of research in what it takes to build an all-star team with simile? >> So, a couple of things. One is the team has to be balanced that there are certain power that I can bring to the team.
32:46 But there's also a lot of things that I don't know. I was a researcher. I was not an enterprise seller. I needed Laney to be my co-founder to lead that part of the game. Um so balancing the team and being able to see where your team is lacking especially as we scale there are new gaps that are emerging actually seeing that ahead of time and making sure that we fill those gaps I do think is a core fundamentals of building a great team at the same time I also do think it's important that the team remains consistent in
33:19 their values and in their rigor this is more of a painters analogy um so when I was a painter. I was a figure painter. So I worked a lot with human subjects, portraits, figure studies. There's sort of this untold secret amongst figure artists, which is it doesn't matter who you paint, your subject sort of looks like the painters themselves in some ways or at least they share the similar vibe.
33:48 I think building a team is actually a lot like that in the in the best team in the team that you care deeply about. You really should see yourself in the team and for me couple of things matters the most. One is and this is the same standard I try to uphold for myself but one is are we the common denominator of success. People live through different stages in their life and they have different careers, different jobs.
34:17 At each stage of their life, were they the reason why that thing was successful? If you squint, were they the common denominator? If the answer is yes, then what that suggests is a couple of things that they have extreme degree of ownership that they are the kind of people who come in and say doesn't matter how everything else goes. I will personally make this successful and it also shows the ability to reinvent themselves.
34:44 So one of my co-founder Michael Bernstein he has had a very interesting career as a researcher where during his PhD 10 years ago he started his field in crowdsourcing collective intelligence. Then very quickly during his early years as a faculty member at Stanford he went into a different areas of AI and then now into generative AI agents and simulations and at each step of the way you could sort of see in the work he's done that this is very Michael you could see that this is the person who led a lot of the success
35:17 that's an amazing signal and another piece the second piece for me is this is a little bit more niche to myself but I found this to be very true at least to the way I look at the world. Do my leaders and my team have two superpowers that's not supposed to coexist in one person? Any expert will usually come in with one superpower or even sometimes multiple superpowers, but they're all correlated.
35:47 You're amazing programmer who happens to be amazing at mathematics, fairly common. Where I found things to be particularly compelling is if people have two superpowers that's really contradictory. >> The most common one here is actually the greatest CMOS which there are very few I count on one single hand are unbelievably data rigorous oriented scientific in their approach and then you blend that with this creative artistry imagination and they are two relatively opposing kind of mental approaches I think.
36:22 >> Yes. And it's very rare to have that in a CMO but when you have that that is the worldass CMO >> and that's magical. >> Yeah. That particular description I actually sometimes have used for my uh board member Shul >> deeply analytical but he's very intuitive and I think that's how he makes investment and that happens to be very successful. So hopefully similarly we continue on the success but in my team the kind of things that I also see as an archetype is and I also categorize myself as one of these kind of people
36:50 on day-to-day basis for instance uh Laney uh one of my co-founders she's paranoid she's somebody who will come to the table and say unless we put everything on our table today and do everything possible we'll lose we fall behind that everything will fail. But long term, she's religious. This is somebody who fundamentally believes the world is stacked for her.
37:18 Then no matter how this goes, we will make this successful. Actually, balancing those two at the same time is quite difficult because if you are short-term paranoid, then you're likely going to be very pessimistic about your future and you're you might be amazing at shorting stocks, but not great as a company builder. If you're religious, you have the opposite problem, which is you're complacent.
37:44 That you sort of feel like ah we don't have to put everything on the table today. Things will be okay. Balancing those two needs somebody who is broken in some ways that somehow they found a way to be deeply paranoid but at the same time ignore all the paranoia of today to believe that the world is going to be amazing. I think it's actually it's exactly me and I think it's actually you believe that the paranoia that you hold today helps that future state be amazing.
38:12 You know, I so often interview the world's most successful founders and they all say I say, "What do you wish you'd known when you started?" And they will say, "I wish I'd known that it would all work out and I wish I hadn't been so worried." And I think it's the worst answer you could give me because the fact that you were so worried and so you did the prep, you put the work in, you stayed up late to do that presentation that led to the success.
38:36 Without the paranoia, Laney didn't hit the quarter. Lady didn't set the urgency in the sales team. They didn't hire those extra people because you didn't know the excess demand would be there. The paranoia drives the success. >> It's a really interesting one. Can I ask you? Brandon at Mccau was on the show recently and he was like, "Honestly, researchers, they're in the tens of millions of dollars.
38:59 It is it is so expensive. Do you find that to be true? And how do you find this intense war for research talent in the Bay?" Absolutely. So the research talent is very sought after today and I have my closest colleagues and friends whose total come does range in tens of millions. Now when they join similarly I I'm fairly upfront with them it is not possible doesn't matter how many hundreds of millions that you raised um meeting them at their base salary is tricky.
39:36 However, the researchers fundamentally care about a couple of things. They care about a vision. If this idea truly come to fruition, like these are people who have literally seen open AI being the laughing snug at you know in Silicon Valley to becoming a nearly trillion dollar business. And these are people who have seen anthropic go to that same state within the past five years.
39:59 So these are people who are fundamentally aware that deep ambitious vision can actually come to fruition. So they care deeply about the vision. They also care deeply about the impact. What are the societal impact of the technology that they'll be working on and is it actually interesting to them? Do you worry about the retention problem in the valley today?
40:23 You see so many researchers move with such promiscuity if you can use that word. Do you worry about the retention problem today? >> Consistently. And in fact, I actually do view the role of leadership to be and that of obviously hiring amazing people, but also providing a platform where individual members can express their superpower to their maximum degree.
40:46 And I do think this actually part genuinely does matter. And retention can be challenging, but it can be done. And one of the core sort of a I have a small sense of pride in the way my career has panned out over the past six years or so as a researcher where Phu students often go from one project to the next and their entire co-authorship would change maybe except for your adviser.
41:12 I've had sort of an interesting career where in the past six years all my core team members never left that we all moved from one project to the next to the next together and now when I said hey I want to do this thing and build similarly I was able to somehow convince Michael and Percy who were they were actually my doctoral adviserss to actually come join me and I think a part of it is you know we worked so closely together that there's genuine sense of trust But at the same time you know I am somebody who
41:45 fundamentally believes that one it is my job to communicate the degree of confidence and trust to a team that they feel like this would work out. Can I ask you a really important question for me which is that I see a lot of amazing people in research in academia who are considering or starting a company in the same way that you did and I always worry that I'm going to finance a science project cuz science projects although great and although interesting and intellectually satiating don't always make great companies.
42:16 You've been able to do that incredibly well in the last you know year to 18 months. If you were an investor analyzing a group coming out of academia, what would you look for that would give you confidence that they would be able to make the leap from research to starting a company? The thing I actually would look for is are they married to a problem or are they married to impact?
42:42 Sometimes researchers are very much uh focused on a problem and something about that problem fascinates them but often times it's just not a good company or it's not a thesis that can really be formed into a company for various reasons. But there are researchers who are fundamentally driven by impact that they can have in the world. And for them it means finding a problem that can actually reach people, finding problems that can actually generate revenue.
43:07 And that's what drives them. you want to find researchers who are in that category. >> So talk to me. It was Shardul that introduced us. Uh huge thanks to Shardul for that. Um but there's a new funding round that's come to be in the last month or so. Can you talk to me about the funding round, how it came to be, and how you think about it for sure? So uh we raised our $100 million round uh about 5 months ago.
43:34 Um and soon after um we were preempted uh fairly recently uh by insiders. So Shu let index uh led our previous round and including Shro and some of the other insiders were looking at the market and Shro has sort of this comment that he everyone once in a while makes where he's seen some of the fastest growing market and his track record does show that he truly you know has seen different markets.
44:00 He has quite never seen this kind of traction in this kind of pool. when that is paired with a technological progress that is being made and also the amount of comput that we can also leverage to even further accelerate our progress that sort of prompted our insiders to go can we actually put in more money now than later um so that's how the initial round conversation came to be and we were now planning on raising at that particular moment but there are a couple of teams that I particularly respected in the um that if
44:36 we were to be raising I wanted to talk to and I found out that the team that I had in mind as my top of list actually was Neil Meta's team at Green Oaks and turns out uh his team actually has been looking deeply into this market and all the players how the market is going and we're actually prepared to make the investment and they're looking for sort of the right time to do so.
44:55 So I reached out and said, "Hey, this is going to be the round. Uh we're now running a process. So if you'd be interested in joining, uh you know, we have a few days to make that happen." And they were excited. So the round came together. So we raised $200 million. So it brings our total funding to be $300 million raised over the past 6 months or so.
45:18 It gives us a very meaningful uh capital to go after this really ambitious modeling challenge and building up this team. It also brings in a lot of really exciting people to the team. Trollul has been a fantastic uh partner. We actually have a lot of index connection uh at simile uh our seed actually was led by Mike Vulpi who runs now his own firm uh and trul and along with uh Neil and Green's team along with Patrick and who is the partner.
45:50 You didn't need the money I take it you raised 100 million 6 months ago. Was there a consideration of we don't need the money why would we take 200 million now? There was certainly that consideration where we netted out was the modeling does take compute and this is one of those areas where you can actually here's a fundamentally interesting part about research with research you really cannot control the outcome necessarily but what you can control is the input and the process and we were sort of at this moment where
46:22 yes we can actually significantly raise the input both in terms of data compute spend to actually meaningfully accelerate this progress. That's when we thought it actually makes sense. Totally get that. What did you not know about fundraising coming from a world of academia research that you now know having been through three rounds? Well, one actually here was coming in I was actually fairly skeptical what the roles of VCs actually were.
46:56 [laughter] What do they actually do? law. What do they actually do? How do they help? And I will be honest, I can't still quite put my finger on it and say this is no way they help. However, if you bring in the right set of people, what I have realized was they can be not the greatest partner and they can also be really strong set of mentors because I never ran a company certainly not one like this.
47:26 uh this is my first real experience building a company and I have a lot of technical experience of doing research but so much of what I need to do on the day day-to-day basis is new if there's someone I can trust then that's an amazing boost um initially um I started to work uh with Mike Vulpi and we also had other firm uh AAR who also helped uh lead our seed and these funds and Mike's team and we also work with Ishani very closely there.
47:59 They really became sort of core mentor as I operated in the field and they also were the Mike actually introduced me to Laney uh who ended up becoming instrumental as I thought about the business and I found a great partner and friend in her which also has been an amazing part of this experience. So certainly one is the VCs can actually help in some magical ways uh they have seen enough that if you're experienced VC they can actually provide the advice that the founders might not have coming in and that is one another
48:37 one here is things always happen a little bit sooner than you would expect obviously coming in I had sort of a you know in my mental model okay well if we raise seed now that means we might raise our A in about a year and maybe B in the year after or something like that all that happened within a year. Uh we raised seed and I think our series A was very soon after.
49:00 Uh and our next round also came very soon after. So I think the market is always moving perhaps one step ahead of where you are in terms of their interest in investing in you and it is useful to be prepared for those moments. That's what I've learned. Do you worry that the market is so frothy that it can get ahead of itself? Like when you announce this fund raise with the people that you have and with the press that you'll get, you'll get more interest for a next round and like it's it's an ongoing cycle and like
49:35 bluntly the hubris is very high right now. Do you think about that? I do think there's parts of market that is actually quite frothy. Yeah, for sure. There's a lot of capital going in. There's a lot of excitement. This is where I actually do care a lot about the fundamentals. Well, where are your customers? Who do you actually work with? What's the market pool that you actually see?
49:52 And what's the technology? One of the most interesting thing about how open anthropics like these companies grew was there were very strong fundamentals they could actually map out. They could actually see, oh, the models are getting better at this rate. Oh, and there's this kind of demand. Some of those they could actually foresee and some of those we can actually see at simile as well.
50:17 does the unit economics vary for you on a per simulation basis and what I mean by that is like you know if you look at say anthropic and open air and model routting some tasks require you know frontier models which are much more expensive much more tokenheavy versus others which are much easier and can have a degraded or older model and a much cheaper model.
50:35 Is that the same for simulations? Do different simulations cost different amounts in terms of compute token usage associated? >> They do. Um usually when you have simulation that is trying to answer something that's much more complex uh or something that's let's say you want to actually understand all the downstream implication of your decision or you want to do market segmentation study across all of the US much more expensive.
50:59 What I also have seen however is in it is in those simulations where we actually get higher ROI for our users because those decisions are some of the most costly decisions if they fail to make the right one. So this is actually interesting for simulation as a field. So you we've always seen as a community the inference costs going up and up and up and we now have these thinking models that are thinking for like half an hour a day and actually start spending like token maxing and you know spending you know a lot of
51:31 money on just running this process. I actually do think simulation could actually be the next frontier of that where in my vision I think there's a world in which in about two three years we're running a single simulation session that's going to take 1020 million to run a single session but it's going to be so valuable that people will pay $100 million for it that's where I see it go and that would be for the world's largest enterprises that would be for a government or whatever that may be >> especially on the sort of
52:08 the high end of the spectrum that's what it would be >> what cannot be simulated today that you think will be possible in 3 years so for me it's actually a little bit less about what cannot be simulated because I actually do think everything that we want to simulate we can actually create the initial uh proof of concept however as we all know one of the core challenges of AI is actually bridging the proof of concept with real value productionizable technology.
52:37 So that's actually the chasm that I see. So interesting thing here is I see the world of simulation going into this world where we are creating this very complex multi-agent simulation or we're running a very long study with many different steps of simulations along the way. But we actually started the field from multi-agent simulation when we created the small game town.
52:59 That was fundamentally that vision and it sort of also makes sense because we did that because we my myself Michael and Percy we sometimes sit together and do this exercise called time machine game. If we were to write a time machine go to 10 years into the future what's going to be the craziest thing we're going to see and can we do that now and that was the motivation for running the smallville experiment.
53:25 So this can be done but the question is can we eval evaluate uh the efficacy of these simulations. Can we actually propose this as a scalable productionable system that people can actually rely on for making their decision? That's the chasm and that's the thing that we see getting uh bridged every day. A huge part of it also is getting you know models to be better creating bigger simulation making the system more scalable.
53:49 All that becomes a part of this. Can we play a time machine game with me and you? Let's do it. In 10 years time, what is the craziest thing that you can see happening? >> A lot of things, but one thing I would actually say is I am someone who is fascinated by history of technology and analogies that we can draw from it. What I see today that's prominent in AI space is what I consider to be the CPU of intelligence unit.
54:20 you have this one language model that's really large that's very smart that can do very complex reasoning tasks that's like CPU what I see coming and what I think simulation as a field can offer is the GPU of intelligence unit as I mentioned before simile does not care about creating really smart super intelligent machines what we care about is creating models that are as smart as we I I fail at a lot of things.
54:51 I want to make sure that the model that represents me fails the same way. But the beautiful part about people is individually we have so much diversity, so much different takes in our world that makes individuals so interesting. But also when they come come together as a large collective the emerging phenomena that we're able to draw out is some of the most wonderful thing that we can see in our world.
55:20 Creating a society creating an amazing process that actually allows us to make all these achievement. Can we actually replicate that in simulation? I think it's going to be quite inspiring. What's the crazy prediction then that every single person will have a replicable twin that acts and behaves like them in a simulated world? I think that's the vision.
55:41 The vision here is again representation at scale. We as a society have found over the years many different ways to represent our members. Sometimes it's a form of government, sometimes is actually companies company. We are as a society allocating capital to make sure that they serve the society and the needs of people. But if we can actually create an artifact and that in a much more scalable and granular way represent all the individuals.
56:08 What are the new kind of policies new kind of companies that can be created on the basis of it? I actually do think it's quite interesting. If I'm a hedge fund, is this not the most obvious buy in the world? If I'm looking for alpha and edge on everyday activity, sign a million dollar contract with you and get unbelievable insight. Yeah, maybe similarly we'll actually own a small uh hedge fund down the line.
56:37 >> That's a cool idea. Would you be down to do that? >> Well, it turns out we actually uh do have uh quants uh in our firm. Um so uh some of the members who have joined actually do have more quant background and I think right now they're joining not because they actually want to start a quant firm uh at simile. They actually join because they actually see the vision of simile very much well aligned with their passion and interest which is to model the world.
57:00 Uh but down the line I think it's actually an interesting idea. Is there a world again I'm just thinking crazy time machine world. Is there a world where you are so efficient and so good that actually stock markets become uninvestable because the world is skewed to Simile's hedge fund or similar providers and actually it is not a fair marketplace. I think especially with obviously if we were to assume that we're going to have some form of AGI and if we were to assume some form of perfect simulator I think a lot of the
57:35 world a lot of the things that we assume to be true about our world I think will change certainly one of this could actually be the stock market what else do you think we assume to be true today that you think won't be in 5 years time >> what I think we generally assume to be true about the world that we live in is that it is fundamentally impossible to get everyone's perspective.
58:00 Therefore, we need representatives of these people to approximate their perspectives. So far, that has worked in some ways, it has failed in other ways. I actually don't think this is a limitation we have to suffer through in the future. I think there's a world in which we can truly create a layer that becomes a representational layer of our society and of our collective intelligence.
58:22 is the future of love not also simile and what I mean by that is like if you were able to create effective simulations of yourself dating itself could be much more efficient if I could I've got a girlfriend so she's watching and but if if you could date a 100 people at the same time >> for the first date of course not not onwards you get in a lot of trouble for that June um [laughter] but if you could it would be much more effective at finding the one for you who could pass through to the next stage.
58:52 Do you know what I mean? >> I I get that. Well, look, I think love comes in different forms and I think just like as we discussed today, people have such degree of diversity. Uh personally, I I am a bit of a romantic. I'll be honest. Um and and this actually goes to the point that I mentioned about long-term religious. I actually do believe that love will sort of uh find at least in my life, you know, find its way in a more organic way.
59:26 Um I actually do think the way I meet the person I personally do care a lot about. Um and the fact that we sort of have shared a journey I personally care a lot about. I think that piece of humanity will I don't think ever change. Actually it is experiencing things together having that shared memory I actually do think is fundamental to the way we form trust.
59:50 Um this is also we talked about process a lot. This is actually a part of it for your users using your simulation. Can you actually bring them along in this process? I think finding love is it's a little bit like you're co-ounding your life with this person. So can you actually find a process that actually would bring them along in this process of living?
01:00:08 I actually do think does matter. Oh, you're so romantic. Unfortunately, you know what I what I'm thinking, dude? I'm a content person. I'm a content person, an investor. Weird mindset actually in both ways. There's a show called Married at First Sight. You might not know it. It's where you marry someone on first sight, but I'm just thinking it'd be the most phenomenal advert for Similey if you could do the perfect marriage at first sight because of simulations that have been run before.
01:00:34 But, you know, I'm just leaving you with pearls of wisdom that I think would be great. Uh, we're going to do a quick fire. I say short statement. Well, who's the most underrated AI researcher today? And there's so many, but I actually do really think um there are some incredible people who are working at these uh larger labs whose names are not known because they work at larger labs and they don't publish.
01:00:57 But I think there are some really incredible people in there. What area of AI do you think is particularly overheated today? I do think Neolabs without a clear vision for how they're going to impact the world. I do genuinely think there's some risk that they will turn out to be interesting research project but not a viable company. >> If you were investing in my seat today, what part of the AI landscape would you say is underinvested and most exciting?
01:01:26 You can't say simulation. >> I fundamentally believe that um for AI companies in the future, you have to have interesting data strategy. Do you have access to data that no one else has has access to? Do you know how to collect data that is very hard to collect? When you see those opportunities, I would invest. Uh right now, aside from simulation, robotics is sort of an obvious place where this has become the case.
01:01:50 Obviously, robotics there's a lot of money already going in. So, I wouldn't say it's underinvested, but I also do think it is a quite interesting area. I also do think uh aside from the core sort of um robotics or AI space, the inference layer but also chip layer, the hardware, I do actually think it's quite interesting and it's a very hard area for people to crack into but there are a couple of teams that have done I think an exceptional job in the recent months or years.
01:02:15 I think they're quite interesting. >> Who do you think those are? >> So recently edged uh came out of uh their stealth uh quite bullish on their team. I think they're going to be exciting. Uh, so that's one >> final one for you. What's the kindest thing that anyone's ever done for you? I think it's a nice note. No, >> I'm somebody who actually needed a lot of help uh throughout my career.
01:02:38 I, you know, I didn't come in knowing everything. Well, certainly I don't know everything now. Uh, but I also didn't come in as, you know, someone who's was an obvious candidate. The one person I quickly call out was when I graduated from college, I moved to PaloAlto living in Samarai's garage. I didn't have a job because I was trying to run a startup that doesn't really go anywhere.
01:03:00 Uh but that was a moment where I really felt lost. I could sense in the air that AI wave was coming and I realized that if you want to be a surfer, you need a wave that you can surf. And I want to make sure that you know I when the AI wave is here I want to be there to ride it and I want to help create the wave in the first place to ensure my seat in it.
01:03:25 But I had no research background. I done no research during my undergrad which is quite rare. I would like reject you know you know if you're a PhD student applicant and have no research experience during your undergrad unfortunately it's very hard. So I actually messaged a bunch of people and there's this one professor at Stanford, Mary Ward. She's a theory professor.
01:03:44 I happened to graduate from the same college as me. She replied and I still don't know why. I think it was truly out of kindness and the fact that we're from the same school and she thought, well, okay, here's a student who is seeking advice. I'll at least spend, you know, half an hour with the student. She very graciously spent a full morning with me and just talking me through like how I should think about AI space or how I should think about research and she actually connected me with an initial set of people that I
01:04:16 started to work with and learn from. So that initial set of people who came together to help give me advice and actually let me have a foot into this area of research it really was not an obvious choice for them. I really didn't think I deserved it but that was the bet that they took. I think truly for you know purely for their own kindness and I'm very grateful that they did.
01:04:40 >> Never forget the first believer [snorts and gasps] uh June from uh quant funds to simulated worlds to love. Uh this has taken many different twists and turns but thank you so much for joining me. >> Thank you for having me.
The best AI companies need defensible data strategies, especially access to unique data or the ability to collect difficult-to-obtain data.
Build models that represent people's viewpoints and simulate individuals, groups, markets, and ecosystems.
Provide organizations with simulated populations and counterfactual analysis to improve decisions, test strategies, and avoid costly outcomes.
Represent populations through simulated people so organizations can ask more questions and test ideas that would be too slow, expensive, or impractical with human panels.
Charge enterprise customers for simulation and market-research capabilities that provide evidence, speed, and decision support.
My fundamental thesis here is for AI companies of this generation, you need to have an interesting data strategy that's going to be defensible.
The world is our ground truth.
Are they married to a problem or are they married to impact?
Identified as an interesting AI investment area, although substantial capital is already flowing into it.
Identified as an interesting area outside core robotics and AI companies.
Identified as a difficult but potentially interesting area with some teams making exceptional progress.
Simile co-founder and Stanford researcher; described as a co-author of work that helped kickstart the AI revolution.
04:14Simile co-founder and researcher; described as the person who coined the term foundation model.
04:14Stanford theory professor who advised Jun Park and connected him with people in AI research.
01:43Investor who led Simile's seed round and helped mentor Jun Park; also introduced him to Laney.
07:19Worked closely with Jun Park and Mike Vulpi's team in an advisory and mentoring capacity.
07:57Company creating foundation models of human behavior for simulations of individuals, populations, ecosystems, and markets.
10:44Frontier-model company used as a comparison for general intelligent machines and company growth.
03:02Example of an enterprise that could use simulation to understand and change future sales outcomes.
08:59Enterprise customer example for surveys, customer feedback, market research, and simulation.
12:21Company founded by Pat Han and cited for the philosophy that asking people to pay provides strong feedback.
03:40