← All transcripts

OpenAI’s Plan to Make ChatGPT the Everything App — Akshay Nathan, OpenAI Transcript, AI Summary & Key Points

Latent Space · 3 days ago · Science & Technology · 01:10:57 · EN-US

📄 Transcript

Searchable transcript of OpenAI’s Plan to Make ChatGPT the Everything App — Akshay Nathan, OpenAI — Latent Space (01:10:57). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Latent Space. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 I think the bottleneck becomes like sort of like ideas and taste, I guess. Um I think because anyone can [music] can build now, I think um it's really is the era of like bottoms-up ambition. And because there's so much to be built, like you're always going to be bottlenecked by the amount of ideas and amount of things that you're doing at any given time.

00:20 >> One interesting part about ideas is like they're not like in a vacuum. It's like not They usually come from somewhere and like, you know, in product development like they're coming from >> [music] >> talking to users or reacting to, you know, friction that you're seeing or feedback. There'll always be value in in these like generalists that we talked about, like, you know, closing that loop and and and coming up with these ideas that are grounded [music] in in that feedback or talking to users, whatever it is.

00:43 >> Before we get into today's episode, I just have this small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way.

01:05 But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better.

01:28 Now, let's get into it. >> [music] >> Okay, we're here in the studio with Akshay from OpenAI. Welcome. >> Thank you. >> And with our trusty co-host Vibhu. We So, you recently launched ChatGPT Work. You lead core product engineering. You know, it's been a long journey into into all this. I find it very interesting that you started with no-code or low-code with Walrus and Airtable.

01:53 And to some extent, ChatGPT Work is kind of like the super app of super apps of well here is the ultimate no code you just write a prompt. >> Yeah. Yeah, it's it's it's funny how things come like full circle. I mean, I think for a long time my career I mean I started my career working consumer fintech but then after that like there's this hypothesis that you know the things that we were able to do with code like as engineers like if we could bring that to many more people in a more accessible way then that would be

02:19 truly magical. We were working on a startup so I actually found it like before LLMs before vision LLMs on how to do automated testing with AI and it was it was kind of jank back then but you know doing what we can and then and then worked at Airtable for a while on you know the same thesis that like if we can bring a database or the parameters behind a database to people that'd be really useful to them.

02:42 Um but once I think LLMs came onto the scene it became clear that like this was like the missing piece like the missing technology required to like bring the magic of code to everyone without them having to know what's going on underneath the hood. And so like I think this launch and you know all of the stuff that we've been up to is like um the manifestation of that.

03:02 >> How was stuff when you joined? So you joined OpenAI in 2023. Now we've got you know so much more stuff. So ChatGPT, Codex app, ChatGPT for work. How have things changed? >> I actually think the more interesting thing is how things haven't changed. Like I guess like one I I joined I remember when I joined it was like 500 people. One thing I was worried about was like I was looking for something you know more early stage and like was it going to >> [laughter] >> start up enough and I joined I was like dude this feels

03:28 even more start-upy than I could ever imagine. And like that that really hasn't changed even till now. I mean I think the like level of like bottoms-up ambition and like the ability of anyone to like you know do anything or have an idea and and and ship it is is really cool. But on the like sort of mission side I think what was really compelling to me is this mission of you know bringing frontier intelligence to everyone like building AGI and then bringing it to everyone and um I think acknowledging even then that like

03:53 that vision is going to, you know, not be a linear progression. Like we're probably going to like try different products and and have different things that that succeed and don't. Um, but the vision has stayed the same and the mission has stayed the same and we're starting to see the pieces um, fall together and and that that's really cool. >> Uh, you worked on enterprise.

04:08 What a lot of people never touched chat chat chat chat GPT for enterprise, God. Um, what is something that you learned from there that you're bringing into your work now? >> I think how there's no like one-size-fits-all solution in enterprise. Um, I remember in the early days of chat GPT enterprise like we would talk to customers and like everyone that was like when I think it was a year after chat GPT was released and everyone was so excited to bring um, you know, AI into their enterprise and like I thought there was

04:38 all these teams that that were being set up as like you know, the AI deployment team with like these enormous budgets and if you asked anyone like what were they excited about? Like what were they excited about solving? Like at first you'd get like you know, kind of like the the baseline answers of like yeah, we have all these contacts and data and all this stuff.

04:52 But then if you ask them like you know, what was like a discrete use case of like they want AI to enable in their in their workplace, you get such a different like variance like explosion of different types of answers and it's interesting like you know, you using like these models and these products, you you have this box and you can say anything to it, which is the magic.

05:12 But it's on the flip side, it also means that like you don't know what to do with it. And in enterprise, I think a big part of that is like actually meeting the users where they are like what use cases were they trying to solve and then actually teaching them how they can use AI to like gain leverage there. >> Do you meaningfully differentiate that from forward deployed engineering?

05:28 Or >> I I think there's like the the like go-to-market side of it and then there's like the product side of it. >> I think >> You need to be more on the product side, yeah. Um, and I think like however good we get at FDE motion, like I think at the end of the day if we have a user who's like looking at their computer or looking at their phone, like it's our job in the product to like be enabling them and showing them where to go.

05:51 Um, so we're really excited about that. >> Do you think there's been changes, you know, over the past 3 years of adoption? So, there've been, you know, step function changes, you have reasoning models and whatnot. There's still the same problems of enterprise has black box to know what to do with it or have things changed? >> I mean, we're seeing now that like there's this huge uptick, right?

06:10 Everyone's like extremely excited about it. It feels like, you know, many people like millions, hundreds of millions of people are using ChatGPT. They understand like how generally to work with AI. But then like every time like a new capability gets unlocked. So, now like we're seeing with agents, like there is probably a contingent of like early adopters still who, you know, truly get it.

06:29 We're like, you know, we you can do anything. You just have to make sure the right context is there. It's connected to the right tools and then that you're supervising it, but like anything is is is possible. But then there's like this like 10x or 100x bigger market or like they don't yet get that or they don't yet see that. Um, and so I think that's the next stage here.

06:47 Um, so I guess to answer your question, like I think the adoption is there and and growing fast, but I think the opportunity is like far, far bigger than that. That's where we want to play, especially with ChatGPT work. >> Yeah. Uh, well, let's let's uh skip ahead to ChatGPT work. Only like a month ago or so announced. What was the sort of decision process that led into it?

07:06 You know, there was this overall merging of the super app. Is Is that what we're officially calling it? You deprecated the browser as well. Just I guess summarize your last like couple of months of working on this thing. >> Yeah, it feels like forever now, but I guess it's only been a few months. I think maybe the one one like impetus that like is most salient is when we released Codex or even internally had Codex.

07:31 Like it was really surprising to us. I think we recently put out some stats on this, but there was this like real inflection of like adoption among non-developers at OpenAI. And I, you know through this product development process like we go to like these UX our sessions to talk to people internally and the thing that stuck out to me is like one like you know you go talk to like strategic finance or marketing or whatever and they're all using Codex for you know their use cases that that that part's cool but the thing

07:58 that really stuck out to me is how proud people were that they were using Codex like how like >> It's like I'm not supposed to be using it but I am. >> was that it was like that they were you know early to this like new thing but it was also this thing of like they felt like they had a superpower right and what we recognized then is that like the the power of Codex power of agents like we already had this massive distribution base of people who have you know come to know and love ChatGPT like how do we show that to

08:22 them like how do we bring it to them which is like a hard product problem and it's like a tricky thing right there's many ways you can go about it and so that's what we called the merge and the super app over time and and and ultimately launched it in ChatGPT work is how do we do that but it came from that initial realization that like the the power was not only for developers like much much earlier than probably even we thought like it could be extended to to everyone.

08:45 >> How do you see the products differently so like who is it for right so Codex started out even CLI then app now there's a merge of ChatGPT Codex and ChatGPT work so is it the opening for the average user for enterprise for work how how do you position it? >> I think we want to get to position it for if you're doing worky related things >> lack of a better word right?

09:10 [laughter] >> I think productivity is like a is is actually what like the pillar that I I support like that's the name of the team and the reason for that the reason we call it productivity and not like you know enterprise or or like work or something like that is because there's also personal productivity right and like I think ChatGPT work is I've seen people do things in their personal lives that you wouldn't classify as like work technically but like these agents are you know super capable for like one one recent

09:34 example that someone posted about on our slack is like someone had like a missed package, like they didn't receive it and then they got like the picture of it you know from Amazon or whatever the courier was and they like asked ChatGPT work to like find out where that package is and like the agent you know is extremely tenacious and like they like took the image and like looked at a bunch of like listings around their neighborhood and figured out exactly the apartment complex in which the package was like gave them

09:57 this information. And so like I think there's all these things that like you you know worky or productivity related things. I think that's what we want the product to be. You asked about Codex. I think we think Codex is you know a durable brand but we have a principle that like the user you know we don't want a user to get stuck in a tab or an experience where they don't get the power of the product.

10:16 And so like basically everything that you can do you know in the Codex version of the product on on desktop you can do in ChatGPT work and vice versa but we made some opinionated product decisions on like you know how much of the get state if you're going to get repo do we want to expose to the end user or how much do we want to make the the experience of seeing the agents thinking like diff forward so that you're getting you're getting exposed to the diff side of the button.

10:40 And then like on the safety side like how do we want to think about like sandboxing and making sure that we have the right defaults in one one state versus the other. So um there's like these some opinions that go behind that but we do we do want we don't want the user to need to choose which experience they're in. >> That is a good goal for AGI right like people don't want like to hide to choose what version of AGI they want they just want the AGI to decide for them.

11:02 Um can I get a answer or like it's not super clear to me is the Codex harness and the ChatGPT work harness the same? Is it just UI affordances or are there actually prompt level or even even deeper differences? >> So the harness is the same the harness is shared. Um on in both of the products we made improvements to the harness to make it good for knowledge work especially as it relates to plugins or computer use or artifacts.

11:29 You get that power regardless of which your experience you're in. On the UX side, there's opinionated takes that we have. When you're in Codex mode, what the UX should be how the UX should behave. And some stuff around the sandbox that I mentioned, but the underlying harness of capabilities should be the same. Actually, I'm just kind of curious, maybe we can uh is there a query that we can run that would look different in the the two modes?

11:51 Yeah, I tried to create like ask it to create like a retirement calculator spreadsheet or something um and in both in both modes and then in Codex mode um you might have to be in a in a repo for this, but you'll see like the diffs of like the the sheet that I was creating and stuff like that um and the file edits. Um but in work, you won't be able to see that.

12:09 I think this that's super clear. And then also the other thing I wanted to dive into was your uh the productivity team. Uh what else is there First of all, you know, what are the top-level teams other than productivity? Isn't productivity everything? So, you know, we have a team focused on uh on ChatGPT like the the core chat experience um for consumer, which is like, you know, not I think all productivity like there's people are using ChatGPT every day for search to, you know, figure out how to write messages to loved

12:42 ones, to think about um how to like learn a new topic, etc. And so, there's so much more inside to create images. And there's so much more in chat that you know, the hundreds of millions of users are using that um you know, obviously that that warrants like a a very dedicated effort. Um and there's teams focused on enterprise and infrastructure and API and stuff like that as well.

13:02 >> I will bring it up. Yeah, so I have them both running. >> Yeah. >> This is work. There's a Codex version here. I picked 5.6 souls, so this will take a while. I think I think we'll just keep it in the background and you know, as as they finish, we'll look into some of the differences. >> Yeah, but immediately I think if you flip back to the Codex version, you'll see that uh that it seems it seems good.

13:21 Yeah, exactly. They're like dynamic on it seems that you're in a Git repo. Um and you might miss some stuff because some of it is like in the actual chain of thought with with those changes and how we display that, but >> Is is there an unintuitive like is there a thing that you wanted to ship and then you got feedback and you're like no let's not do it.

13:40 Like what's the thinking behind that? >> In uh ChatGPT work? >> Yeah. >> I think one direction we could have gone with this is like keeping the experiences like completely separate. So it's like why why exactly like different apps or even in the same app like different completely different experiences. Like why merge it all? Like what is you know Codex obviously people love.

13:58 Like why why bring these products together? And I think the intuition here is that like all of our jobs are like changing dramatically with AI like frankly every few months. Like I feel like I wake up and then I'm like doing a completely different thing that I was doing a few months ago and my my hypothesis here is that or I should say our hypothesis is that like part of what we're we're building in this this technology is giving people leverage.

14:20 Like you know the things maybe it's the more mundane parts of your job or or parts of it that like if you were able to automate you'd be able to share more ideas faster or whatever like you're able to do now. And because of that like that might actually blur the lines between someone who's like only writing code or creating strategy docs or you know planning events or um helping with marketing or doing podcasts or whatever, right?

14:41 And so like these things are going to get blurred over time. And so like trying to draw a hard boundary based on like the who you are is is is going to be is going to be tough and like we should enable users to choose, but we shouldn't box them in. And so a lot of the work that went in here like you know keeping the primitives the same like for example plugins are like unified across um this product and ChatGPT and the cloud was because of that.

15:03 It's this this thesis that like eventually things are going to come together and and we don't want to be like we want to be prescriptive about when to be in either experience, but we don't want to box anyone in. >> I wonder if there's users who are very tuned to the old ChatGPT harness that is effectively now replaced by the the Codex harness. I can't imagine what that was, but maybe they're more the more conversational side.

15:27 Can you compare and contrast the the two harnesses cuz only you've seen it? >> [laughter] >> Yeah, I I mean, I think ChatGPT the the existing harness like still exists today. It like exists in this app. >> The classic, right? >> You just start a new chat and you don't go under work, right? >> Yeah, if you start a new chat and go to chat, then you're you're talking to ChatGPT with the existing >> do another Yeah, I guess you know, instance.

15:50 >> Yeah, so this one's not going to code. Or it's going to be in line. It's not in a in line of sandbox. >> Yeah, actually we do try to push you to to go to work if you're creating a spreadsheet. >> And this is a router decision. Sorry. >> Is it a router decision? >> This is the decision that, you know, the model is making and then like, you know, it sees that you're able to or you're trying to do something that would be better served in work mode.

16:08 But I I think your question was like, what what are the advantages of like the the chat like ChatGPT chat harness? >> It's more broadly like I wanted basically do an oral history of harness engineering. >> Mhm. >> Right? You know, the ChatGPT harness lasted us from let's call it the 01 era until now. And now it's being replaced by the Codex harness effectively.

16:29 And they're they're overlapping somewhat, but I'm curious what changed if, you know, if there is. >> [laughter] >> My perspective on this is like there's there's there's sort of like a a constant process of like divergence, convergence, divergence, convergence. And and chat, like many of the use cases I was talking about before like, you know, search or learning, I think we're we're really optimizing for latency and optimizing for personality and like different things that over time like the product The reason people

17:01 love ChatGPT is because we've been optimizing for those things and working on them for so long. Codex, what we learned was that like if you give the agent access to this infinitely flexible environment as a computer, you can do really, really powerful things. And so when we think about like, okay, well, for for knowledge work, like what is which mode should we choose?

17:21 It was like it felt more natural to us to bring that to this like computer environment, and you know, maybe abstract some of the details of this computer away from users who might not be used to that, but like give them that same power. Um but ultimately, I think that we want the power in in all places, right? We want to meet people where they are. So, I'm sure there'll be work down the road in order to to get things to be um equally only capable in in all scenarios.

17:45 Um but it's just a question of like what we've been focusing on the product on historically, and what we're focusing on now. >> I think alongside that, outside of just Harness and when to use Codex and GPT at work, there's also the new models you've released, right? Um any guidance there? So, people love to min-max what to use, like only use Terra on high reasoning versus uh for this, you know, you want to use Soul here, ignore >> There's 32 options.

18:14 >> Yeah, yeah, yeah. Um >> [laughter] >> but um that being said, you know, for people that are expanding, so uh productivity trying stuff for work that don't have the breakdown of what all this is, um what's what's the advice, right? >> Well, I mean, I think before the advice, like the first thing is like none of this would be possible without these models.

18:31 Like I think you asked earlier like, you know, what was like the inspiration for work, and like, you know, earlier on, like I I mentioned like what we were seeing with Codex, but that was also because of the the models were getting infinitely more capable. That's happening again. I think it's like another step function jump now. And to answer the question on advice, like we want the default to be the best possible.

18:51 Like we want to be opinionated about the default, and so we've we've chosen a default that we think is going to be the best for everyone. And you know, we have for power users options under the hood. We could one could argue that there might be too many right now, and we're you know, working on simplifying it. Um but you can extend, you know, the reasoning level, and you can change between the different model classes if you need to, but the default should be the best for for most most use cases.

19:15 So my advice to most people would be to stick to that. And then, you know, if you reach a situation in which you think that you could you want to try um a different configuration, if you're not seeing either the the efficiency on the on the cost side or or the the quality on the intelligence side, then you can change the defaults and see if you can get something better.

19:34 But but we think that the default should be good enough. >> I have uh I I'm just going to run something by you since you you have way more experience than me. I've recently been doing Soul light but on with goal. With the idea that the goal basically augments the reasoning effort but with more terminations and turns. >> Mhm. >> Is that a good way to think about it?

19:53 As opposed to Soul ultra or Soul, you know, extra high. >> Yeah. It's hard to say >> because it's like an interaction effect. >> Exactly. It's like there's a preference on, you know, for you as an individual, like how do you like to collaborate with the models? Like how many of those like terminations, as you call them, do you want where, you know, you can uh steer or make sure that it's doing the right thing.

20:15 I think generally people should try whatever works for them. Um I think that like using ultra or the like multi-agent setups are best for like when you have like tasks that are either incredibly complicated like open open explorations or very parallelizable. I think even for tasks, using goal I think is best for for tasks that you know that you'll be able to make consistent progress in a way that's verifiable over time.

20:44 But I think for most tasks, they actually don't fall into either of those buckets. Um and so like at least when they're starting. And so that's why I think the best first step is like trying it with the the default configuration and then seeing like where you want to go from there. >> Right. You guys worked on a slider which is actually super helpful for reducing the amount of panic.

21:04 >> Yeah, yeah. [laughter] >> It's nice on mobile at least. There's a nice ladder here. >> I haven't tried it. >> So, you you have you have the advanced view there, but if you click advanced view Yeah. Yeah. >> Just a simple ladder, yeah. >> Very very pretty, very colorful. >> Yeah, the idea was here here was like reduced to like one dimension, even though there's multiple dimensions, right?

21:21 Try to project it onto a single dimension for the user. Yeah. Like, you know, something from that represents like, you know, speed and efficiency on one side Yeah. And then like sort of like quality and thoroughness on the other side. >> I am just puzzled that it uses Soul so much. Like, the lower bronze >> I think it's Spider, if I'm not mistaken, is Oh, it is.

21:39 >> Yeah, you see our side. So, they they preset Terra to only be the light one. >> I see. >> But like, I think a lot of people actually would More people should use Terra. One, because Soul keeps running out of capacity. >> [laughter] >> I'm the reason, you know. Here's 10 minutes of our >> There you go. >> retirement calculator. >> Oh, that's the Excel thing we're Oh my god, look at that.

21:57 >> is work, and then Codex is still cooking, so we'll get back into it. I think it'll be interesting to actually see the thought process, the reasoning. And also, you know, I guess this is 8 minutes on work. Codex is still cooking. >> Yeah, and by the way, I So, I Do you know Gabriel Chua? He's part of the Open AI safety team. He showed me this, and I was like pretty shocked that this looks like Excel.

22:19 Yeah. It It edits Excel files. You never paid an Excel license. Right? Like, but somehow this is This is like kind of workable, and it's agentic Excel. >> Yeah, I mean, one of the big like pushes that we made for this launch was like artifacts, right? Like, both on the on the model side, like I think if you compare this with with 5.5 and and and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of these artifacts.

22:42 And then also on the product side. >> The UX side is also crazy. Like, hosted sites and whatnot, no no longer needing to host your own little webpage. Like, >> Oh, I have a story about that. I can I can do a separate thing. I'll I'll need to take the the the visuals here, but we'll we'll we'll cut to that later. Was there code training, I guess, because you were moving making this big move and you launched 5/6 on the same day as ChatGPT work?

23:08 Was there influence between the model training teams and the harness teams or did they launch days just happen to line up the same day? >> I think the you know, we collaborate heavily with the research teams and I mean I think that's like one of the most magical parts of the job, like the most fun parts of the job. But yeah, I mean just just using artifacts as an example like you know, a lot of what you're seeing like underneath the hood, there's a lot of work that went into making sure that like you know, we had the

23:32 right infra to be able to train the models to get better at this and then on the product side like had the right experience for users to be able to collaborate with the model on artifact like this. In fact like this whole viewer like the intuition here is that like you know, it's it's not necessarily that you wouldn't need an Excel license. This is stage one, right?

23:49 Like this is probably not what you meant when you're like making a retirement calculator. You want to iterate and like when you're when you're seeing it and if this thing is high fidelity to like what you'd actually see and and or what your your co-workers would see if you were to send this to Sean. Like that that I think makes it so easier and makes you trust the the the product in in terms of iteration.

24:07 >> When you say co-workers would see, do you see a multiplayer multi-team collaboration with artifacts? Any any things you guys think about? >> share it, right? Yeah. >> Yeah, it's something that you know, we're actively thinking about. One thing that you know, we've noticed internally without talking too much about the road map is that like I there's many times when someone will ping me about something and I'll ask ChatGPT work the question and then I'll ping them back the answer.

24:32 >> Like the simplest would be you know, the three of us are just all on one hosted. >> Exactly. And I'll think about like you know, like was I required in this loop or maybe it was you know, rephrase like what they were asking or pulled from some context or whatever. But like you know, when I gave them back the answer, that process was also lossy, right?

24:47 Like I gave them just like my interpretation of what ChatGPT work cooked up. But like underneath the hood, there's so much context like in the real world and stuff that that could be interesting. So like the answer was preemptively respond to every inbound request. >> No, it's just like literally like this is what I do sometimes as my job. >> you copy paste and then you you know, you're just a message forwarding service from AI to AI.

25:09 >> I think it's interesting, right? It helps people understand the capability of what you can ask and delegate that often times people don't realize until they try or someone shows you and then you're like, oh okay, okay. >> I see. Yeah. I think it's also there's also like a light security issue where like basically you're the permissions layer. Like yes, I could query everything that you query and I could get an automated response, but maybe I'm not supposed to see it.

25:30 Yeah. And then yeah, there's no way I would know because I'm not supposed to know what I don't know. >> Especially as like, you know, with try to be work for for asking you to connect your plugins and you know, it's pulling from your local files and stuff like that. Like the the amount of context that the agent has access to is like deeply personal and like that's something that like we need to preserve.

25:49 So that'll be definitely a challenge. >> There's Excel, there's PowerPoint, there's Docs, you know, the grand trio of work. What other formats of work do you do you think about? You know, like obviously you worked on Airtable. Is there a future where there's like OpenAI Airtable? Like, you know, like what what what does that look like if if you ever ended up doing it?

26:09 >> It's a really good question. I think um I mean, one that you didn't bring up was Sites and I think that was a big part of this launch. There's one one side of Sites that I think people commonly talk about especially on Twitter and stuff or X of like, you know, this sort of like prototyping tool. Actually like we saw that happen with this launch even the the model slider that you guys were referencing earlier, like that was developed almost fully in a site.

26:30 Like, you know, the the collaboration between design and engineering and product on that was like on a site where we could play with, you know, the the the affordance and and and figure out how it feels and and all of that. Um but the other aspect that that I think is a little bit less talked about is like Sites has like an an artifact for for knowledge work.

26:50 I was actually talking to someone the other day who was on like our our corporate finance team and like we were mentioning how like now when they have these reports that they're they're working on as a team month-to-month, historically those things were in in slide decks and in spreadsheets and now they're just in sites. Like sites is the mechanism that they collaborate across the team.

27:07 And the reason is cuz it's like it's like somewhat higher bandwidth. Like, you know, at some these tools like PowerPoint and Excel are like infinitely flexible, but at some point you reach the boundary of like either as a human you may not know how to use some feature or something or the product itself doesn't support it. But in the case of a site you can do anything you ask for anything and you can get that.

27:27 I mean, once people see that magic I think it's been really valuable. >> Yeah, let me show you my my case study. This involves all the hot topics including ChatGPT work but also 5.6, token billionaires and token maxing, and sites and auto research. I'm a fan of this game called Strata. It's a basically it's like a little board game that you that you play with a physical blocks that come on top of it like that.

27:48 So over the weekend I I I took like 30 photos and just threw it into ChatGPT. 1.7 billion tokens later outcomes this site with a fully playable thing with with 3D block placement and everything. Because it requires physical blocks and I needed friends to train on it so they can get better so I can play against them. But also I could also do things like train an AI on it.

28:11 And that that's that's your auto search. That gets into auto research. So you want to train your own AIs and then make sure they self play against each other. I need to set both the AIs. So this is AI versus AI and they're they're going to self play. Obviously that the AI start out bad and then you want to define a loss function and and get good. I wasn't going to supervise all this.

28:33 I was at I was down in San Mateo attending a conference. What I ended up doing was auto researching and and on this and creating benchmarks. And there was just way too many parameters for me to to read. So I started asking it for a site and it's created this uh uh this this this lab panel. Where is there is there a is there a shortcut for a site that it's created?

28:57 >> Uh you should be able to go in the sidebar to sites, top of the sidebar. Uh left sidebar. >> This one? Oh, left? >> Yeah, just scroll all the way to the top. >> Oh, oh, it says sites. >> Yeah. >> Oh, there you go. Yeah. >> Ooh. >> So, it create it creates the sites. Um I don't I don't think this is a it is exactly what I what I wanted. Um but let me let me show you what it popped up, right?

29:16 Like I think as a as a research artifact, um it is very important to communicate uh exactly uh what is being done. Outputs this uh this thing which I eventually started publishing. So, I moved it off of sites because I wanted more uh database and infrastructure than sites afforded me. Uh but this is this is like research output that you can start to mess with and like try to think about like what hyper-parameters are you tuning for training your AIs.

29:40 Uh and like I was trying to make like scaling laws and everything and doing all sorts of like game optimization stuff. Uh and the fact that you can just kind of throw this up as a research artifact, like I no longer need to read ChatGPT output. I read site output. But then there's also huge sprawl. Like look at how long this thing is. There's so many numbers.

29:58 Uh it is pretty overwhelming. So, then I have to start putting in from there. But it's an interesting transition from markdown >> Yeah, actually that you're putting out to you're you're putting a whole functional site. >> I think markdown just isn't that optimal for people to read, right? Might as well just write HTML website and I don't know. I think you can do a lot with customizing this, right?

30:19 You have your skills that explain what you want. Like I noticed they're quite verbose. I don't need a lot of this information. >> verbose. >> So, and then the nice thing of having a site side by side is, you know, you just iterate on what you want and what you don't, right? >> Yeah, I don't know if that any that triggers any stories for you of how it's run internally.

30:37 Am I doing this right? >> Yeah, I mean, I think that this is like a workflow that we're seeing like all different types of teams use where like the the canonical artifact that was previously a DAC or something is now becoming a site and like with a site you because it's just HTML you can like it's infinitely flexible and so you know if you if you want to give more prominence to a certain thing that like in a slide deck would you know feel like it was braid like you can do that you can have it be like the hero image

31:06 right and so I think that like um people are starting to see that um there's obviously more work to be done to make these things like much more easier easy to collaborate on um you mentioned that they're very they're long and verbose could be broken up I'm sure there's there's something we can do about that. >> long. >> Yeah yeah but I think we're starting to see that like there there is this aspect of this is a really interesting uh format um for for people to use um that's like much more flexible than than what they

31:35 were had before. >> I think your job also comes becomes kind of meta you're not designing the products you're designing a product to make products and I'm curious how you manage that. >> I I think one thing that we've been like when we look at the UX like that we've that we've been thinking a lot about is how can we balance like simplicity with capability like if we if we're designing a product like you said that like is is made to make a build out of the things right you can build so many different things but we can't

32:07 put that all in front of you because you'll get overwhelmed. >> Yeah. >> And so we had similar problem or similar challenges even with ChatGPT but especially now like when there's so much that can be done I think the balance that we're constantly trying to strike is like how can we give the user enough of a UI surface where you know they can be expressive they can they can tell the the the agent what they need they can verify that it's using the right tools it's pulling from the right sources etc.

32:33 but then it gets out of the way and then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is going to be like how do they discover the next use case and the next one after that if they really want to you know to to be super powered by the AI. >> Yeah. >> It's interesting. I feel like everyone also just has a different way to do it, right?

32:50 I made a similar version of this, same game. I didn't take any pictures of board or rule game. I threw in at goal, 18 minutes 53 seconds later, a lot of tokens later I've got a similar version. Obviously not with all the auto research and whatnot, but you know um >> You got to do all the latest trends. >> And yeah, I I did it with uh did it with codex not work, but it's interesting, right?

33:14 >> Yeah. And this is obviously GPT image generating the avatars. Very good for game design. Like a lot of game designers were like really into GPT image. >> I I will say like the the broader takeaway probably is the reason that we do this is more so just to test the tools, right? Like um this was also a test before 5.6 came out. I had done the game on 5.5, right?

33:36 The ability for me to no longer need it to I had to feed it the rules. It's it's a pretty niche game. It couldn't find how to do this on its own. >> Oh, yeah. >> Uh 5.6 >> auto distribution. That's is why I was also very keen on testing the 5.6 capability. >> But you know, this is just as as work comes out as new things come out. These are just our sideways to test things, right?

33:56 >> Yeah. I It's some kind of private email, I guess. Which that is not all that private. But also valuable cuz now you can send this to your friends and I mean I learned about this game through seeing this. >> Hard game. He's very good. >> [laughter] >> Uh uh it's good to win when no one is competing with you. Uh but yes, it's a classic RL problem of like self play bootstrapping your game AI.

34:15 Um yeah, you see how easily work becomes personal and personal becomes work because the thing I do for personal, it actually directly informs people I work with cuz I showed it to them. They were like, "Oh, you can do that with GPT?" Which like I I imagine is the growth strategy. >> Yeah. The the show not tell is a big piece that, you know, I think we've not still not fully cracked of like, you know, showing people all the the things that they can do with the product versus like trying to teach that to them through

34:43 like, you know, articles or onboarding or whatever. >> Yeah. So, meeting them in the moment. >> It's a career risk for me uh because I used to be in developer relations, right? Where your job is to show. And then, you're like, "What do you mean it you you don't you don't need uh it actually your job is to tell." >> Mhm. >> And then and then but the product people are like, "Well, we don't need you if our product is intuitive enough."

35:02 So. >> [laughter] >> Yeah, I mean, that's the magic of the the models. So, you can tailor the telling or the showing to like specifically what the user needs, like what what they care about, what they've done in the past, exactly where they are on the adoption journey. So, I think that's like going to be a super big opportunity. >> Seems easier and easier now to tailor custom showing, right?

35:22 People have different use cases. As much as you said you don't want to segment different people into different buckets, right? It's also not that hard to for people that are in different categories. But, the question I guess is you said your team is more broadly on What What was the term you used? Productivity? >> Productivity. >> Productivity, which is now work, basically.

35:42 >> Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different than ChatGPT, Codex, or work? Is there Is there more that the mass isn't targeting? >> I see it as like a sequencing, like, you know, the the vision is like bring useful agents to everyone. We started with like developers.

36:02 Like developers historically are like early adopters. They're willing to put up with more friction, set things up, etc. Like that's where, you know, Codex started. I think the next opportunity is like sort of we call it general knowledge work, you know, all the other functions around developers. I think when when you go from developers to this segment, like there's inherent challenges, obviously, with like, you know, this this show not tell thing that we're talking about, um making the product more understandable, um

36:32 bringing in new capabilities that matter more for for this cohort, the matter for developers, things like artifacts, things like computer use, etc. And then I think like the the same learning is like similarly how we took the learnings from developers and brought it to, you know, general knowledge work, the next stage will be like taking the learnings from general knowledge work and bringing it to everyone, no matter what they're doing in their lives.

36:51 Um and we're already seeing that a little bit, like this this game example that you have is, you know, something that's it's like on the border of like fun and personal life to to, you know, your your professional life. I use ChatGPT at work full-time at home for everything, like for for whatever I'm doing. I used it the other day to come up with a meal plan and like, you know, save that on um on the like computer environment that it has and something that I can continue going back to.

37:13 Like, is everyone doing that yet? Probably not because of things that we work on it, but eventually, you know, we want to get people there. >> ChatGPT life. >> Yeah, exactly. ChatGPT cooking. Um but I think there's a lot of uh there's a lot of opportunity there, but I see it as like, you know, we're we're built we built the foundation in in in in software engineering and we're going to take this same learnings that we take from software engineering to knowledge work to knowledge work to to everyone.

37:35 >> Do you have any power user advice? I feel like um there's a group of people that will live it, use it for everything, stay on it 24/7. Yeah. Um and then there's a bit of a gap between that crew and people that, you know, okay, I use it for work, I use it occasionally, sometimes I pipe questions. Uh any advice, any learnings, anything you recommend or just, you know, takeaways that you found that help bridge that gap?

37:59 >> I think a couple things that I've seen is like one that it really helps to broaden your imagination of what's possible. And this has been a learning even for me, like, you know, the technology has progressed so fast that, you know, there's something that like even 3 months ago like, no way, no way the models can do this. Like, now it's like, wow, it's like it actually can.

38:19 Like, um >> Give you have >> We're going through right now that are like review cycle internally and you know people always talked about this as like kind of a a thing that the the models are good at like you know there's a cliche of like okay like I don't want to be writing reviews and like we just use AI to do it but I mean and then it needs to be evaluated as well.

38:37 Yeah exactly. In all seriousness before it was like just like slop basically and like I I think it was helpful but you know not super productive. Now I found that like the model can do a much much better job than me especially in this environment of like pulling contacts on like what people are up to how they've like the things that they've done to make a difference highlighting like you know wins that they've had that like I might may not even have seen you know it has access to like everything right like the code

39:03 like you know things that they've caught reviews slack everything and so it's like incredibly powerful in that domain and like just like six months ago the last time we did this cycle like I didn't even I tried using it but it was not at all helpful and this time it's been like incredibly helpful and like so I think continuing to push the the frontier of imagination what's possible even if you tried something before I think is maybe that my biggest piece of advice.

39:23 The other I guess thing is like the more the more you put in especially in this environment where like you know the model has access to to everything on your on your computer or in chat GPT work like you can create you know artifacts over time and save them in your library and like the model continue having access to those like the more information you give it about whatever domain you're in whether it's your life or your work the more valuable it becomes and it'll become more valuable in like ways that might surprise

39:46 you like it might pull from contacts in a way that you know maybe proactive in that you might not even have thought about but it needs to have access to those so that those tools or that context first. One thing I was just wanted to talk about the review stuff because I still that's a very sensitive thing and you're you're founders manage people you've hired people as manager myself I'm very reticent to put out any LLM generated things especially when it comes to people cuz it feels like you don't care.

40:15 Presumably at Open AI people are obviously more open to being basically rated by GBT. Uh but are there any unofficial rules around this? Like what's the etiquette? >> Oh, I mean I think the etiquette is that like I would never write something via like only via AI and like present it as like a review for someone. What I was talking about is more like gathering context.

40:36 >> Yeah. That's the place where it's >> So it's just search. It's agentic search. >> It's like agentic search, but you know, that you can tailor and steer much more capably than you could before. And cuz like the the thing is it's all there's sort of a flywheel happening, right? Because of Codex people are able to do them because of ChatGPT more people are able to do so much more now than ever before.

40:54 And if you're able to do so much more, it's easy to miss things as well. Um and so like I think we need to use these same tools to keep up with all the impact that people are having. And and and understand, you know, where where we can be helpful. >> I think that the the thing like obviously I I run a small company, so easy to search. But at the scale of OpenAI, with the mono messages you guys put in Slack, do you think that it misses things?

41:19 >> Probably, but I think that I also miss things. >> Like it doesn't matter, right? Like it's it's as it needs to be human level. >> relative, right? Yeah. Sometimes it's nice when it finds things you wouldn't, right? Like right now my Codex system prompts they're set up in such a way that every project I have has a secret separate notes MD. And it just writes learnings to there.

41:36 And then the the global one can pull from all these. So sometimes it'll be like, "Oh, there's this project you did like 4 months ago. Here's a note that we had." And it it randomly pulls it back in the context that I would never do I haven't thought about. And I'm like, "Okay, this is quite superhuman, right? Like stuff that would And you know, it'll save like hours on chunking of stuff or find something that's already been done.

41:56 I'm like, as much as it might miss stuff I would do, but it's very useful when it finds stuff. And I have like a very, you know, non-super engineered solution to this. It's just markdown files that get pulled whenever they want. >> Yeah, I actually have a funny anecdote about this. Like recently gearing up to this launch, you know, the team has been you know, really cooking on it for for for a couple months and over that time like there's so much conversation and chatter going on in Slack and Docs and elsewhere.

42:21 And uh one one of the members of the team set up this um scheduled task like automation to like look at everything that's going on and like come up with the best memes and then post in one of our shared channels. And like there are two cool things about this. Like the first is like I think the models are you know, over time like actually starting to become like funny.

42:42 Whereas like you know, a year ago like that was not at all the case. The second is that it was what you were saying like they find things that in surprising ways that you may not have not have thought of and like create connections that you may not have thought of and that really helps with like the meme generation because then you can see something that you know, genuinely surprises you and then and um is funny in that way.

42:57 So yeah, I mean obviously that's like not like the most productive uh use of this of this technology, but it does it doesn't cover this like this capability that's emerging which is just like defining information that you otherwise would not know. Talking about the the launch, I I think uh I have pretty much said this is the most successful launch in a long time.

43:17 I think even more successful personally than 5.0. And you're announcing 10 million users. Does it feel different? You've been through a lot of launches. I think it feels like a culmination. Well, I think two things. One it feels like a culmination like I was mentioning earlier like this like vision vision that we've been on for a long time. Like I said, we saw the magic of Codex internally and then we're like extremely excited to bring this to many more people and to see it working to like see us reach you know, the

43:44 distribution goal in um numbers that you mentioned like I think that's like huge and and and super exciting. The flip side of that is like there's so much more to do too. Like that's also really exciting. Like you know, ChatGPT as a whole like this product that you know, everyone almost equates to AI and like loves you know, has hundreds of millions of users.

44:03 And so like 10 million is really cool, but like we we need to get this to everyone. Like we need everyone to feel this magic. Um and so that's the next step from here, but yeah, I think extremely pumped about how it how it's going so far and then the opportunities. >> Uh awesome. Uh I did want to also, because I have I've I've been tracking the the number closely, it transitioned at some point from just Codex users to Codex plus ChatGPT work.

44:23 Obviously, because the same harness, the whole point is that you don't you you can't uh comment separately. Do you have roughly a billion um ChatGPT users? When did it just jump to 1 billion right away? Like, isn't that the default on ChatGPT or no? >> We don't default you into ChatGPT work if you're on ChatGPT or free. Yeah. It's also only available to it to paid users right now and I think there's like a process if you know educating users of what is the value of this product, having them try it, learning from their

44:54 feedback, and making it better over time. But, I mean, the goal is to, you know, get as many of people who who love ChatGPT today to like feel the feel the power of ChatGPT work, but I think it'll be a a journey. >> Yeah. And Codex will will still be alive as a brand for the foreseeable future. >> Yeah. >> Um and we'll just toggle between them as as as needed for UI stuff.

45:12 >> Yeah, I think it's even even stronger point than that. Like, I think we fully intend to like, you know, treat developer like developers have been, you know, core market for us for so long and like there's there's so much more that we can do to make Codex great specifically for um software development and we'll continue to do that. This doesn't take away from that at all.

45:28 If anything, it should increase the utility of something like Codex because now like you can move seamlessly between writing a def to creating an artifact or, you know, doing a search over your code. >> I do wonder how much this terminology leaks to the non-technical user. Like, do they have to learn to say artifacts if I want artifacts or, you know, >> It's funny like we call artifacts internally cuz that's what the teams call them, but like externally like no one says that.

45:53 No one calls it an artifact. I But, I think that people like often like describe things whatever they're used to, right? So, if, you know, ChatGPT work is good at creating slides, they'll say ChatGPT work is good at creating slides, and that's actually what we want. >> One big another, I mean, it's July of 2026. One big thing that also happens in for opening AI was open claw, and that's I think a lot of people's first time really maxing agent for personal stuff, but also crossing over to work in in that sense same way.

46:21 As far as as far as I understand, open claw is still independent, but did you go through your own open claw moments? Were there any lessons you took from open claw to Codex or back, whatever? >> I think there's a lot of inspiration. I did go through my own open claw moment. I um >> Yeah, tell the story. >> Me and my um my my wife like set up an open claw to like try to manage everything in our house.

46:46 Not that there's like a ton, but it was like actually quite useful. It gave it a calendar, started, you know, creating events for us and stuff. At some point the the laptop they were running on it died, and then they never got a chance to to pick it back up, but there's a lot of inspiration there. Like, um you know, in ChatGPT work in web and mobile, like you you get access to this like persistent computer environment where, you know, you can store files, and those files stay around between sessions.

47:09 And the idea is to be able to enable use cases like this. Um one of the members of our team actually uses ChatGPT work for what they used open claw for them before, and I feel like it has like completely transitioned, which is like uh work out planning and like meal tracking, um which again, it's like a worky thing, right? It's like not work necessarily, but it's like in this personal productivity space.

47:28 But it has all the same primitives. Like it has scheduled tasks. It has the ability to store files on a file system. It has the ability to like reference those things over time. Um and so you start to see the same types of use cases emerge, which has been really cool. >> Is there a point that ChatGPT work completely replaces open claw? Obviously, they're independent, so >> Yeah, I mean, I'm I'm I'm not close to it, so I can't speak to the the open claw road map, but I I don't think so.

47:52 I think that there's going to be, you know, there's always a need for like this like incredible like open source technology that that team has built and I think that we can draw inspiration in the product and you know, chat GPT I think many more people have like heard about and used chat GPT than have you have used open claw and if we can take the magic from open claw and bring it to them, I think that'll be a success.

48:14 I think that like one thing on the chat GPT work side that we feel strongly about is that like the core experience is that you come to this product and you have a conversation, start a session, whatever you want to call it with this agent and the magic of the product is that you can do anything in that moment and we would like to create a product where you don't have to click a button or to go to a different place, whatever and you can get whatever functionality exists in you know, your your finances app or or any

48:41 other product like in this in this one place and so that's the goal. It's like it'd be we want an extensible system with plugins where you can connect to the tools that you need in order to be able to accomplish like a financial task where you can you know, if if you're doing like science work like we have an ability to like extend the system in such that you can like write the tech and and it performs well.

49:02 There'll always be like products that we support that are best in class at those things, but we want as much of the magic as possible in that core experience, you know. >> Do you think that you can do everything you used to do with Wolfram's in chat GPT finance? >> I actually tried it. I mean, like chat GPT doesn't yet custody cash and and assets for you, so so that part no, not yet, but I I mean, there's like a whole component of like retirement planning and um sort of like financial planning and budgeting and stuff

49:35 that that um we were looking into when I was there and like with the finances plugin like that's all possible with chat [laughter] GPT today, so um I feel like at least that component's replaced for me. >> I haven't really plugged it in yet. Um I'm somewhat scared to look at the answer. >> [laughter] >> Like that's honestly like the same reason for health and finances.

49:52 Like, I'm like, I don't know. I don't know. >> [laughter] >> It's really good. I mean, it's it's really cool how I mean, we were talking about like the agentic search aspect a little bit earlier, but like it's really cool how like, you know, in in conventional UX, like if the more power you want to give to a user, the more like knobs and bells and whistles you need to add.

50:08 Like, you know, for like these finance and budgeting apps, like there's always like a bunch of the different filters and like search bars and stuff like that. But like now like with the right connect connectivity to the right data, you can have whatever you want. You can ask any question you want and and into that box and and get the answer and I think that's super powerful.

50:26 >> I think it's also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, all these different things. It's just nice to centrally collocate it. >> Which is, you know, part of the whole thing of Open Claw, right? Like that that you would have a personal OS, which presumably ChatGPT wants to become.

50:44 I do think that just relying on like just-in-time pulling of data for, let's say, through via MCP, CLI, API, whatever you whatever you do, still not enough. Like I I like I come from a bit bit of a data engineering background. Like you still want like a data warehouse or some kind of caching or semantic layer. Um, do you [snorts] feel that or do you already have that?

51:09 >> I can't speak to like all the details on how everything works, but I think it depends on the access pattern, right? Like, if you want to answer immediately, then yes, it's very difficult to do that if you need to pull from all of these sources. But, a lot of the like use cases that we want to enable on ChatGPT work aren't necessarily something that you need immediately.

51:26 It's more like a task that you want the agent to go and do. And that that's going to take a certain amount of time. And, you know, with things like programmatic tool calling and stuff now, like some of that time and sub agents and stuff like some of that is also parallelizable. And so, it's possible I I I think it's very possible that there's a the ceiling on what can be done, you know, with MCPs and like calling out to these third party services has been raised substantially.

51:49 So, we're really excited about that. >> You mentioned sub agents. I I I got to double click on that. Ultra is a new mode. Um you have special affordances in ChatGPT itself to show off the the agents. Can't really do much with them to be honest. I just just just watch. >> [laughter] >> Um what have been your what have been your experiences? Any design issues that you would call out to other builders building with sub agents?

52:14 >> I think it's sort of goes back to the balance that I was raising earlier. Felt like, you know, showing builders the power of the tool, but also creating enough of an abstraction to not overwhelm them. I think with sub agents the thing that we wanted to show is that you can take a task that, you know, has many parallel tracks or um is is complicated in a way that, you know, sub agents can handle and this product is for you.

52:40 Like the model can can accomplish those goals or try to accomplish those goals. Um and so like that's the point of like showing them in the product and and that's where we we've gone with the design. There's another, you know, iteration of this where like you can see exactly what they're doing and and things like that, which I think is like, you know, could could verge on like overwhelming um with information.

52:59 And so this is like the double edged trade off that we made for now. >> you you you do display quite a lot of transcripts. >> Right. Right. >> I think it's you want to display more than that? >> No, no, it's fine. >> Some some people could want more. So, I'm one of those people that will basically throw a lot of stuff at goal and pretty much every goal I'll tell it to use sub agents.

53:16 Seems redundant, right? But every time I'm like, "Okay, use sub agents where possible." And I have a lot of people a lot of friends that recommend and do the same. Whereas I'll sometimes talk to people that are like, "Okay, this is where I want you to use sub agents for this sub task." And I'm sure they would appreciate seeing into how they're being used.

53:34 For me it's primarily like two things, right? One is net time efficiency. So, span out across sub agents. Two is probably cost, right? Uh don't use big expensive model, offload to a lot of smaller, cheaper models, and some people want that level of control. So, if you have repetition in what you're doing, right? Say I want something built where I want it to consistently do this every day, I might want to go in and fine-tune subagents here, subagents there, so you can see both, but I think if I'm not mistaken, it's

54:04 hidden by default. There's a drop-down that goes a lot where I'm like, "Okay, I'm just going to keep it on." >> You can You can change the model that they use? >> I know I tell them to be steered. I'll say my I know Anthropic offers this in Claude code. You can tell Fable to use Sonnet or Opus to use Sonnet as a subagent, so pretty trivial thing, you know, you tell it to span out subagents with Sonnet, you know, it's cheaper, faster.

54:25 I would assume if it's not there, it could be built there, but I think there's a side of >> Too many toggles. >> Mhm. It's not a toggle, actually. It's just a You tell it in chat. The way I do it is prompt it, right? And I think this is something that gets abstracted unless it's something you built for repetition, right? So, if I'm building something, say that's a podcast prep, right?

54:47 Research into people, do a very, very deep, extensive research, um that I might want to configure to cheaper, faster model just for web search, right? I can see a world in which you want both. I think the default is actually pretty good right now, where it's hidden, but you can drop down and get some more info into what's done. I know people talked a lot about it on 5.6's launch.

55:06 Uh this thing loves to use a lot of subagents and causes the ChatGPT app to just crash because it's so processor heavy, but um >> So, what is your personal experience? Yeah, I mean, yeah, you know, I haven't had it crash from subagents. >> I I haven't either. I have We both have big laptops, but I know I know people brought it up. There was a topic of discussion that we didn't see the same, but it is another vibe eval, right?

55:32 People are like, "Okay, the amount of subagents all this running is crazy." And I'm like, "I think this is okay. I think it's good, but just stuff people bring up. >> I think when we launched the product too, we we weren't as opinionated about like who is Ultra 4 and like when when should they be using it and since then we made some changes to like, you know, require you to turn it on and and and find it in the advanced setting cuz that's who it is for.

55:54 It's for like power users who understand what's going to happen because it also, you know, depending on your use case can can use more of your your limits as well. >> Yes. >> So that's where I think a lot of the the feedback was coming for us. >> Okay, reset the limits. >> [laughter] >> Always reset the limits. >> Well, it's you know, today we're resetting because it is I want to change topics to one last piece of the harness, memory.

56:12 A lot of people are commenting on memory recently. ChatGPT's new memory system used to suck is not very good. And then this guy also basically the same thing. And Samir who you presumably work with talking about memory. What can you say there? >> I think that, you know, Samir and the team have made a ton of and and and the research teams have made a ton of updates and and improvements over time.

56:33 I think when I talk to friends, family members about what they love about ChatGPT, like the fact that it knows them, that they feel like their ChatGPT is is their ChatGPT, I think comes up probably >> Yeah. >> number one. And ChatGPT work in the in the cloud like by default all conversations like inherent from ChatGPT memory so you'll know they'll know context about you.

56:52 They'll also be able to write back to this memory. >> With the like a like a like a small text right. Like you tell me when you're writing, right? Is it >> No, it's part of the same like memory V3 system that that we we launched. >> What do you mean V3? Yeah. >> So I think that's been really powerful because, you know, going from ChatGPT to ChatGPT work feels like an extension of what I've already been doing with the product for sometimes many years.

57:15 Um so that's been awesome and it's awesome to see that like people are recognizing um the the improvements here. >> Is there So it's it's basically a retrieval problem, right? Like are you retrieving the right things? Are you over focusing on the wrong things? Is there like uh more false positive or false negative, you know, if if that makes sense? Like what's the bigger problem?

57:34 >> So, I don't work on memory directly, so it's hard to say what the bigger problem is with like certainty, but I I think you're right. I think that like, you know, the there's two sides of it. It's like, you know, making sure it knows things about you, but then also having the EQ to like bring those things up at the right moments, proactively or surprising you in ways that are positive, not negative.

57:50 Um so, I think it's a very challenging problem, but something that I I think we feel very is a huge opportunity get right, which is like why we made like big investments in it. >> How do you see the side of, okay, when you're building ChatGPT for work different than the regular chat app, different than Codex, managing memory across different projects, um collaboration and whatnot?

58:08 How do you see the the side of what's separate from the harness, right? So, if I have four threads on one project, um any learnings on how to build memory systems there, you know? >> For background as well, I guess, to steer it a bit is when you do chat style applications, I'd say you have a lot of one-offs, right? When you switch to work, it might be something you're doing for a month, something you do a lot, right?

58:32 Now, as I add more sessions, there's a lot more than just single-threaded, right? And there there might be memory there. >> I mean, I think first I'd challenge that like the depth of the memory or the like value of it is like fundamentally different across chat and work. Like, it it is true that like, you know, there are a lot of like shorter sessions on chat, but I think, you know, the ChatGPT the product has had like a ton of longevity, um and you know, as long as this this technology has been around, and and people

58:58 use it for worky like productivity-related things already today. And so, I think we've found that there's a lot of value. I mean, I found this even in my personal usage, like all of these one-offs add up over time into something like quite durable and like like quite a good representation of who I am. I know like from time to time something will go viral on on X about like, you know, ChatGPT telling you everything it knows about you, and people are always surprised like how how deep that is.

59:24 >> The fun roast me, you know. >> Exactly. So, like I think like the That's all to say that like I think there's a lot of depth there in the existing, you know, ChatGPT product, and so that's why I think we think it's valuable to bring into the work product. But, the other reason I brought that up is because I think, like, hopefully we can use some of the same fundamental primitives and systems to extend memory here as well.

59:43 And I know this is something that the the team that focuses on this is like working working through right now. >> I wanted to bring up one element of memory, which I honestly don't really use much, and I'm curious if you do, Chronicle, which was is up on screen right now. Uh it's kind of a super memory, or like what what is it? >> [laughter] >> I mean, the idea is that, like, it can learn from, you know, how you're using your computer, and like it's an another input source um into memory.

01:00:08 And um I think it's, you know, experimental right now, and something that, like, isn't default off, but I'd recommend that you try. I think that it's like quite interesting how it goes back to a conversation we were having earlier on, like, you know, you were asking, like, does it Can ChatGPT miss things? Like, does it, you know, on Slack, when it's searching, does it miss things?

01:00:29 Cuz there's such a volume of stuff, right? And like it's I You can ask the same question about, like, everything that you're doing on your computer. Like, is it going to know everything that you're doing? Is going to capture the intent and stuff like that? Probably not, but like it probably will find things that you might not know about. And then, if it can surface those to you in relevant times in proactive ways, like when you're doing tasks, then I found, at least, that it can be quite helpful.

01:00:51 So, it's worth trying. >> So, mostly for insights and longer term? >> Yeah, exactly. Like, insights and it builds context that that makes that can make you more productive on certain tasks, but it's it's it's hard to describe without feeling it. Um >> [laughter] >> I will say you can feel it pretty well. Like, the idea of what they're saying here, right?

01:01:10 Just check through my memories, or check through my logs, and add skills. >> Yeah. >> Pretty underrated, right? >> But, that's that's automations. You you can repeat that using a cron job. >> Checking through your memories and creating skills. >> Yeah. >> But I think the creation of the memories from Chronicle itself is like what's different. It's like you have much deeper memories because you have Chronicle on.

01:01:29 >> It's there. I don't use it much, but maybe I just I need more examples. I I imagine you guys use a lot of it internally, so I'm always fishing for use cases. >> Yeah, I would just try turning it on and then like it just auto works. Like it >> Yeah, and seeing like where where it might start helping you. I think you'll be surprised. >> Yeah, amazing.

01:01:48 >> I think that was the about it in terms of like the the overall coverage of ChatGPT work. I think there's been a lot of like good progress and discussion on building and all these things. There's a lot of like ex-founders in the in the community and in OpenAI as well. Do you think that things have changed a lot? I I guess like your overall reflection of building pre-AI and post-AI.

01:02:13 >> I mean, I think things have changed a ton. I think it's it's like super exciting to see how quickly you can go to from idea to something real today. Yeah. Um whereas like even before like I think you know, 5 10 years ago like it it was fast if you were scrappy and you know, like willing to build the the minimal viable thing, but like now the extent of what you can build is like much much broader.

01:02:35 And I think that also like what we've seen internally building is like that gives you an opportunity to validate much more quickly, to talk to users, to talk to to internal doctors, etc. and like make sure you're on the right track. And like that loop I think has been has become more closed than ever before and that's like a win for product development.

01:02:54 I think it's a win for for consumers and users too cuz ideally that means they're getting much more better much better products out the gate. >> Does it mean your team's a smaller? >> I think there's much more to do now. Um so I think people can accomplish more individually or in a small team than they were that would require more people within before, but there's at the same time there's also more to do.

01:03:13 So I think the teams are much more ambitious. >> You seen any changes in scopes of roles and building teams and how we used to have teams say a few years ago versus what ideal teams look like now? >> I think we've seen a blurring in the lines between like the typical product development functions like between like you know EM, PM, engineer, um designer, etc.

01:03:34 like >> Yeah, when I bring up this quote, there will be uh only four jobs left in tech. There's AI uh slop cannon, the people who just like uh they'll burn a bunch of tokens uh and then there is there is SRE uh who people who are more responsible. There's grown-ups who sell things and then there's hot people. >> [laughter] >> It's an interesting take.

01:03:55 I think my my suspicion is that there's everything everyone will be like T-shaped in a way and not like AI will enable everyone to become a generalist like you know things that like I I never would be able to like come up with a design before and like even now I don't have maybe like the visual taste required but I can iterate on something with the help of AI.

01:04:18 But then people will have a specialty and that's like the the like it's a straight line in the T or the the upward line in the T. Um and so like you have a specialty that you're interested in with the help of AI you can go deeper and become better at over time but then you'll also be a generalist and so with that foundation the what you can accomplish is like almost limitless.

01:04:35 >> What are you bottlenecked by in terms of specialties? Like do you need more designers? Do you need more slop cannons? Do you need more hot people? >> I think the bottleneck some becomes like sort of like ideas and taste I guess. Um I think because anyone can can build now, I think um it really is the era of like bottoms-up ambition and because there's so much to be built, like you're always going to be bottlenecked by you know the amount of ideas and amount of things that you're doing at any given time.

01:05:06 >> Do you think models help solve that? >> Models? >> Yeah. I mean I have the example of like I have a front end design skill that's like they give me four drastically different examples of what this looks like. You sure it burns a lot of tokens, but you know and then I'll mostly just condense down. Okay, I like this part, I like this part. Let's draw these together and it's like yeah, I had a vision, but like I don't know.

01:05:29 I would say that the one automation that I would love to work and it doesn't work is bring me new ideas, right? Uh somehow LLMs are just not it. One interesting part about ideas is like they're not like in a vacuum. It's like not they they usually come from somewhere and like you know in in product development like they're coming from talking to users or reacting to you know friction that you're seeing or feedback building on some foundation that you already had planned out before whatever.

01:05:54 Um and so I think that's where like I think there there always be value in in these like generalists that we talked about like you know closing that loop and and and having coming up with those ideas that are grounded in in that feedback or talking to users or whatever it is. >> Cool. Uh you were going to you lead the productivity team. How do you define productivity?

01:06:15 >> I think our mission is to make it possible for people to do things that they weren't able to do before. And right now we're thinking about it from the the perspective of knowledge work. And so when looking at knowledge work I think about people are no longer siloed by their roles, they're no longer siloed by maybe the the um background or training that they have.

01:06:32 Like no matter what function you're in you can suddenly build things and suddenly get access to data that you otherwise might not be able to interpret etc. Um and then I think that extends to your personal life where we want to give you leverage at the end of the day. Like we want the models in the product to be able to give you leverage so that you can you know create time for yourself to do the things that you love.

01:06:54 >> Does that also translate to a way to measure productivity? Like what is how do you measure leverage? >> I think we haven't figured this out yet. Part of the reason is it's so diverse. Everyone has different goals and really the true measurement is like their ability to achieve that goal. Did we help you or did we not? >> Yeah. >> And it's very difficult without knowing what that goal is up front and also tailoring it for every individual.

01:07:17 >> And the thumbs up and thumbs down from ChatGPT doesn't give you anything, right? >> Right. I mean, you don't know if they're thumbs downing the the content of the answer, the vibe of it, whether or not it helped them with their goal. I think that's difficult. Um but it's something that I think we will need to figure out and the industry at large will need to figure out because, you know, that's how we we measure success if this is what we're >> for.

01:07:35 >> Do you think it's changed productivity and how you measure it? Basically, you said there's a lot more work that can be done, a lot more scope. Um has it changed? >> I think it was always true that what you really wanted to measure is like you know, was your team, was the individual, were you personally able to hit the goal or are you closer to hitting that whatever your goal is, right?

01:07:53 But I think previously we used proxies for this. So like, you know, code commits or lines of code or whatever. Um >> Story points. >> Yeah, exactly. Story points. And like >> Coming back, by the way. >> [laughter] >> May maybe. I mean, but but that is for part of the change and like I think with with AI now, those proxies starting to fall apart. Like you you know, the number of tokens you use or the number of pull requests you make or like no longer like maybe as hyper correlated with the that is your team able to hit

01:08:24 the goal or are they on track to hit their goals. So I think we'll need to come up with with new um measurements. >> For the managers listening, uh give them one thing to try. >> I think for me, what's important is like at-bats. Are we as a team building the muscle to have not just quantity of at-bats, but quality? Like, are we able to go all the way from like generating an idea, building it out, getting the feedback, reacting to that feedback, actually validating or invalidating the hypothesis, going on to the next

01:08:55 idea? Are we able to do that really efficiently? Like that goes to like, you know, the actual like code that's being written or the designs that are being made or the specs that are being written whatever, but also the culture of the team. Like do we have the humility and and and and um are able to like go through that process many many times and stay motivated and excited throughout that.

01:09:12 Um so that's the thing that like I think is important now, especially when we're on the frontier of this technology and like there's so much to build, there's so much to do. That's probably the most most important thing that we look at. >> Any traps people fall into around measuring productivity with your team work on? I feel like there's a lot of okay, we added a lot of LLMs, we have dashboards for this and that, but not much has changed, right?

01:09:35 >> [laughter] >> That is the trap, yes. >> And you know, the the broader source of the question is for for the managers and teams building, you know, how how should they approach this? >> I think maybe the trap is like conflating motion and progress. I think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be like very prescriptive and deliberate about like what you're actually trying to achieve.

01:10:02 And it goes back to our question of measurement, right? Like you were we were talking about like can we OpenAI like figure out how to measure productivity for our users? That's that's a very hard problem because of the diversity, but like as a team like you should have a really prescriptive and deliberate view on like what progress looks like for you and for your team.

01:10:21 And if you don't have that, then it's very easy to conflate these two things. I think that batch is a really great thing. I'm I'm really glad I I like the >> discussion between motion and progress. I think that's a quote that we're going to feature in the write-up. You've been very generous with your time. Thank you so much and congrats on 10 million. >> Yeah, thank you for having me. >> The next to a next one at 100. In two months. >> [laughter] >> Thank you. >> [music] [music]

💡 Answer

ChatGPT is being developed as an extensible personal and work productivity environment where agents can perform tasks across connected tools, files, data, and applications without forcing users to choose separate experiences.

🧠 AI Summary

ChatGPT Work is designed to bring Codex-style agents, computer use, artifacts, sites, plugins, persistent files, and memory into a unified productivity experience for personal and professional tasks. The rollout is intended to move from developers to general knowledge work and eventually everyone. AI reduces the time from idea to working product and blurs traditional product roles, but progress still depends on ideas, taste, user feedback, validation, and clear goals. Productivity should be measured by progress toward goals and the quality of iterative at-bats, not by activity proxies such as tokens, commits, or pull requests.

🔑 Key Points

  • ChatGPT Work merges the power of Codex agents with ChatGPT while preserving a shared underlying harness.
  • The product targets personal productivity, professional knowledge work, and eventually broader everyday use.
  • Agents become more capable when they have the right context, connected tools, and user supervision.
  • Artifacts and sites provide flexible, interactive alternatives to conventional documents, spreadsheets, slide decks, and markdown outputs.
  • Persistent files, memory, scheduled tasks, plugins, and computer environments support longer-running workflows.
  • AI enables individuals and small teams to build and validate ideas much faster, increasing the importance of taste, ideas, and user feedback.
  • Productivity should be evaluated through progress toward explicit goals and the quality of repeated idea-building-feedback-validation cycles.
  • Managers should avoid conflating increased activity with meaningful progress.

✅ Actionable items

  • Start with the default model configuration and change reasoning levels or model classes only when quality or efficiency is insufficient.
  • Use Ultra or multi-agent setups for highly complicated, open-ended, or highly parallelizable tasks.
  • Use goal-based configurations for tasks where progress can remain consistent and verifiable over time.
  • Connect relevant tools and provide the agent with domain context, files, and prior artifacts to increase its usefulness.
  • Use sites to turn research outputs into interactive artifacts and iterate on their structure and presentation.
  • Test new AI capabilities through concrete side projects, such as building a game or research workflow.
  • Use agentic search to gather context for sensitive work such as performance reviews, rather than presenting entirely AI-written reviews as personal judgments.
  • Create separate project notes and allow a global system to retrieve learnings from them across projects.
  • Turn on Chronicle and evaluate whether computer-use history produces useful memories and insights.
  • Define explicit team goals and assess whether the team can repeatedly generate ideas, build, gather feedback, and validate or invalidate hypotheses.

📣 Marketing

Sales

  • Meet enterprise users in their specific use cases and teach them how AI can provide leverage.
  • Avoid assuming that enterprise customers have a single common AI use case.

Branding

  • Keep Codex as a durable brand for software development while integrating its capabilities into ChatGPT Work.
  • Describe capabilities in terms familiar to users, such as creating slides, rather than requiring internal terminology such as artifacts.

Distribution

  • Use ChatGPT's existing distribution base to bring Codex-style agents to more users.
  • Expand sequentially from developers to general knowledge workers and then to everyone.

Customer acquisition

  • Use a show-not-tell approach by demonstrating useful capabilities in concrete personal or professional contexts.
  • Let users discover capabilities through the product rather than relying only on articles or onboarding.

🔍 SEO & discoverability

Other channels

  • Use concrete demonstrations to help users discover capabilities they may not realize are possible.

🧭 Frameworks

Idea-to-validation at-bats01:08:47
  1. Generate an idea.
  2. Build it out.
  3. Get feedback.
  4. React to the feedback.
  5. Validate or invalidate the hypothesis.
  6. Move on to the next idea.
Motion versus progress01:09:35
  1. Define explicitly what the team is trying to achieve.
  2. Measure progress toward that goal.
  3. Avoid treating increased activity as evidence of progress.

🧰 Tools & AI usage

  • ChatGPT Work — Unified agentic environment for work and personal productivity.01:40
  • Codex — Agentic software-development product whose capabilities are being extended to broader knowledge work.07:13
  • Airtable — Database and low-code product referenced in Akshay's career history.01:52
  • Artifacts — Interactive outputs for spreadsheets, sites, and other knowledge-work tasks.22:29
  • Sites — Flexible HTML-based artifacts for prototyping, collaboration, reports, and research outputs.26:00
  • Chronicle — Experimental memory input that learns from computer usage.01:00:00
  • Open Claw — Independent open-source technology used for household management, calendars, work planning, and meal tracking.46:34
  • Excel — Spreadsheet format that ChatGPT Work can edit through an agentic experience.22:01
  • Slack — Source of workplace context, conversations, and scheduled meme-generation inputs.42:13

AI is used for

  • Automated software testing — Use AI to make code-level capabilities accessible to more people.02:23
  • Creating a retirement calculator spreadsheet — Demonstrate artifact creation and differences between Work and Codex experiences.11:51
  • Finding a missed package — Use an image, neighborhood listings, and an agent to identify the apartment complex where a package was left.09:43
  • Generating performance-review context — Search code, reviews, Slack, and other sources to identify a person's work and contributions.38:37
  • Generating memes from workplace information — Use a scheduled automation to review Slack, Docs, and other conversations and create memes.42:26
  • Building and optimizing a board-game AI — Use self-play, loss functions, benchmarks, hyperparameter tuning, and auto-research to improve game-playing agents.27:47

📊 Numbers mentioned

Costs

  • Subagents can be used to improve time efficiency and potentially reduce costs by offloading work to smaller models.
  • Ultra can use more user limits depending on the use case.

Growth

  • Akshay joined OpenAI in 2023 when it had about 500 people.
  • The opportunity beyond current AI adoption was described as a 10x or 100x larger market.
  • A board-game demonstration used 1.7 billion tokens.
  • Another board-game build took 18 minutes 53 seconds.

Traffic

  • ChatGPT has hundreds of millions of users.
  • ChatGPT Work reached 10 million users.
  • ChatGPT Work is available only to paid users at the time discussed.

⚖️ Advantages, risks & lessons

Advantages

  • AI shortens the time from an idea to something real.
  • Agents can work across connected tools, files, and data sources.
  • Sites are more flexible than conventional slide decks and spreadsheets because they are HTML-based.
  • Persistent context and memory can surface information users would not otherwise remember.
  • A shared harness lets users move between coding and broader knowledge-work tasks without choosing separate experiences.

Risks

  • Users can become overwhelmed by the number of capabilities, models, settings, and outputs.
  • Agents may miss information or retrieve context at the wrong time.
  • Connecting deeply personal files, tools, and data creates security and permissions challenges.
  • AI-generated performance reviews can appear insensitive if presented without human judgment.
  • Increased motion and activity can be mistaken for genuine progress.
  • Ultra and subagent workflows can consume more usage limits and may be computationally demanding.

Lessons

  • The default configuration should be good enough for most users.
  • Users benefit from broadening their imagination about what current models can do.
  • Generalists remain valuable because ideas arise from users, friction, feedback, and existing foundations.
  • AI expands people's capabilities while preserving the value of human specialty, taste, and judgment.
  • The most useful productivity measure is whether people achieve their goals.

💬 Quotes

The bottleneck some becomes like sort of like ideas and taste.

Summarizes the claim that building capacity is shifting the constraint toward judgment and invention.01:04:45

I think maybe the trap is like conflating motion and progress.

Captures the central warning about measuring AI-enabled productivity.01:09:35

Our mission is to make it possible for people to do things that they weren't able to do before.

Defines the productivity team's stated mission.01:06:17

👤 People & companies

Akshay Nathan

OpenAI leader of core product engineering and the productivity team, discussing ChatGPT Work, agents, product development, and productivity.

01:32
Vibhu

Co-host participating in the discussion and product demonstrations.

01:40
Gabriel Chua

Member of OpenAI's safety team who showed Akshay an agentic Excel capability.

22:10
Samir

OpenAI colleague mentioned in connection with ChatGPT memory.

56:22
OpenAI

AI company developing ChatGPT, Codex, ChatGPT Work, models, agents, artifacts, sites, and memory systems.

01:32
Airtable

Company where Akshay worked on bringing database parameters to more people.

01:52
Amazon

Company referenced as a possible courier or source of a missed-package image.

09:43
Wolfram

Provider referenced in a comparison involving finance capabilities.

49:16
Anthropic

AI company referenced for allowing users to specify models for subagents in Claude Code.

54:13