Yes, databases have a future, especially relational databases for structured enterprise data and durable workflows, although AI may substantially change database systems, software development, and computer science itself.
Run long-running, computationally expensive agentic-AI workflows with important state stored in a database, allowing them to resume from completed steps and logically undo committed updates when later steps fail.
Amazon Neptune is Amazon's graph database service. The service represents graphs internally using node and edge tables. In the discussed workloads, the video argues that relational databases have so far outperformed Neptune.
Claude is an AI assistant developed by Anthropic, positioned as 'The AI for Problem Solvers'. It is a general-purpose AI system used for tasks such as generating code prompts, refining requirements, creating advertising strategy and copy, and processing creative content like storyboarding and video prompts.
DBOS is a database-oriented operating-system project and application environment that stores important system state in a database. Its practical focus includes durable, recoverable workflows, particularly workflows used by agentic AI systems.
Ingres is a relational database system created by Mike Stonebraker. It used the QUEL query language and competed with Oracle in the database market.
Neo4j is a native graph database. In the cited discussion, it is used as an example of a graph-database implementation, which Mike Stonebraker characterizes as less performant than relational representations of graphs.
Postgres is an open-source relational database created by Mike Stonebraker and subsequently maintained and enhanced by a volunteer community and companies. It is presented as a widely used database whose open ownership contributed to its adoption and continued development.
Searchable transcript of Do Databases Have a Future? | Conversations in Action — Imagination in Action (01:19:31). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Imagination in Action. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 I think there is a definite chance of that of that being the right thing to do. >> You are and >> [music] >> correct me if if this is a mischaracterization, you are in some sense the Godfather of the modern relational database. Everyone I know is building agents. You've said today's agents are mostly read only. What breaks first when agents start writing [music] to the world?
00:23 >> Well, if you ask me that question, I would say, well, health care and the building [music] trade seem safe. Everything else seems to be a risk. My take is you should [music] learn Chinese because I think we're like post-World War II Britain. [music] >> Welcome everyone to conversations in action. I've known Alex Westerberg for a third of his life and he's been consistently brilliant the whole time.
00:51 He ingest everything that's going on in the world and he thinks about it and I love having conversations with him. So, I thought I would collect some of the most interesting people I know to have conversations with Alex and with them. And today we have Michael Stonebraker, the 2014 Turing Award laureate. For those who don't know, as of 2014 it was a million-dollar prize and it celebrates the top computer scientist.
01:18 Many people refer to it as the Nobel for computer science. The Nobel family when they created their prize, CS wasn't part of the you know, the the list of prizes they were given out. Serial founder across five decades. He's five decades young and still going. He's a founder at in his 80s. He's creator of Ingres and Postgres, co-founder of DBOS, and he's connected with MIT CSAIL.
01:49 Mike Stonebraker is the closest thing software has to a living founding father. On a PDP-11. Uh he ended up outbuilding IBM. Um and uh he's really special guy. He's proved one size does not fit all. We're going to we'll go into that. And he said he's he's quoted as saying recently, "I'm not sure I'd recommend an 18-year-old study computer science." You know, this here's the guy with the the the top award in in that field.
02:20 He's blunt, he's fast, and he's funny. Uh and he spent uh 50 years uh being early and um and being right. Uh and uh and he likes to tell us uh that and and you'll hear some of that. Uh and I had he he and his wife over for dinner not so long ago and we talked about a banjo and a tandem bike that he and his wife biked across uh the United States on 3,600 mi uh from Seattle to uh Quincy.
02:49 Mike, uh welcome to the show. Alex, I have tons of questions, but uh want to give you the the first uh first bunch. >> Mike, thanks for joining. I have to ask you the obvious question in my mind, which is you've spent decades talking about how every other model other than the relational model is being dissolved in some sense by relational. So, no sequel and XML and object stores and map reduce, etc.
03:16 all being dissolved by relational data stores. And yet and yet we find ourselves in the middle of a super intelligence intelligence explosion where arguably the fundamental data structures look nothing at all like relational databases. And I've argued in past that the solution for the intelligence explosion, the the fundamental data structure is simply compressed knowledge and it's usually end-to-end differentiable, but looks nothing like a relational model.
03:44 Do you think that LLMs and foundation models are the first exemption to all these data structures otherwise collapsing into relational models. >> Uh So, first of all, uh let If you generalize that slightly to a genetic AI, that then essentially all the all the data sources uh are relational that feed a genetic AI. And so, I would argue that the having having done a bunch of a genetic AI, uh if you the traditional way to view a genetic AI is, you know, with quad-like systems or open claw-like systems.
04:36 And if you have two data sources and you need to join them together in quad or open AI or, you know, open claw, uh they will get turned into text and you will do the join as a text. You know, in text to text. And that throws away structure, which is clearly a terrible idea. And so, in my opinion, you want to do you want to turn If you're joining two tables, you want to leave them as tables and join them table to table.
05:13 If you want to join table to text, in my opinion, you're better off converting the text to a or the document to a table and doing a table to table join. So, I think uh foundation models, uh you know, in particular, you know, having having known uh a bunch of people who work for Anthropic, I think one of their big flaws is that they are text text only.
05:43 Uh and that that will prove a a big weakness. So, I think uh uh I I think it remains to be seen uh how all of this is going to evolve. And so, I think, you know, we're we're in the the bottom half of the first inning of a nine-inning game, and we'll see what happens. Uh specifically, you know, I'm working on a version of Agent AI. Uh we've been working with uh we've been working with the Department of Transportation of the city of Munich, Germany.
06:26 Uh and they Department of Transportation run, you know, runs the trolleys, uh and, you know, manages the stop lights and all that. And so, they have four full-time people who who are charged with answering citizens' question/complaint. And a very typical complaint is, "I don't have time to cross the crosswalk before the walk signal turns to red." And so, they have to they have to access the following data sources to get the the answer to that question.
07:11 First of all, they have to access uh rules and regulations both from the city and from the state of Bavaria. So, that's text to typical LLM stuff. Uh but then they have to access the diagram of the intersection. That's CAD data to figure out how wide the intersection is. Uh the rules and regulations say how how fast you can assume pedestrian can walk.
07:44 Uh then they have to access uh the trolley schedule because if you're crossing the crosswalk right when the trolley goes by, the lights all get screwed up. Uh they also have to access a map of of Munich to figure out which intersection you're really you're really talking about cuz they're they're probably going to say I I don't have time to cross the crosswalk in front of my house.
08:12 And you have that's an ambiguous question. So, you have to access a map to figure out which crosswalk you're talking about. And then lastly, uh it turns out that if the time of day is while school are opening or letting out, uh then the walk signals get shorter or get longer to accommodate uh the school kids. So, all of that goes into trying to figure out whether the citizen's complaint is viable or not.
08:49 So, that's that's one data source, rules and regulations, the text it's text. All the rest of it is various structured data. And so, they've tried traditional agentic AI and it failed completely. >> I I I out of curiosity, Mike, what is traditional agentic AI given that long time horizon autonomy is relatively recent? Is there a tradition of agentic AI that goes back more than a few months?
09:22 >> Well, my view of agentic AI is people figured out that LLMs are not by themselves sufficient to answer a lot of questions. So, you put a a bunch of junk around an LLM uh and turn it into a workflow. And that's labeled a genetic AI. >> So it's scaffolding, what what everyone else would call scaffolding or harnesses. And but maybe let let me put us on the original point.
09:52 So I my original question was whether in some sense LLM's foundation models were the ultimate counterexample to the thesis that everything would ultimately at least every data structure would dissolve into the relational model. And you told a story about Munich and about one narrow application, but I I really do want to press on the the core question of whether ultimately I I can pose it differently.
10:22 It's sort of which came first, the chicken or the egg? Which is more fundamental, the cell or the table? Where or the string, the the plain text string or the table? If I understand how you've thought for decades about this, you've been on a a quest almost to to reduce or let's say reformulate most of the data structures in the world into something that looks like relational tables.
10:47 But now we find ourselves in an era when foundation models are arguably not being pre-trained off of anything that looks remotely like relational tabular data. It looks like long strings of Unicode plain text and that's sort of that the native even the tokens aren't database tokens or table tokens. They're they're just Unicode sub-word tokens, BPE usually.
11:15 Is there in your mind, is there a future in which tables themselves are really no longer the primitives of data structures and instead everything reduces to to some sort of layer abstraction layer that sits below the tables that looks like BPEs or plain text in the end. >> So, you're making in in Let me Let me give you a long-winded answer to that. >> Go for it.
11:45 So, that that's that's a really strong I I think claim that there's there's some capability ceiling on text-to-SQL by foundation models. Would love to understand where you think the ceiling is. >> Sure. Uh so, there are four reasons why this stuff doesn't work. Uh number one, uh data warehouses are never in the in the public domain. They're always behind serious firewalls.
12:11 So, LLMs can't train on them. So, all of your fancy text stuff, uh you know, if you if you're asking a query, uh like what what department is Mike Stonebraker in, unless you're lucky enough to have some public data uh appear in some some news article somewhere, uh that's not going to be in the training set. And so, you're not going to get the answer back.
12:41 So, that's problem one. Problem two, which is, I think, even more fundamental, is that if you look at Spider and Bird, uh they have uh table names like employee and and department. And they have column names like salary and and {dot} {dot} {dot}. If you look at the MIT data warehouse, uh it has columns like underscore XYZ. Uh and so, non-mnemonic uh column column and table names, uh overlapping semantics, materialized views, sort of what I would call, just to generalize the category schema rot.
13:31 Uh so there's schema rot in all existing warehouses that have been around more than a few years, which is almost all of them. So, you've got to deal with schema rot uh just like you have to deal uh with you can't you can't train on until such time as things get cheap enough that I can actually train on the actual MIT data warehouse, uh then uh things are not going to work.
14:05 Uh moreover, there's lots of video idiosyncratic data in in the MIT data warehouse. Sort of things like J-Term turns out to be an MIT-specific thing. Turns out to be a 1-month term in January, uh which uh hardly anybody else has. And so of course if you ask questions about J-Term, you're unlikely to to find that from an LLM. Uh also if you ask uh what's the what's the name uh of the uh uh you know, of the of this computer science building at MIT, well, uh buildings don't have names at MIT.
14:54 They are numbered. Uh and so idiosyncratic data, uh and then real real data warehouses have four and five-way joins as the typical typical query. So, they're more complicated than the student-written queries that are used in the public benchmarks. So, the combination of all of those things is going to make it a long time before text to SQL is going to work in practice.
15:27 And if you have >> I just have have have to ask Mike, you say a long time. What does a long time mean to you? >> Uh well, I think uh you'll have to tell me when it will be when it will be cheap enough for MIT uh the MIT administration to train a large language model on exactly their stuff and produce uh a workable data, you know, data dictionary that has that has the semantics of each column and has the provenance so that uh when you have columns that are derived from other columns uh you can uh actually figure that
16:16 out. So, when is that going to happen? Uh I think in in the case of the MIT data warehouse, it's probably it's probably never. >> Incredible. Well, I I I I think it's fair to say then I'm a good deal more optimistic on this front than than you are. If if I were to ask myself the question of when do I think frontier capable models will be able to to auto-document a 1,400 table relational database, however pathological it may be, I would guess either now or in a few months.
16:54 Uh so so maybe somewhere between either now and never the the truth probably lies at some intermediate point. May maybe one more question uh just before dropping this topic. It one of the I I would say essential >> So, we're back to So, we're back to uh in my opinion, LLMs fail badly on a number of important uh application areas. One of which being uh enterprise data.
17:30 And putting putting together enterprise data. Because if you look at the Department of Transportation, uh you know, it wants to integrate a bunch of disparate data sources, only a minority of them are are text. And uh that's And there's all kinds of corner cases. Uh and that's very typical of enterprise data. And so uh what I'm interested in is real-world problems.
18:05 Uh and I completely agree that if you're trying to write marketing literature, uh LLMs work great. If you're trying to my daughter teaches high school math, if you're writing recommendation letters, it works great. Uh it's been much less successful in going into production use in enterprises. And I think the in my opinion, uh there are a whole bunch of reason reasons for that.
18:36 One is uh you can't you can't just give an you know, a gigantic AI what the text that corresponds to a four-way join. To a real data warehouse, uh you're not going to get the answer. Just not. Uh now, maybe maybe if the schema is well structured, maybe if you uh have uh very capable AI folks uh inside the enterprise or you're talking about Google. I mean, there are a bunch of cases where uh, I I would not be so negative.
19:22 But, if you're talking about, uh, you know, for example uh you know, the uh price if you're talking about, uh, you know, let's just say uh, a large drug company and any traditional enterprise I think the things things are not nearly as rosy as you assume. Now, I'm sure that I'm sure the truth is somewhere in the middle. But, uh, for everything I read uh, everyone on the planet and if you restrict those to people inside enterprises uh, who have uh, everybody has built pilots, uh, AI pilots in all kinds of, uh, ways.
20:16 And if you look at the percentage of them that have gone into production it's like it's like 20%. Uh, and so, uh, the technology uh, enterprises are having trouble adopting the technology. >> Maybe just a closing question before perspective John. It would seem to me, Mike, that what you're describing, if it exists at all, is at most maybe, uh, training data challenge.
20:48 Maybe there's a an opportunity for frontier labs to do a better job of mid-training or post-training or reinforcement fine-tuning of their models on hard challenges in relational database management and maybe their information diet or their post-training diet could include a much larger lump sum of database management versus computer use assistance or other modalities of interaction or training.
21:14 Do you think that this issue that you're articulating, if it exists, could be cured through just more aggressive post-training frontier models on database tasks? >> Well, for a starter, there there aren't very many enterprises who are willing to entrust their data to OpenAI. So, you have to be talking an on-prem model, which means an open-source model.
21:45 Uh so, you have to start there. Uh and there there's serious issues with access control because my salary is in the MIT data warehouse. And your foundation, your model of any sort, is going to have to adhere to uh to any access control that that uh is is uh in place. So, there there's a bunch of challenges that you guys tend to gloss over uh that are, you know, about the rubber meeting the road.
22:23 >> I see. So, so maybe if I were to just try to uh caricature, maybe the wrong word, to to try to distill what I I think you're you're saying, is it is it your position then in some sense the foundation models, including at post-training time, have been trained off of the most accessible, most public or public-adjacent data, whereas there's almost a dark matter of training data, the most closely held, the most uh secret in some sense, training data live inside relational databases everywhere, and there's almost a blind
22:57 spot by frontier labs to the known unknowns of what's inside those enterprise databases and how they're structured that creates a sort of blindness by the the even the the state-of-the-art frontier models to being able to interact with those data sets. Is that the idea? >> Yes. >> That's fascinating. Back John, back to you. >> Yeah, sure. Mike, you launched D Boss 36 months ago with Mati Zacharia and raised 8 and 1/2 million dollars on the argument that the cloud has outgrown 33-year-old Linux, that the operating
23:34 system should be an application running on a database, not the other way around. What does D Boss understand about the cloud that Amazon and Google don't? >> Well, first of all, uh what you what you what you described uh was our our research project. And when we started pitching the idea to VCs, they immediately said uh it's a fool's errand to try and replace Linux.
24:12 Uh that's a 10-year sales cycle if if it's possible at all. Uh and so forget it. But they said, "However, uh what you have is really uh besides that is really valuable. Because what what you basically your premise was that uh you know, all all important state uh inside an operating system belongs in a data management system cuz that makes it, you know, recoverable.
24:48 Uh you can write you can write a scalable scheduler. Uh dot dot dot dot dot everything everything is much easier to deal with. Uh so what in the process of the of the research project, uh we wrote an application environment for Dboss. And that adopted the same principles to the programming language. Which is put all state all important state in a database.
25:20 Uh and you'll be much better off. So, what essentially and and the VCs all said, "That's a terrific idea. Run with that." So, this was about the time when uh Agentiq AI was coming on the scene. And Agentiq AI, uh as we pretty much see it from Dboss's point of view, is a bunch of a bunch of stuff in a workflow and a graph around one or more LLMs. And so that that number one is a workflow.
26:01 And as Alex will no doubt be happy to point out, it can have loops and all kinds of craziness in it. Uh but it's a workflow. It's not a single application. It's a bunch of stuff. Uh number two, uh it's very long running. It's computationally expensive cuz LLMs are computationally expensive. So, you have an important class of applications that are A workflows and B computationally expensive.
26:35 So, uh much of the Agentiq AI community has come to realize that what they want is durability. So, that in this workflow when you complete a step, you want it you want it if bad things happen, you don't want to back up to the beginning, you want to back up to the last completed uh set of steps. So, they want durable workflows defined as uh if if anything bad happens, uh you can back up to a known state and you know, at the end of every every step that's successfully executed.
27:23 So, this looks this looks a well, this looks a lot like it is uh the D in acid and you want the programming language implemented exactly the same way that the database implements it, uh which is put state in the database and life is great. Uh there's a whole bunch more complications, but uh so, basically, what D Walsh is selling or is supporting is uh durable durable workflows, especially when applied to agentic AI because that's that's a very important use case that that folks are working on.
28:10 >> Great, thanks. Everyone I know is building agents. Everyone you know, many people listening today are building agents. You've said uh that today's agents are mostly read-only. What breaks first when agents start writing to the world? And how much of the fix did your field already build 40 years ago? >> Uh uh somewhere between most and all of it. Uh well, the first problem is that uh let's say you have you have a workflow that's adjusting uh Alex's salary cuz you think he's well smart, he's underpaid.
28:53 So, you have a a workflow uh that's somewhere in the middle going to adjust his salary. >> Or just maybe maybe I might drop my table entirely. >> [laughter] >> Yes. So, so then uh what happens is you've adjusted his salary in a step in the workflow. Uh and that's a transaction. So, it basically the transaction is committed once that step finishes. Let's decide that you get to the downstream somewhere and you decide that the raise you're giving Alex, you know, is is too great and you want to unwind everything.
29:38 Uh so, you have to be able to back up past a committed transaction. Uh and if somebody else has updated Alex's salary between when you did it and now, then you can't just physically back out his change. Got to logically undo it. Uh and that was studied in the database literature in the '80s, the concept called sagas. Sagas. Uh the the seminal paper is by Hector Garcia Molina and Ken Salem uh in I think it's the 1985 Sigmod.
30:26 So, Google sagas. Uh sagas basically are how to back out updates over uh completed transactions. >> [cough] >> Uh [clears throat] and then there the rest of ACID. Isolation doesn't really apply because uh once you have something bigger than a transaction uh unless you make the bigger thing a transaction, which no one wants to do. Uh then you can't have isolation.
31:01 Uh and then there's atomic. Uh and in my opinion, most of the time you want this agentic workflow to either run to completion or look like it never happened. And so if the agentic AI is uh what's behind uh a website that sells bicycles. Uh well, the customer picks out a bicycle. Uh says, "I want to buy that one." And then the workflow is you first of all see if you've got it in stock.
31:35 Uh if if not, you probably want to check an alternate warehouse. Uh if you have it in stock, great. Then you go on to check his credit and decide whether you want to sell him the bicycle. If his credit is okay, then you want to accept payment. That's another transaction. And if he gives you a bad credit card, you want to unwind that. Uh and finally, if everything goes okay, you then want to ship him the bicycle.
32:08 Uh and if he gives you a bad address, then you can't do fulfillment. You got to unwind the whole thing. So, buying a bicycle should be atomic. Uh and so in all of these uh in spite of all these corner cases, so in my opinion, you want to support atomicity uh for workflows. Uh and that gets a little bit tricky. Uh but we we wrote a paper that was in uh uh uh in the cider uh conference last year uh on this topic.
32:45 And so, if you just look up the cider proceedings from 2025 uh you'll you'll find it. >> Great. So, I I have very quick questions about your past, and then I want to uh pass it back to Alex. Um Mike, you grew up um in a mill town on the Maine border in New Hampshire. Your dad worked at GE, and later ran a shift in a fiberboard factory. And your mom taught school.
33:12 And your father moved the family to Newbury Mass for one reason. The town would pay a boy's tuition to go to Governor Dummer's I I Dummer Academy, uh the oldest boarding school in North America. Uh take us from that kid uh to Princeton, then to um uh Michigan, and then to Berkeley, and now MIT. When did you first suspect computers were going to take over everything, given that, you know, you're a historic figure in the uh computer landscape?
33:47 Well, when I when I went to Princeton uh I decided to major in electrical engineering. Well, prob- probably cuz I had uh uh you know, an 800 SAT in math, and a 500 and something SAT in English. So, I was destined for uh something in the sciences. Uh and uh you know, for no particularly good reason I be I became an engineer. And once I was an engineer my father was an electrical engineer, so why not?
34:25 Uh >> [snorts] >> so by the time I got to junior year at Princeton it was completely obvious that computers were going to take over the world. Uh and that wasn't that wasn't that Well, maybe I didn't view myself as prescient uh at that time. It's just obvious that if you wanted to get a job uh being being computer-skilled was a really good idea. Uh and so back then there were so you could just call yourself a computer science uh expert and start taking computer courses.
35:10 So, [snorts] uh that that uh that got me uh into Well, that that I guess the other piece is I'm exactly the age uh where when I graduated from college, the draft and the Vietnam War was going full blast. So, my choices were uh go to Vietnam go to jail, go to Canada uh or go to grad school. So, that was that was kind of a no-brainer. >> Great. So, so let me fast forward to the mid-80s.
35:45 A lot of people thought uh your Ingres was better than another system that starts with an O. If I mention the word Oracle, do you get upset? No, I can mention Oracle, yeah. Um I know the Ellison family owns Paramount and uh they're producing Star Trek, so I'm I'm tracking them. Um what happened? What were the sales tactics of that era and who actually won uh in the market?
36:09 Uh the best technology or something else? Why is Oracle more known than what you created? And if you met Larry Ellison on the street, what would you say to him? You have 40 years to have thought about this. Well, Orac- Oracle started 1 year earlier than than Ingres Corporation. And so, by 1984, uh which is 3 years in, uh we were we were gain we were growing faster than Oracle.
36:41 And we were poised to overtake them. Then the the event that happened was IBM released DB2, which ran ran SQL. And uh Oracle supported SQL uh out of the box. And even though QUEL was a better better data sublanguage in everyone's estimation, uh that was the end of QUEL. And so, in the next uh 18 months, uh we well, next 12 months, we re-engineered uh Ingres to support SQL.
37:23 But by then, uh Oracle had leaped ahead, uh and the game was pretty much over. So, uh you can blame IBM for the demise of Ingres. Uh or or not the demise, the fact that that Oracle became the dominant player. Uh and I think they they engaged in all kinds of very disreputable, in my opinion, disreputable sales tactics. Uh they for instance, one of their favorites was uh was a concept called referential integrity.
38:00 Uh that became important uh in the in the early '80s. So, we had it, and we said, you know, yeah, it works the way it's supposed to work. So, if you looked up in the Oracle manual, they had manual pages for referential integrity, and a little footnote at the bottom that said not yet implemented. Uh so they were happy to confuse present tense and future tense uh and uh got away with it.
38:36 >> I have one last question on the past uh but before I go to that, if you saw Larry on the street, you know, have you seen Ellison since then or or no comment on this one? >> Uh I would rather not comment on >> Yeah, okay. The trillion-dollar industry that got away from you. All right. Uh Berkeley's rule was that everything you built was public domain.
38:55 So, uh the Ingres tapes went out to anyone who asked and the uh Postgres uh shipped under a license that let anyone do anything. Then you walked away, the two graduate students bolted on uh SQL, a volunteer community adopted the orphan, and today Postgres is the most used database on the planet. You've called it the epitome of open-source software because it doesn't belong to anyone.
39:24 So, um be honest, was giving uh it away a strategy or by accident? And what does Postgres's win teach us about who should own the infrastructure of the AI era? >> Well, first of all, uh it was not a Berkeley policy to declare stuff open-source. That was my decision. Uh and you know, all all I wanted was to I was trying to get tenure and all I wanted was to get Ingres to be better known and the obvious way to do that was to give it anyone who wanted you know, a tape.
40:11 Uh so that that became a tradition at Berkeley. Uh you know, basically starting with it with Ingres. Uh, I think >> So, you set the precedent. >> Yes. And I And I think, uh, I think it's it's it's complete happenstance that you know, that that Jolly Chen and Wei Hong, who were the two students you're talking about. So, there was post SQL uh, and this pick up team of volunteers with no relationship to Berkeley whatsoever.
40:53 I didn't even I didn't know any of them. Just picked up the code and started supporting it and enhancing it. Uh, and they've been doing so all these years. Uh, you know, for the last 30 years. And I think what I find totally miraculous is that the there are 20-ish uh, the committee of of people is about 20 20-ish people. Uh, and the Microsoft uh, EDB uh, and Google among others probably.
41:32 I mean, I you know, what's happened is that uh, various [snorts] enterprises have basically contributed to supporting and enhancing post gress uh, under the the rubric that no one owns it. And I think that way, uh, and I think in fact the reason post gress is so dominant is that Oracle bought MySQL and it became owned by Oracle and the developer community fled.
42:08 And so, I think, uh, I mean, open source has wonderful advantages. And I, you know, it it it's sad that uh that OpenAI and Anthropic don't take that point of view. And I think one of the we'll we'll see what happens because in my opinion, what's quickly going to occur is that it's not going to, you know, that all all the foundation models are within a few percent of each other in terms of uh accuracy.
42:47 So, pretty soon it's going to be uh answers per dollar, not just answers. Uh or at least in a whole bunch of domains. It's not it's not clear that will be true of Claude code, but that'll be true in a bunch of AI. And once once that happens, then open source is going to be wildly cheaper uh than closed source. Uh and especially uh DeepSeek is a great deal cheaper than than uh than the American foundation models.
43:24 So, we'll see what happens, but uh in my opinion I hope it it turns out to be open open source wins. >> Great. Well, thank you, Mike. Uh back to you, Alex. >> All right. So, um it during all of that, Mike, I think we glossed over a bunch of things I'd like to go a little bit deeper. So, D boss, to start with, I remember as as I assume you do, Microsoft Longhorn, uh Bill Gates's quixotic attempt to rebuild Windows ultimately resulted through twists and turns in in Vista, but the original concept, as I recall, was
44:01 actually to rebuild Windows from the bottom up as an object-relational database or around an object-relational database, and it was just total failure at least from a market perspective, even if some of the concepts were impossibly elegant. What do you think went wrong with Longhorn? >> Uh so, well, first of all, uh I've talked to a whole bunch of the people who were there at the time.
44:28 Uh and essentially all of them said uh Longhorn itself was a really good idea. Uh and Microsoft up the implementation uh big time. >> That's the technical term, I assume. >> Yes. And and they kept, you know, they kept changing the specs. They kept, you know, it it was subject to feature creep and and got so unwieldy big that they finally had to kill it.
44:58 So, I think the it was it was mismanagement, not not the fact that not it was a good idea that was badly managed. >> So, do you think it's not intrinsic? Like I In your description of Dboss, if I understand the history correctly, you can correct me, it sounded to me like what you were saying is originally Dboss started as trying to realize what one might call the Longhorn dream, which is an OS maybe even at the kernel level that's based predicated on a relational database.
45:30 And then, if if I understood the narrative correctly, it sounded like through pivots and twists and turns, it became a little bit more about a hardened durable relational store for agentic AI and less about literally replacing good chunks of the kernel with a relational database. How Is that the right way to think about it? And if that's even remotely the right way to think about it, when do we get our Longhorn?
45:55 >> Yeah. Uh I think so so, from the VC's point of view, Uh, the sales cycle for an you know, replacing Linux uh, you know, is a decade long. And and so, uh, if you're trying to displace somebody, uh, you know, it's a very long sales cycle. Uh, also, you've got to have uh, all of the junk that that goes around uh, that goes around Linux, you know, you have to have device drivers for every conceivable gizmo on the planet.
46:42 Uh, and so, there's just a ton of infrastructure. Uh, and I think if we just had said, uh, you know, we want to we want to run Linux plus plus, uh, and say, you know, keep keep Linux, you can call it Linux. Uh, we're going to swap out some of the guts uh, and get a better system. Uh, that would have been a much better strategy than saying >> Let let me then push on that a bit because we are in the era of super intelligence miracles and great programs that historically would have been onerous to the point of being
47:21 non-viable that are suddenly becoming viable especially for example, the Linux kernel is now being rewritten step-by-step in Rust. And historically, that would have been Well, for first, it was socially unacceptable and then became more socially acceptable, but now in an era when even Postgres, as I'm I'm sure you've seen recently by the community was there there was a complete port of Postgres to Rust.
47:44 Surely, in the era of super intelligence, it must be possible for someone without enormous resources to undertake a project to say take a Linux kernel or a BSD kernel or you know some applicable open source kernel and just do an architectural rewrite with a relational database at the core. Whereas previously that would have been borderline impossible for the reasons you mentioned like device drivers and all of that the other hassle filled elements of a kernel.
48:14 >> Yeah, well remember that what I was just describing was 2023. Uh >> That that's the stone age, right? >> Yeah. So I I think one of the one of the fabulous things about Claude is that it it enables just throwing code away and rewriting it uh becomes a feasible alternative to pat patching uh and and you know extending what you've got. And so I think every enterprise on the planet should take your advice and try and try to use that tactic to retire their their legacy code step by step piece by piece uh and get
49:03 something that that's modern uh and maintainable. So uh I think it uh you know Claude enables new software strategies that weren't feasible 3 years ago. >> You're I mean I'll I'll direct this even more pointedly. You are and correct me if if this is a mischaracterization. You are in some sense the godfather of the modern relational database. Who who better than you Mike to be the one who finally maybe it's third times the charm.
49:36 First time was Longhorn, second time was D Boss or or other similar attempts. Third time empowered by Claude and and other frontier models and agents to literally finally definitively give the world the database based operating system that we've been trying to build for decades, but maybe weren't smart or capable enough. Now AI or Claude or both give us the ability to do that.
50:00 Have you ever thought about taking that big swing? >> Uh well, right now uh the VCs who are backing me boss would kill me if I said I'll do that. But I I think >> You don't think it's just like a weekend project that you could do on de minimis budget without venture capital financing? >> Uh well, I think get getting getting a prototype up and running is as you say doable fairly quickly.
50:35 But then then you you have to make it really work. You've got to make it fast. Uh you've got to support it. And so I think you know, it will take it will take it will take a fair amount amount of VC money to make this happen. And I'm in I'm in the conflict bad conflict of interest position right now. But yeah, I think that that would be a really nice thing to to try and do.
51:06 >> Maybe it's a call to action. I I I guess the the title of this or the name of this podcast is conversations in action anyway. So maybe Mike, what I think I heard you say is call to action for the community. Since as as you say you're you're you're being steered by venture capital in certain directions. If someone else wants to go and create that that final definitive database-powered OS kernel, I assume that would be something that you'd find greatly interesting.
51:32 >> Sure. >> Is there another holy grail that that you'd like the community if if many many people see this if if you have a call to action to everyone other than that that final definitive database-backed OS kernel at the end of the database rainbow that you'd most like to see or is that as holy as the Holy Grail gets for you? >> Well, I think the if if the thing I would like to see even more than that is that every enterprise that I know of uh is just dragging along a huge cinder block called legacy code.
52:10 Uh and the worst of it is uh that's it's in COBOL and you don't you can't hire developers anymore and dot dot dot dot dot. Uh and what seems to be happening is uh so your enterprise buys somebody and it comes with a whole suite of of of systems that keep that acquirer running. And so no one stops and and goes through the effort to uh integrate the new guy's stuff with what you've got.
52:52 So you just create another silo. And so enterprises are just full of data silos that impede their progress. Uh I think the the great possibility of of sort of quad at al is the fact that instead of creating a silo you could actually uh do the integration up front and start to get rid of silos rather than building new ones. Cuz I think uh if you're spending 95% of your IT budget on maintenance, you just don't have you don't have much money to to do new stuff.
53:39 And and so I think, you know, the the legacy is just weighing weighing people down big time and it would be uh I would yearn for a world where that wasn't uh such an impediment to progress. >> Call it maybe a technical debt holiday. You'd like to see a technical debt holiday for the entire world? >> Yeah. That's good way [clears throat] to >> Tax free holiday this weekend and we have a new one.
54:07 I'm also curious just another thing we we drove by pretty quickly Mike you mentioned the ability to reverse transactions. There's another data structure, another paradigm in in software that superficially looks relatively little like RDBMSs and that is call it get style DAG based is usually plain text software but maintaining a directed acyclic graph for commits and branches and merges and so on call it get DAG style data structures.
54:41 How in your mind assuming the question makes sense and if if not I can repose. Seems to me these are almost two separate structures that semi-peacefully coexist but really in an ideal world we'd want to just pick one. On the one side we have tabular RDBMSs in the style of yours including Postgres. On the other there's a world where everything looks tree-like or DAG-like and you can branch state and you can merge state and it's distributed natively a little bit more natively than RDBMSs including Postgres are and you
55:22 can roll back commits and all of these interaction metaphors that I think you were gesturing at earlier as being still somewhat onerous in in the very best RDBMSs like Postgres are just natively built into Git. Do you think there's a world assuming you even buy the premise where the future of databasing looks more like Git and less like Postgres or do they somehow merge?
55:51 >> That is >> Do you think there even is any sort of future for first-class graph databases or at best sort of passing passion that that ultimately gives back gives way again to conventional RDBMS? >> Well, Andy Pavlov and I wrote a paper uh 17 years ago called What goes around comes around. Uh which was an argument that uh that all the all the data all the data structure ideas uh you know in be prior to 2008 were were basically bad ideas.
56:34 Uh we wrote another paper called What goes around comes around and around last year. It's in SIGMOD Record sometime last year. Uh and it basically argued uh that all all the data structures that have been proposed uh in the last 17 years are also bad ideas including graph databases. Uh and so the problem with graph databases is if you're looking you can model the graph uh as you know as an edge table and a node table.
57:11 And when you start running queries against uh the tabular representation uh of a graph uh it tends to be faster than with native graph implementations. Uh and in fact Neo4j is not very performant at all. >> I know. >> And so still So the question is uh and in fact Amazon has a graph database and it's implemented under the sheets exactly the way I just described.
57:45 >> Neptune specifically I think you're referring to. >> I Yes. >> Yeah. >> So I think it it's so far has been beaten by relational databases. Uh the one case where it ought to work is if you want to find the shortest path from node A to node B. You want some algorithmic thing. Uh and there the problem is that uh if you can represent that graph uh as you know, in main memory in some proprietary data structure uh then uh Sheri I can't remember his first name.
58:30 Uh has written a whole bunch of papers that say, you know, complicated algorithms you don't want to use a database system at all. Uh you want to you want to you you want to use some main memory fancy representation and and some and some very application specific algorithm and you win by two orders of magnitude. So I think it it graph databases data use case on which they are the they are the answer and so far I don't see it.
59:04 >> All right. Let me jump in here. So um Mike you you received the um the Turing prize. Is that right? >> Yes, I >> Do you do you have it in your office? Can you show us? >> Uh I'm not in my office. It's uh it's actually on the on the mantle in my in my home. >> Okay. So today it's been given 60 times, I think, to 81 people. 60 years from now, who's getting the Turing Prize and for what?
59:34 And I know a number of times two or three people got it. Will an AI be getting it? And will people who don't even study CS be getting future Turing Prizes? What do you think they'll be celebrating in the decades to come? And when you look back, and I'm sure you're tracking who got the the Turing, is there anyone that you think is undeserving? Or are there people that that the committee missed that you think should have gotten it?
01:00:01 Uh I know I was with Bob Metcalfe around the time that that he got his prize for his work that was 50 years prior. He was very emotional. Um how important is the Turing Prize uh in this moment that things are moving so fast? Uh last 6 months have been, you know, you know, not not something I think people anticipated. And what do you think the next 24 months are going to look like?
01:00:24 And what role does the Turing Prize play in helping to shape or celebrate or challenge or provoke? >> Uh so, you asked about four questions. Uh number >> grade you on your answers. Good luck. >> Uh the 50-year the futurist uh 50-year forward >> 60. >> I have absolutely no idea. And I think uh if you if you were to predict uh in 1990 uh what were the main advances uh in computation during the '90s?
01:00:59 Uh it turns out that the fax machine and the internet you know, were the answer. >> A little known fact, the post office was was offered the fax machine, but they they didn't see any use for it. They could have commercialized it much earlier. And around that time, Bill Gates came out with a book called The Road Ahead, which I have, and it didn't mention the internet once.
01:01:22 >> So, I think I think uh I think in these times speculating even 5 years ahead is I think a fool's errand. So, I mean, I don't know. >> All right. Well, what about an AI getting the Turing who who was passed over, who shouldn't have gotten it? If you won't answer what's happening 60 years from now, take on some of those. >> I think I think personally that the 1979 winner, which was Charlie Bachman, was undeserving.
01:02:01 So, so I think he he was he was I think it was maybe it was earlier. May- Anyways, it during the '70s. I think he he he he in my opinion is not deserving. I think the committee in the in the in the odds was stacked with programming language types who gave it, you know, disproportionately to programming language folks. And so, I think now now there's a great deal of political, you know, fairness being being you know, being rendered.
01:02:48 So, I think it's become somewhat political. As to uh how important it is, I mean, I think it it recognize it recognizes you know, incredible achievements. And I think the problem is is it became easy to was easy to pick people when it was, you know, Knuth and Minsky and, you know, the founding fathers. It's now a great deal more difficult cuz the field is way bigger and there are way more way more actors.
01:03:27 And I think I think it's it's a really important memento for people. My bigger problem though is in the age of AI, what are computer science departments going to teach? Cuz the people at MIT are having are gnashing their teeth saying, "AI can write all our programming assignments way better than anybody else can." So, what are we going to have for assignments?
01:04:07 And what are we going to have for tests? And what should we be teaching? If AI is going to be the major if Claude is going to be the major Claude et al. are going to be the major writers of software. And so, I think I wonder how that's all going to turn out. So, >> Do you even even think, Mike, that programming language that you've created a number of programming languages at least maybe more domain-specific languages?
01:04:42 You created as you mentioned Quel. I think there was post-Quel. Maybe there were others. Do you think humans even in the near-term future will ever need to learn any programming languages at all? Or will this just be handled under the hood by Claude and ChatGPT and frontier models in general and humans learn English for at least a few more years at least and that's it.
01:05:05 Programming languages as a whole are cooked. What do you think? >> That can can said about most everything in computer science. >> I agree. >> And and so >> computer science cooked? I mean, should should should the departments fold up shop in a graceful way over the next few years and just say, "All right, field is over. Let's move on to something else."?
01:05:26 >> I think there is a definite chance of that of that being the right thing to do. >> What do you think a graceful wind down of the discipline of academic computer science would look like? >> Uh everybody everybody would transition to some other department. >> Do you have any favorites? What do you think I mean, math also in in in the the throws of similarly being cooked by AI?
01:05:52 We see that every day. Do you have a favorite successor discipline for people in computer science to migrate to? >> Let me let me put the question more broadly, which is if you had an 18-year-old son uh or daughter, what would you advise what career path would you advise that person to take? >> You want my view or is this a rhetorical question? >> Well, if you ask me that question, I would say, "Well, uh health health care and the building trade seem safe.
01:06:25 Everything else seems to be a risk." >> For a few years, but a career presumably Yeah, as John says, robots and many other advances coming. Do you Do you think, Mike, that health care and uh the trades, quote unquote, are on the time scale of say longer than 5 years are a safe bet for a career? >> Um that's a good question. I think uh That's a good question.
01:06:52 I have a uh Will the building trades succumb to robotics? Uh that's a good question. And will health care succumb to robotics? We'll see. I mean >> I do think we're going to be finding ways to make better humans and use our you know, what's unique about humanity to to help with that and we're going to also have a symbiotic relationship with technology like we've never had before and everything is going to be different and we're at a transitional point and we're going to build on, you know, what is but the future is
01:07:30 going to be very different. >> So, you you have an 18 you have a high schoolish collegeish kids. >> Yeah, I have a I have a son who graduated college. He's now working at Whoop. I have a daughter who's pre-med. I I try to tell her medicine's going to look very different. I and you know, I I think she's just trying to get through the way they teach medicine right now and I have a young son who wants to study business and finance and I, you know, Cornell and Tulane and and I I do wonder you know, what what those skills
01:08:06 are going to offer and and what's the way for career advancement? Is everyone going to be doing startups and and, you know, working with AI? Uh and I think, you know, these consultant firms and business schools and you know, I I I think they're they're in for a rude awakening. >> So, what would you tell your 18-year-old offspring to to educate >> I I tell them to to track everything that Alex Wester Gross is saying and uh make sure you understand it and then understand what it means for you.
01:08:42 I I think it's important to follow these trends that uh that we've never been in a situation like this. You can't really pattern match to the way things have been and uh I think find some best friends that you want to build a startup with and and um you know, understand the landscape of where things are going and I think there are a number of people that are going to just throw up their arms right now.
01:09:04 A number of people are going to be bridge for the old economy to the new economy and then they're going to be leaders and people need to decide who they want to be. Uh but Alex, you know, you're you're very thoughtful on this and you're you're also very shocking. Sometimes you say things that I think aren't mainstream just yet. So curious for you to uh add to the dialogue and have Mike tell us what he thinks.
01:09:28 >> That's my job, I think, to drag the Overton window and Mike Stonebreaker kicking and screaming into the the near future, which will be just wild. Yeah, I I think it to the extent the question of what to advise an 18-year-old now in in today's world, gosh, assume that we'll have brain computer interfaces that work reasonably well in the next five years.
01:09:51 Assume that >> Hey, just on that one, if we have brain computer interfaces, are you are we even still using English? Like are we are we communicating through other ways? >> You know, I I I think the QWERTY keyboard is going to be with us through the heat death of the universe. So yes, I think we'll have English as as well. It'll just be faster and you'll be able to do more with it.
01:10:09 I think we'll still have English, absolutely. But I I think people will be able to do much much more than they were able to do before that. This is already the case now, so that's borderline truism. What would I advise an 18-year-old to do? Take your favorite sci-fi concept from Star Trek or whatever other sci-fi you like and just implement it today.
01:10:29 I I think this is the era when finally, for the first time in human history, it was possible to take concepts from decades ago that people thought were either centuries or millennia out, if ever, and just implement them today. And I I think my advice to any 18-year-olds listening is if if you're not making sci-fi happen, what are you doing with your life?
01:10:54 Aim much higher in in short. uh don't don't be distracted uh or disheartened by certain disciplines, many disciplines getting succumbing to automation. That that's just the the prelude. That's just the appetizer to what is about to happen, which is I think basically every sci-fi trope happening everywhere all at once, and that will be well worth uh any um melodramatic destruction of fields or, you know, to Mike's point, if if computer science as an academic discipline has to be gracefully wound down in favor of other
01:11:29 things, that's a small price to pay for enabling many sci-fi concepts to suddenly be brought into existence. >> All right, I want to hear Mike. You you just heard two people's take on this. What what what's your take or react to what was just said? >> Uh my take is is uh you should learn Chinese. Because I think uh we we are we are in an in the US, we are we are like post-World War II Britain.
01:12:00 Uh you know, we are we are going to be overtaken if it hasn't already happened uh by the Chinese. And certainly being accelerated by many of the policies of the current administration. So, I think um get ready get ready to be in a world dominated by someone other than us. >> I'll I'll challenge that one. I I think it it's relevant to the database discussion and to AI as well.
01:12:28 Taken literally, I think and I I maybe be curious to hear Mike your response to this. I I think, you know, other than perhaps for purposes of uh cultural cross-pollination or intellectual stimulation and uh overall mental health, there's approximately no point at this point to learning another language. AI translation, ironically, the transformer itself arose from attempts to improve statistical machine translation.
01:12:58 AI is quite good at translating to other languages. That's taking your comment about learning Chinese literally. Taking it metaphorically, I I'd be curious. I Do Do you really think that Western technology or the West more broadly is is somehow on the decline in an era of superintelligence and that China really has something that's a durable advantage over the West that will inevitably, Ray Dalio style, perhaps, suddenly or over time precipitates the West decline.
01:13:31 What What do you see as the reason why anyone should bother learning Chinese either literally or metaphorically? >> Because uh because I think the chances of this hypothetically 18-year-old working for a Chinese company in 15 years is fairly high. Uh I mean, I think the the dominant companies in the world will become Chinese. You know, that Temu is the answer, not Amazon.
01:14:11 Et cetera, you know, et cetera, et cetera. >> That's interesting. I mean, so with perhaps remaining time, I'll just again push. How do you reconcile on the one hand, assuming you agree with the premise, so much AI automation entering the world, and on the other hand, the the premise it it it almost I mean, you're evoking for me the the sort of outlook sort of worldview that I remember from the 1980s where there was concern by some in the West that Japan at the time was going to overwhelm the West and everyone would be
01:14:44 working for everyone in America would be working for a Japanese company and for a variety of reasons, that future most certainly did not play out. And uh so >> CEO had, you know, helped uh egg that on. >> Yeah, so why why is this time different? Why is why why do you think it would be the case that uh Americans in 15 years from now are going to be working for Chinese companies rather than what to me seems far likelier, which is to the extent Americans have recognizable jobs at all, they're working for AI CEOs, they're
01:15:14 working for American AIs, not Chinese CEOs. >> Well, I think uh the we we have a on the global stage, we have a competitor uh which is totally focused on STEM, has 40 extremely good uh universities that are turning out uh that STEM talent that has 10x the population we have. And if you look at the grad schools, including MIT, uh in the US, they are largely uh they're like I don't have the statistics in front of me, but uh my suspicion is uh they're probably 50% uh you know, Asian.
01:16:12 Uh if you look at the SIGMOD proceedings, uh 20 years ago, SIGMOD proceedings were probably 80% US paper, US-based papers, 10% European, and 10% Asian. If you look at SIGMOD proceedings of, you know, last year, I think it's probably about 2/3 Asian papers, uh and the rest is scattering from everywhere else. So, if you just view conference proceedings as a leading indicator, uh I think we are being overtaken by the Chinese.
01:16:59 >> And you're you're not at all I mean so so that my counterpoint to that would be um so at NeurIPS this year I I've remarked publicly that the most popular language I heard in in the hallways was Mandarin. So, I I accept that part of the premise, but the part I'm not buying is that somehow that's the terminal story for the time scale of 15 years. Uh superintelligence is so strong you can ask superintelligence right now to write SIGMOD papers and they'll probably be pretty decent.
01:17:27 So, why isn't it the case that forget about Chinese authorship at SIGMOD or some other ACM journal or proceedings overwhelming American authors. Why Why don't Why wouldn't you expect it to be the case that superintelligence is authoring the vast majority of papers in these proceedings 15 years from now? >> But if if that actually occurs, then it's going to render So so far the major advances have come from a relatively small cadre of very smart people.
01:18:03 If this if this cadre becomes unimportant, then basically we live in a world where machines are in charge. And >> Exactly. >> I don't I don't want to be in that world. So, maybe I'm glad that I'm old. Uh and >> Then then you better hope Mike against that longevity escape velocity. Longevity escape velocity it sounds like you're you're betting against it again.
01:18:31 I I would for what it's worth I would I would bet for the machines and probably for humans merging to some extent with the machines as well and not at all worried about the the 15-year time scale. >> Okay, well. Then >> Maybe a happy note to end on. >> You You will probably be alive. I I don't expect to be alive to to see super intelligence. >> Well, I I'll end on that.
01:18:58 I find you to be super intelligent and I really value um you know, your leadership in the field and uh you know, keep keep going and you're an inspiration to many and uh thank you for your contribution and uh I'm glad uh today people could get a better sense of where you've come from, where you are, and where you think the future's heading. So, thanks for joining us on Conversation in Action.