Not fully yet. Current AI models can perform powerful calculations, combine known techniques, and sometimes produce creative or unexpected results, but they generally lack autonomous mathematical intuition, theory-building ability, big-picture understanding, and reliable evaluation of long arguments. They may develop these capabilities with continued progress and different training or reinforcement-learning environments.
🔒 7 more in the full analysis
🔒 7 more in the full analysis
Searchable transcript of Can AI Learn Mathematical Intuition? — a16z (01:03:12). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by a16z. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 The goal of mathematics is not to produce mathematics papers. It's to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that's like pretty unsatisfying. >> Comparing anthropic with open AI, do you detect any differences in how that is similar to human reasoning? >> They definitely are not good at it autonomously, but with some hints, you can kind of get them to do something interesting.
00:19 A lot of progress in mathematics comes from like letting a thousand different flowers bloom and people pursue their own curiosity and then the boundaries of knowledge expand in some fairly uniform way. >> What has been the most impressive result so far? >> My favorite fully autonomous result by an AI so far is the solution to the Irish unit distance problem.
00:37 There was some lema I wanted to prove none of the frontier models could do it. So I like worked out a ton of examples on my own and I realized oh well maybe like here's some reason why it could be true. Once I had that statement, the models were able to very quickly prove that sort of better statement. >> How should like the mathematics community kind of best adapt and benefit from this.
00:57 >> Um >> I am so excited to have you on Daniel. Um and so Daniel is a professor of mathematics at the University of Toronto. Um Toronto is my hometown. So, uh also very exciting. Um but the thing that is most um special here is Daniel's an actual practicing mathematician and in addition he's been incredibly vocal about um his evolving views of AI in math.
01:22 And so um I feel like every if I just don't check in with you you know like in two weeks something you know different has been revealed and then you're you're um very kind of like what do you call it? Um so you >> I have a lot of opinions. >> You have a lot of opinions. Exactly. So I want to get into that. So, I mean, you know, one of the things um that I'm most interested in is not just like a discussion of how the capabilities of advance.
01:46 I feel like in math it's, you know, that that's definitely the the headline, etc., but also, um, you've been very thoughtful about how practicing mathematicians should respond. Um, and so that kind of gives us a chance and opportunity to talk about actually what is special about math, you know, like it's not just like, hey, AI has been really making progress here, but like delve into what actually mathemat mathematicians do.
02:06 Um and so maybe like with that um uh arc in line we can start with um what has been the most impressive results so far uh given all the recent progress um uh for you and then maybe yeah we'll kind of take it from there. >> Yeah. So okay so there have you know now been a lot of results some of them produced autonomously some produced like semi-autonomously some who's like where the AI contribution is just not at all clear um they're in a lot of different areas so anything I say is kind of you know I can only really
02:37 comment on things that I you know have some expertise on so it's quite possible that if you talk to a different mathematician you'll get like different answers here uh so my favorite uh like fully autonomous result by an AI so far is the solution to the Irish unit distance problem which I think was announced in midMay. Um, so at least what I liked about that is it seemed to me that like it was in some ways a little bit creative.
03:00 So uh I think some of the results uh we've seen have had kind of the flavor of like you know you kind of take some known techniques and apply them in maybe a clever way or you uh I don't know they have kind of been some kind of results I would characterize as like last mile like where some recent work was done like quite deep work done by a group of human mathematicians and then the AI kind of took the final step.
03:24 >> Yeah. Um but yeah with the sort of you know distance problem I I think it was something where the result was like a little unexpected. So first of all my sense was like that people working in the area thought it was true and then there was a counter example found but then also it like brought in some techniques from from another area. I think those techniques were like not especially uh you know like deep or new.
03:45 they were sort of classical ideas from the 60s but they were new to this area of studying um you know point configurations in the plane and so that was pretty cool and then afterwards we got to see like it was kind of fruitful so >> a bunch of mathematicians took those ideas and used them to find counter examples to a bunch of other interesting open questions so for example like the sum product conjecture over the real numbers >> um so that's at least what way I like to think about how cool >> result is like you look
04:13 at it post hawk and you see like oh well you know we're whatever new ideas that were introduced, if any, like kind of useful to do other things like did they improve our understanding of something and I think that's maybe so far the the the main example I know of of a of a result of that form. Yeah, I think it's really meaningful that you're commenting on this because in um you know that result came out to your point in May and there's been so much kind of like so many headlines uh so far and it's kind of probably hard
04:40 for somebody who's not a practicing mathematician to appreciate the differences in these headlines and so you kind of already started laying out sort of a tonomy of like what is different in that proof and so it' be kind of interesting maybe to use as that as kind of both an excuse to talk about where you sense the model differences are and what like math mathematicians actually do.
04:59 So in this case, I mean the most kind of maybe naive understanding of what mathematicians do is that we're pushing around symbols in a logical manner. And this is why, you know, RL is so successful at this because you can kind of both verify it somewhat cheaply um compared to other domains and then also you know because like the rules are quite um uh quite neible and so you can kind of like you know if you're if you're super human at that you might be good at math.
05:24 Um, but I think that of course betrays most of actually what is miss what is interesting about mathematics which is perhaps I think you said this as well but I think anybody who's tried to do math is sort of like it's about the understanding and sort of getting at truth remaining confused and developing intuitions and very little I mean the tool to do that stuff is of course in having really good very uh strong abilities to push out logical you know implic applications.
05:54 Um, but maybe if you can kind of talk to speak to like, you know, when you say it's most impressive and creative, like decoupling the um just the inhumane maybe feats of just like logical um implication from like where is it being creative? What is kind of um what is it helping in gender in in terms of like mathematical activity as well? >> Yeah. Okay.
06:19 So, first of all, I mean you you kind of characterized it as like inhuman in some way. I actually think the argument was very human like you know I'd love to judge or sorry open AI like released some chain of thought it was very recognizable. It was like kind of you know if if I tried to imagine like my chain of thought >> uh in trying to solve a problem like it might look kind of like that.
06:38 >> Yeah. >> Um you know we haven't seen the raw chain of thoughts maybe that >> maybe they cleansed it a little bit. Yeah. >> Yeah. Or maybe you know maybe the model like to swear a lot in the middle of the chain of thought and they clean that up or something. We don't know. But like at least the summary seemed prettyable. And I would say that's actually like kind of typical of uh most of the results that I've studied.
06:56 Like they don't seem inhuman at all. They seem absolutely like something a human mathematician could produce. Um and uh they're like typically understandable if like not so well written if you just like look at raw model output. It's not like there's some you know move 37 or whatever. Um it's like a human mathematician doing math. It's like a human mathematician doing certain types of math.
07:18 So like they're definitely the models are still like not very good at some mathematical activities. And here I don't mean like by field but just like certain things you do when you try to solve a problem the models don't seem to be doing but certain things they're very good at. So um the ways they might be a little bit inhuman is like they don't get tired.
07:37 They know a lot. Um but it if you have to just read the final output it doesn't seem kind of inhuman. >> It's interesting because I mean obviously we can't get too much of information from the labs who are producing these models of like why um you know how of what the training recipes are or how they're kind of advancing in reasoning. Um but at least one of the things we do know um and I think kind of open AI spearheaded this is this like reasoning in natural language is actually what they happen to scale up and it's
08:03 not actually pushing a lot of lean verified proofs as like the training corpus and that's kind of amazing. Um on another side, what's kind of interesting is it's not clear that a lot of the mathematical um training data, if they use that to a large extent at all, is reflective of how mathematicians think, if that's fair, because a lot of it is not like legible as traces, right?
08:25 Like most of the papers are crisp and like polished. The textbooks certainly are just show very little motivation of how something is developed, which is why it's like usually easier to kind of follow research direction by actually talking to the researchers and how they're thinking about it. So kind of curious before maybe even going to the texonomy as an excuse like when you're examining these uh models and their results comparing anthropic with open AI, do you detect any differences in in how that is similar to
08:53 human reasoning? And then also if you have any comments on like insights on perhaps why natural language scales so well that way even though it's >> okay. So so first of all I really like your point by the way that that they're mostly doing natural language reasoning rather than lean. Like I think that that suggests to me that like these you know you hear a lot of people say like math is a verifiable domain like that explains these the progress whatever like my sense is that because they're you know primarily scaling
09:18 in formal reasoning like probably the techniques are going to generalize to other domains pretty well. That's just my guess. >> Okay. Um, you asked a little bit about like Claude versus CHBT. >> Um, my sense is that they're pretty similar in terms of capabilities. I' I've played a lot around a lot more with CHGBT than than Claude Fable, but you know, I it seems like there's a lot of cases where, you know, uh, Open AI will drop a solution to some problem and anthropics will say, "Oh, you know, too."
09:46 Exactly. >> So, it actually seems like they're they're solving a very similar collection of problems and it's like kind of a small, you know, uh it's like a relatively small portion of what human mathematicians do. Um so yeah, one thing um that is interesting is like we see the solutions have a certain flavor, right? There'll be things that this isn't surprising like there'll be things that rely on the model strengths like their ability to like grind out along computation or like you know pull together kind of
10:14 technical ideas for many areas um or like many maybe many papers that you know a human my mathematician might not have read but they're not like they seem like weaker in things like intuition or like having some big picture point of view like a lot of what I do as a mathematician is like I have some kind of philosophy that's like very non-rigorous like maybe I think this thing is kind of analogous to thing and then like a lot of what I'm working out is like figuring out how to make that precise and like you know trying
10:40 to measure the extent to which I've succeeded in understanding that by like can I solve a problem or whatever like can I find an interesting phenomenon which I don't understand and then now I understand it and like so far um I haven't you know you can even the best results the models are producing you don't see um that much of this kind of reasoning it's more like they're very very good at applying some known techniques >> um which is to be clear that's like a a very powerful thing to do to be very good at applying
11:05 like all known techniques. >> Uh there are mathematicians who have had great careers uh doing very high quality work of that flavor and I think a lot of what the models are producing it's high quality in that way but it's it's like some kind of fairly narrow band of what mathematicians care about um so far um I do think there's like signs of both um you know all the frontier models starting to be able to do more fuzzy things um so like I've I've tried to get uh both both Fable and um Chad GPT 5.6 six, I guess, to do
11:38 some kind of theory building. And it's like they're not good. They definitely are not good at it autonomously, at least with like whatever scaffolding I've set up. But with some hints, you can kind of get them to do something interesting. Um, you know, when you do when you give the models hints, it's always a little hard to tell like what part is from the model and what part is from you.
11:56 But my experience is that like if they can do it with like a hundred bits of hints or whatever in six months, maybe they can do it without hints. Yeah, >> you know, I do think there's signs that they're also kind of somehow picking up some of this like implicit and unwritten mathematical knowledge. Okay. I would love to so go into the um the intuition part and where it um sucks at um to put it in a very basic way, but um you actually mentioned a small detail which is that you know from the public you've gleaned that
12:27 anthropic and open are probably neck and neck but you personally are using a lot more um chatbt like why is that or 5.6? >> Well, why is I don't know. I mean I just I think it's just in our show like I have >> pretty sure of it. um it got one thing is that chat GPD got better at math earlier. So >> like for a long time um the claw model models were just like not useful for research math and then I think maybe around opus 4.5 or opus 4.6 they like more or less caught up.
12:56 >> Yeah. Uh but you know some experimentation suggests to me they're pretty neck andneck and so for my own work you know except when I'm just experimenting I mostly just stick with one and sort of random >> as people outside of labs like us it's really interesting just to compare how they differ on the frontier and I mean you know to your point it might be a little bit of momentum I do think that at least from my anecdotal experience 5.6 six has been like a lot more clear in exposition and it just like there's a little
13:24 bit and this might not be true. I mean obviously the the models are just incredibly jagged at the frontier. Um but like it's in the explanations of uh results to me. I I always find that 5.6 is giving a more accurate theory of mind of what it assumes I know and don't know. whereas Fable might be explaining something very trivial but then just like jump at like well you know obviously you should know these things and >> I'm not I find they're both pretty bad at theory >> okay great so then see from in your eyes it's
13:55 you're probably asking much deeper questions um okay so um the thread that I really wanted to pull on was when you're talking about um the models maybe starting to get better intuitions or even theory building so maybe before we even dive into that it would be useful to kind of talk through like what is your primary ary you know activity as a mathematician um especially in your area of like algebraic geometry probably has a very different flavor than a combinatorialist or you know um some other areas.
14:21 So if you can give a little maybe brief lay of the land and then kind of explain what your what your ma ma mathematical activity was preai and then maybe how it's kind of like changing with AI. >> Yeah. So so I I mean I think there are a lot of different kinds of mathematicians. Um there's like a lot of different uh you know spectra on which one can put a mathematician.
14:40 Um so uh definitely a lot of mathematicians like solving open problems. >> Mhm. >> Uh and then I think I'm one of those like I like to solve an open problem. That's uh I think of myself as a problem solver as opposed to like one you know one other taxonomy you could you could have is a problem solver versus a theory builder. Mhm. >> Um at least for me, uh the point of an open problem is it's like supposed to measure your failure to understand something.
15:07 So it's kind of like a benchmark. >> Um right like uh you know one problem I really like is the gross decap curp whatever that is. It's it measures something about our our failure to understand differential equations. So it's like there's some very basic object we would like to understand. If we can't answer this conjecture we know we don't understand it.
15:28 >> Mhm. Okay. So in practice like how do you get at a problem which is supposed to be measuring something you don't understand? Well like of course you try to understand the thing better. Um and and in practice what that means is like well you try to find the smallest situation where you can't understand something and and and fiddle around with it and then you stare at like what once you win you stare at like what you develop to win and try to turn that into some theory.
15:50 Um so that's one thing you might do. You might try to like solve a problem and in so doing develop some kind of new understanding of the situation. Uh you might also just like have some feeling that like this thing is related to this other thing. Um and you know you might start building a table like oh this property a is related to property a prime property b is related to property b prime and so on so forth.
16:12 So uh for example uh in in my work a lot of it is motivated by some analogy between homology of algebraic varieties and representations of fundamental groups. Okay, so that's some some fancy stuff. Uh, but it's that this analogy is like very very fruitful and like really any phenomenon that appears on one side, you can find an analog on the other side.
16:33 And so you know trying to uh to realize that dream has led led to a lot of beautiful mathematics of the last like 30 or 40 years um by people like Carlos Simpson and and Takuro Machisuki and others. Um, and so here it's really just like someone noticed like here's an analogy and then that analogy has like led to a huge amount of developments. There's not really like an open problem at the end although of course like as you develop this you come up with lots of open problems.
16:59 There's just like a philosophy that you're trying to uh you know you're trying to realize and that philosophy is like super not rigorous actually. It's not like symbol pushing at all. >> Um yeah so that's another kind of activity that I like. Um yeah, beyond that uh you know a lot of uh what you do when you try to you know a lot of the activities you like you're actually trying to figure out what the right question is even like here you're you have some object you feel like you don't understand it and like figuring out
17:34 what you don't know is like actually a very challenging thing to do. So, so um you know there's like a list of you know you can go online and list open conjectures or whatever and this like doesn't really capture in a lot of ways what we don't know like often finding the conjecture is like really really hard. Uh so I don't know a good example of this is like the burton swerend dire uh conjecture which is one of millennium problems uh is this like beautiful relationship between l functions of elliptic curves and uh the
18:03 the set of solutions to the corresponding equations of the rank of the of the group of solutions. sort of what that means. Um it was discovered by like it was like the first big data conjecture. So Burton Dyer had like found all these statistics on elliptic curves in the 60s like this one of the first ever computer aated bits of mathematics and they like graphed these statistics and they noticed that you know some the slope of some line on a graph was related to some other algebraic variant they knew and that was the
18:29 source of this conjecture. So a lot of the time you're just like working out examples and like trying to like you're doing kind of science like you run an experiment and you try to figure out an explanation >> for that um experiment. And so I mean I think going back to AI I think like these are things where AI seems so far to you know help a lot more in some things than other things.
18:47 So like uh you know I think the the more vague uh uh a phenomenon is or like the the less you have a precise question in mind the less useful h happens to be. And so you were asking like how I use it in my daily life like actually what I found is that the projects that I have that kind of predate AI like the projects I've been thinking about for three or four or five years um it's just not that useful.
19:13 like it's primarily kind of a substitute for Google or something like I might use it to like learn about some related topic or like something where I would have earlier Googled something and then like read a paper like maybe I'll discuss it with AI instead. So it saves some time for sure. It's like not really doing deep intellectual work for me. But then you know because I you know I'm now like well I I suck at coding and so now I have you know my good friend who's really good at coding and now I have all these coding
19:36 projects because like suddenly well I if I had a question where coding would have been really useful I would have procrastinated on it for six months until I >> tireless PhD student. >> Exactly. Yeah. So, so yeah, now I picked up all these projects where yeah, coding is really useful. Like >> the models are like very good for kind of massively parallel things.
19:53 Like if you want to find an example of something, you can just ask it to find, you know, work through a thousand examples in parallel or maybe 10 examples at a time, you know, 10 different sub aents and like that's really useful. But these are like kind of different activities uh which are like I don't know on top of what I was doing before. >> Um >> I'd love to dig in.
20:13 >> So yeah, it's Yeah, go. >> Yeah. Sorry. cuz you were mentioning you there's projects that you've been working on for like three fi you know 3 four five years and um I don't know if it's correct to say like those are more of the theory building aspect of it cuz you did characterize yourself as like an open problem solver or like what is the thing that is because like the deep thinking part maybe maybe like you know kind of for this audience it might also be useful to kind of say or explain you know most of pure math
20:36 is it's not like it's being motivated by you know nothing on applied math but applied math is like at least there's some external motivation for why a certain formal structure might be interesting to study. Whereas this one it's like purely it seems almost sociological and to some extent I think like thirsten made some comment you know or point about this in the 70s that it is a sociological phenomenon of more and more mathematicians start examining something then you'll maybe you know converge on some interesting
21:02 structures but it's just like it's not it's there's some reason why people would prefer to study something or they think it's beautiful like what is it that drives maybe you in particular and then maybe you can make a more general comment about the profession. Yeah. I mean, so definitely some people are like motivated by like beauty or or or kind of aesthetic considerations.
21:21 I try not to be motivated by that. >> I think that's a bad way like Well, one thing like that sort of limits you, right? Like one >> thing kind of a failure mode I see among young mathematicians sometimes is like you you have something and you kind of think you know how to prove it >> and then like the proof feels really ugly and you decide but like okay I mean what if you're wrong and it's not ugly like you why limit yourself you should >> for this audience like what is ugly because I have an intuition of what's ugly
21:50 but like what is that for kind of spell >> out I I don't really know I I don't really have aesthetic feelings but people people sometimes feel this way like maybe it involves a lot of grinding calculation that's not illuminating or >> but you should like win by any means necessary in my opinion like I I like to think of what I'm doing like doing kind of physics except with concepts >> so um you know I instead of beauty I try to think about maybe like what is kind of fundamental what's going to open up further
22:17 understanding most >> will you know am I introducing a new idea that will be like broadly useful to understand this object >> and I guess you know in some sense there's there's like some aesthetic consideration there, but it's I try to I think the the orientation of like trying to do good science rather than trying to do art is what I prefer. Uh but there's a huge I mean there's a huge variety of of opinions here and like lots of mathematicians think of themselves as being like closer to poets or something.
22:45 >> Um yeah, what I what I do think is broadly true is that like progress comes from people like it seems to come in mathematics from people sort of pursuing their personal curiosity. Um and that it's sort of been crucial historically that there's like a lot of different people with different views of what's interesting and then the frontier of knowledge expands and expands and then suddenly you know you have these opportunistic situations where a new idea has been introduced and you can kind of now find you now can
23:10 suddenly like cascade through a bunch of other things that we didn't understand before. >> Yeah. I mean, so then going back to the three, four, five year um problems and where are the models sort of not useful? kind of ask it another way like when you're doing the deep thinking like is it just that it's just not clear that you formulate it as a um problem and it's more that you're thinking about these like what are you know what what are the fundamental kind of physics of you know >> yeah so in some cases I mean the
23:41 the some cases there is like a wellstated problem here like you know you know I said I've been thinking about things for se you know maybe 10 years now at this point for certain problems sometimes there is just like a conjecture >> that I would like to prove that that is is there. Um I think one reason the models might not be useful for some of these things is like the conjectures are true.
24:03 Um so like uh I think that certain you know for for example um you know this unit distance conject problem like the the general belief uh in the community was that it was true and then it turned out to be false. >> And so what that means is that there's like a specific construction you can do to >> to um uh to to refute it. Um on the other hand I think a lot of the things I I think about like I know okay maybe I'm about to you know that someone will come up with a counter example to the curvature conjecture hypothesis
24:31 tomorrow and like uh I'll look like a fool. Uh but that in general like these conjectures fit into some very broad theoretical framework which means that like we actually have a lot of evidence that they're true. Um and so uh you know yeah so that's part of it like there's not like a construction you can do to refute it. um you need to somehow kind of uh you know there's we have this giant framework where certain certain pieces of it are only conjecturable and you probably need a to resolve some of those conjectures to
25:01 win. >> Um we also have like a pretty good sense I think that like very serious new ideas are needed to resolve those conjectures. >> So like um of course you can't be sure like maybe there's some very clever construction that will let you I don't know uh avoid having a big new idea. >> Um I don't know. uh uh maybe you know it's quite possible we'll find this out but my sense is that for at least a lot of the things I've been thinking about um there are they're just not accessible to um you know applying known
25:32 techniques very very techn in a very technically strong way um so you need to develop a new technique just be clear like I'm not saying the models won't be able to do this it's just like so far they seem not to >> actually that that's exactly the point that I wanted to delve into because it's um you know to your point to your point is like okay we can try try try to calibrate and forecast like why they would get better at this but it is true like you know that just providing a construction for a counter example they
26:00 seem to be strong on if you have to start developing either new theory or to your point techniques machinery to solve to to prove why conjecture is true it struggles more and it's probably because a lot of what it's drawing on is also just like techniques that have happened in other areas and they're porting it over and and to your point that's why maybe the unit distance problem was such a more creative result because it was maybe doing more of that on its own.
26:23 >> It was like from another at least it was like drawing in something from unexpected area. >> Yeah. Exactly. Exactly. And so like maybe then to kind of ask the question it's not that because now we're like okay fine AI is getting so good so fast we can't count it out but like why what do you think it has to um yeah I guess it's it's you're spelling out what it has to get better at but maybe some more kind of feelings on like why it's kind of particularly hard to then develop that new theory and technology.
26:46 Um, yeah, it's a good question. I mean, I think you just need a different like my my guess actually is it's probably totally doable. Um, and it just hasn't been done yet. Like maybe you just need a different RL environment. Um, I don't know. Yeah. So, I at this point my expectation is just like that the trajectory will continue upwards. I'm not a skeptic of continued capabilities growth.
27:07 >> Um, yeah. You know, I think what is definitely true is that like the skill of like developing a theory um or like building your understanding of some poorly understood object is like a fuzzier one. Mhm. >> Uh so it might be harder you know uh I guess you can try you can tell it you know uh develop your understanding of zeta functions then once it proves through your hypothesis you give it a reward but it's like kind of harder I think to come up with some like intermediate no things that you can reward um >> yeah
27:37 that said you know I do think uh mathematics as a whole provides a lot of conjectures of varying levels of difficulty so you know maybe this explains why there seems be a little bit of progress in these areas like presumably they are trying to you know get it to solve lots of problems and some of those problems develop at least some of the skills of theory building and humans are able to develop these skills um you know I guess sometimes they get rewards from their PhD adviserss and their advisor says oh that's a good
28:06 idea or something based on some element of human taste or whatever >> uh and that that might be something one can do too but yeah and my my expectation is that just like as part of continued capabilities growth we'll see growth in these areas too >> yeah yeah And I like the framing where you're sort of casting these increasingly difficult conjectures as a as a you know a form of curricula for both humans and of course AI.
28:27 And it it you know maybe kind of getting a little bit too philosophical for some people's tastes. It's like it it it gets at the question of like why are we particularly good or maybe particularly bad at math too because it's um you know what is it that we're either struggling to do or or some people are particularly good at when you develop new theory because it's it's not again it's like not really or I maybe it's related to like why can we formulate good structures for physics as well.
29:00 That's just like it's not obvious all parts of the world are kind of understandable um and legible that way but some parts are and therefore we try to do it because there's maybe compressive pressures on our our you know our our minds because we need to uh we can't understand anything uh beyond you know compressing uh stuff uh more finely. Um, I don't know if that's also like the correct interpretation of >> I think that's that's I mean I I'm a little bit skeptical as like compression as a metric of interest, but it's
29:28 definitely like as a anthropological reason for like why we do a certain theory building. I just think it's true and in fact I think it's like kind of like our our inability to just grind is kind of important to our ability to make discoveries. So and just like as an example uh so I have I have one paper out so far where uh the models were kind of useful.
29:47 So they like like proved a couple lemas. >> Um so this is some some situation where like we had kind of proven the main result and then there were like some lemas I was unhappy with like they seemed not to be optimal and so I like worked with actually this was with um with Gemini uh deep think which at the time was also on the frontier no longer man.
30:08 >> Um I I where I kind of worked with it to to prove the lemas. And so what happened here was there was some lema um I wanted to prove uh and the models couldn't do it. So the no none of the frontier models could do it. And so I like worked out a ton of examples on my own and I realized oh well maybe like here's some reason why it could be true. Like so I like I found a better statement of the lema.
30:30 And then okay once I had that statement like okay probably I could have done it pretty fast but also the models were able to very quickly prove the pro prove that sort of better statement. Um, so like our inability to prove it like led to an improvement in the result. Okay, now you can take the original lemma I had that the models weren't able to prove, put it into chat GPD 5.6 Pro and it will output like the worst proof you've ever seen, like 10 pages of like just brutal calculation with no insight whatsoever.
30:54 And so, okay, this would have been a perfectly fine proof, but it would not have led to this discovery of like I think a kind of beautiful conceptual explanation for why this thing we discovered was true. So like we found a better proof because we couldn't do I mean okay I I say we couldn't do the calculation like what actually happened and here I I understand that I'm being hypocritical about my complaints about ugly proofs before is like I realized like this horrible grind proof would work and I like could not bring
31:20 myself to do it >> uh and so I like for another argument uh and so and it's not yeah but you know now the models can do for very long technical calculations pretty reliably. >> Yeah. Yeah. Well, I mean to your >> credit. I don't know. >> I think I understand why you're saying it's like don't, you know, just shy away from trying to prove something just because it seems ugly because it's like you have to take the first step, right?
31:42 And eventually you all try to work towards insight. And so, you know, you might be shy about admitting it, but it is still maybe driven by whether it's like aesthetics or just like pure, you know, like I want to understand. And if understanding just means it's a little bit simpler or, you know, more compressed, then then I've gotten understanding. I mean, it's just like a deep phil philosophical question, too, is like what it is, >> right?
32:01 Like I mean of course any you know there's a reason we do informal mathematics rather than like writing out long formal strings of symbols of CFC or whatever even though you know besides length it's just like somehow we're trying to put it in a way where we are actually getting some non-rigorous understanding out. Yeah. I mean, and >> I think somehow >> Yeah.
32:18 And even how that even relates to like why it helps. And again, it's like is is it because of our inability to grind or is it because like if we if we are pushing towards something that's more compressive, it also hopefully also ascends on the does it help understand other fields as well. And it it's just it is somewhat magical that when we try to optimize for two the the both that it it tends to coincide.
32:40 I don't know actually you know if that's a fair it's actually a quantitative statement but like why you know that is is is kind of magical >> it's at least sometimes true yeah >> yeah exactly or you know we we certainly bias towards the cases in which it is that's why we call that a good theory to build um um but it's it's uh yeah it's very much I think you know um it's like this is what we can study and understand and therefore it's also very convenient that it was rich um in in kind of mathematical uh results there.
33:09 Uh yeah, I mean okay so I think there's two things that we can like you know pull on which is like what um given you've seated that or not seated but you've believe that mathematical ability of AI is going to continue advancing and can do some of the things that you're you know doing nowadays. How sh how should like the mathematics community kind of best um adapt and benefit from this?
33:37 I mean as a just kind of saying as somebody who's you know doesn't have the time to kind of practice mathematics anymore this is kind of ve very great because I can maybe dabble more there's a lot of results that can come out but I can also see where you know you've made the more precise point of we can not motivate the right kind of behavior of understanding and development and so I would love to kind of hear more about um your views there.
34:05 >> Yeah. Yeah. So, okay. So, first of all, like it is clearly really exciting like that there are increasingly capable models that are like producing high quality results. Um, you at least, you know, some high quality results also also a lot of slot. >> Uh, yeah, some some good stuff. Um, and so, you know, as the models get really capable, you know, my hope is that they will answer a lot of the questions that I've been like, you know, kept up at night thinking about and like I I'll get to learn the answers.
34:30 I think that's really exciting. And you know, a lot of people got into math uh like largely because they enjoyed learning math. >> Yeah. >> U like the first thing you do as a a math student is you like learn stuff that other people did and you do that for like 20 years before you before you start doing maybe not quite 20 years but >> 15 years >> if you're lucky 20 years.
34:50 Yeah. If you started early. Yeah. >> Yeah. Yeah. Um yeah. So um that said like the goal of mathematics is not to produce mathematics papers um like it's to produce some kind of understanding uh so maybe okay maybe that some of that understanding resides in model weights or something to me that's like pretty unsatisfying like >> uh my own you know my motivation for doing mathematics is like I would like to satisfy my own personal curiosity um I think people should be uh have the capability to do that and like that
35:23 requires a pretty substantial apparatus like it's like simply not the case that you can like study the questions I think are fundamental unless you've invested a huge amount of time and effort um kind of getting to the point where you can meaningfully do so >> and then moreover like that you know even the people like you know you have the small group of people doing like fancy research math on the frontier or whatever um that relies on like a huge apparatus of like you know thousands and millions or billions of people
35:50 who are trying to learn to think mathematically like you need a an entire mathematical community to support a small group of people who are on the frontier like you just without the pipeline then the pipeline doesn't exist. >> Uh so if you think that's important like development of human capitals that can like meaningfully engage with frontier mathematics I think that uh you know you still have to incentivize those people to actually invest their time and effort getting to the point where they can engage and then do so
36:15 in a in a high quality meaningful way. Uh so right now I think like the existing incentive structures for math research do not uh do not encourage people to do that. Um so you know right now if you're like a posttock on the market you want to get a job maybe for the next couple years before the community adapts uh the best way to do that is like you want to produce a lot of papers which maybe prove you know old conjectures or whatever and you can do that by playing the slot machine until >> um the model produces a
36:43 hopefully correct proof of such a result. Um, so you don't even have to pick the the the theorem in advance. So like here's an experiment you can do. You can take codecs, you can say go online, find five recent conjectures in algebraic geometry and prove them. And uh, okay, I've run this experiment and with some back and forth, I was able to, you know, in an hour get like three, you know, quite bad papers, >> but correct papers.
37:07 Um, which, okay, are now sitting on my hard drive waiting for me to email the relevant people, but this is not like a good use of my time to invest results. But yeah, you you definitely see people doing this. Um, so there's, you know, been a huge uptick in post to archive, mostly not very interesting. Um, some of it is interesting. >> Um, but, you know, a lot of it is kind of clearly low quality, even if in so far as the like even if the result is something that would have been like kind of highly rewarded a year ago,
37:37 just like the there's sort of no evidence that a human being is actually engaged with it. Like there's no development of human capital or understanding. Like sometimes, you know, we've seen examples where like three or four or five papers with the exact same proof of the exact same theorem have come out in within a couple days of each other, which is clearly, you know, some situation where someone's playing the slot machine.
37:57 Yeah. >> Chat GPT is kind of consistently finding the same. >> But that's also interesting. It's like not it's kind of mode collapsed on like certain paths of reasoning. >> Yeah. >> Yeah. And I think it's like not obvious that this problem goes away as the models get better. Like maybe it does, but maybe it doesn't. Like I think a lot of what I was saying earlier is that a lot of progress in mathematics comes from like letting you know a thousand different flowers bloom and people pursue their own curiosity and then
38:19 you know the boundaries of knowledge expand in some kind of fairly hopefully fairly uniform way and like this really highdimensional space of mathematics >> and then like opportunistically you you suddenly get some applications or like answers to old questions we found. Uh and it's like not clear to me that if if you know we kind of subordinate mathematical exploration to what the model want to pursue like if what you're getting is like one mathematician duplicated a thousand times or like actually you know a million
38:47 different mathematicians doing a million different things. You know, I think actually that's pretty I mean, you know, aside from like okay, math has to adapt um by changing incentive structures, I think this is actually a pretty um um maybe dangerous or something that the labs have to pay attention to because you know on one side what's been successful for them is that this emergent reasoning capability is obviously incredibly powerful and we've been I mean it's creating a lot of like great PR headlines.
39:15 But >> you know to your point and also like where my interests tend to is that like human mathematicians they are coming from all sorts of weird intuitions that like you know and that is why you end up developing you know to your point the frontier that is so diverse that you actually can draw from and then be able to connect and produce a lot more.
39:34 And so it if it's true that most of these like proofs that are being pushed out by the labs are converging on very similar things because they are technically drawing on the same body of literature. Um and that's where they're strong at now. It's it's not clear that I mean one you know where are they developing their intuitions? It's from practicing mathematicians.
39:55 It's not clear where that other emergent you know stronger diverse intuition might be coming from. It might come but I don't have a good theory for where it comes from. And it's also just not clear that you know just with like test time compute and various post-training scaling that you can even I mean it's a huge debate whether you could introduce new capabilities and it's like precisely I think it should be studied in the math context because like where it's good at and where it fails um and where human
40:20 mathematicians are good is exactly a very kind of like precise um question to study it on. Um, so this is like a long way of saying that I it is unclear to me that we'll maybe produce um if we don't incentivize enough people to interact with it, maybe unlike other domains, it it might not actually uh continue producing uh a superior uh uh result uh without human aid.
40:47 >> Yeah, let me actually push push a bit further even. So like let's suppose the models become like really robustly superhuman like even like we're not even adding like meaningful cognitive capacity. Okay. >> I claim like still actually we we still won. >> Okay. Great. >> Human mathematician. So, right. So, why? So, like it's because the you know there's like a it's like there's a question about how we've we're designing society, right?
41:07 Like maybe uh >> the optimal uh situation is like you have the models doing all sorts of math research and like that leads to very in you know in this like highly you know whatever non-uniform diverse way like maybe you don't need humans to add that kind of diversity. It's like maybe that's the optimal thing to do and like in the end that leads to lots of applications and lots of imp proofed understanding and so on.
41:30 It's like despite like something being optimal doesn't mean you do it >> right like there's no reason to think that you know if we hand over control of whatever math research to the models and just let them do their own thing that it will do the optimal thing. Um and in fact like if you you know if we are if we kind of instrumentalize what we want them to do like we want to say like oh you know make our life better or whatever >> it might not be the case like be that that like what they decide to do that is like is is
41:57 pursue a wide variety of interesting research right like they might just try to take the direct path we don't know what's going to happen >> um so I don't know if you believe that there is any value in sort of this broad-based like uh fundamental research which I do look I think that's like one of the most valuable as human humans or or whatever or the models can be doing like I think the easiest way to guarantee it happens is like to keep a community with like broad interests who are like pushing the models to do and
42:23 like helping us to design a society where like that's what we're what we're pushing for. >> Yeah. Like I think at least you know my my hopeful vision for the future is that like humans are not like totally disempowered like we have some control over where we're going and like if that's the case like what we end up doing is going to be driven by human interests >> and so you want to have people who have lots of different interests and also the capabilities to actually pursue them like you want people who are like smart
42:46 and engaged and like you know well can do not just like mathematical thinking but all sorts of thinking. >> Yeah. Yeah. I mean, yeah, a great fear is like as AI advances, we don't develop the right ergonomic um kind of interfaces to actually encourage us to also be good continue to be good thinkers and and it's so easy to kind of um relinquish that because you just off and and the models aren't even good at that level of like you know thinking where it's like the top level structure but despite that it's it's so easy
43:11 to and so especially for math I mean just like you know another selfish reason is um if you kind of math max you know I think it's actually a great pedagogical um uh excuse to actually get just really rigorous at thinking about various things. I mean this is not why mathematicians do it. Um but as just somebody that's part of what we do >> Okay, great.
43:31 Because it was kind of, you know, I thought it just really helped give me a very um good framework to think about many things, not just mathematics. And as somebody who's also, you know, now a parent of a 2-year-old, I kind of think about this a lot, too. you know, it's it's not about kind of grinding, even though whatever it's still good, but like not to not to um on grinding too much, but you know, just hearing I I used to collaborate a bunch with some Hungarian mathematicians, um Bery among them, and I I heard that
44:02 um in in Budapest, they would just teach group theory when you're in primary school. And I'm like, well, we should definitely do that. We should continue doing that. now that AI is so good at, you know, some somewhat good at explaining, but it's far more accessible, we should actually, you know, probably proliferate uh that even more. Um, and so maybe that helps with bring more people to the frontier rather than, you know, just >> Yeah.
44:25 I mean, this is something I'm concerned about, right? And of course I mostly talk about math because that's like where I live. But like I think you know one nice thing about thinking about this is that we're one of the first uh professions uh to to kind of you know have a significant impact of high quality models. Although I think um maybe we're one of the first professions for it to happen so publicly.
44:45 You know, my sense is that there are plenty of other professions that are >> 100% >> coding. But also just like I mean I think that you know anything you do at a computer like probably a huge amount is being done by the models at this point and then like there not having a public reckoning about it but >> um you know because uh you math capabilities are are useful for the company for the labs to talk about I think probably a bit more publicly than everyone else.
45:08 Um but yeah, I mean I think uh you know one reason to try to maintain like human capital in this area is just like it's a model for all professions like presumably we still want people who are like meaningfully engaging with the world and like experts and you know have like talents and trained skills and so on and like um yeah so at least yeah it's actually very convenient that the math profession is so entwined with education here because like I think we're also seeing like you know some amount of uh challenges um you
45:38 know among uh college and and uh like secondary education coming from AI too. >> So as you said I mean it's also an amazing tool to learn. So >> you know people have have I I've heard he heard people start talking about a biodal distribution in their classes where there's some people who are really like figuring out how to take advantage of new tools and other people who are just like letting them do their homework and then bombing everything else.
46:01 >> Unfortunately I don't think that adapts fast enough but it's like we definitely want to be living in a world where we're producing better thinkers. I think it's just, >> you know, even if we talk about just the pure kind of optimization game, I think that's better for us. And so um but as you know, just like a human being, I'm like that would be pretty >> um inconvenient if we became worse thinkers just as AI ascends.
46:20 And it's too easy to let that happen. So we should kind of be thinking hard on how to actually take advantage of this and harness it for our own improvement, right? Well, >> yeah. I think it's like it's sort of interesting to observe like the models um at their current level of capabilities let you do a lot more. They let you do a lot of things you wouldn't have done more cheaply than uh you know cheaply enough to do them now.
46:42 But it's not clear to me that like in many cases they're actually improving the quality of outputs. >> Yeah. Um, and I think this is common like you have a new technology that's doing something a little bit worse than was previously done but much cheaper and so you get a lot of suddenly a lot of like lowquality outputs that are displacing previous high quality outputs.
47:01 But I think it's possible to use the tools in a way that actually like improves you know the quality among along every dimension. It just refers requires some thoughtfulness and some some redesign of institutions to actually incentivize that. >> Yeah. Well hopefully capitalism works there. I do feel like that the most high value things do require people to use it effectively and and right now the models are not good enough without like the human experts to actually participate.
47:26 Um but to your point there's a vast majority of maybe like more junior and entry level um kind of you know it's it's harder for um for those uh roles to adapt as well. And so the thing that would be a mistake is to use the AI models in a way that doesn't >> basically you need to be ascending and using the models to deepen your understanding. And it's so easy for human nature just to be lazy and you have to resist that because that is the moment that you will kind of kind of lose basically.
47:55 And so you kind of you know kind of forgive the very competitive language but it really is that like it's just so easy to to kind of relinquish the thinking to the models. The models can't really think. And so as things are ascending so fast, it's like critical that um you continue developing those facilities and actually leverage it to to to improve those facilities rather than relinquish.
48:15 Yeah. >> Yeah. Yeah. I mean, one one thing I I think has been nice about, you know, sort of this vast increase in in like semi-expert attention or like model attention on math problems is like now there's been, you know, okay, I've been complaining about slot papers or whatever like, you know, people who are not um you know, producing really high quality stuff.
48:35 often that's actually coming from professionals like it's not I'm not saying like like you know there people who you know like within academic mathematics there are incentives to produce like a lot of a lot of stuff and and that's you know that's one place the slop come from so there's definitely also like slop coming from non-experts but that I I kind of actually don't see as a net negative like okay there's a lot of now there's a lot of like documents on the internet one might have to comb through to figure out if a
49:03 problem has been solved or not Um, but to me it seems like just the fact that there's lots of people excited about math is like kind of a positive. So that's like a nice like >> I totally agree. I know I get to talk >> and not even kind of a positive obviously. >> Yeah. Yeah. No, exactly. It's like suddenly there's a spotlight on it and I can um nerd out about math more.
49:18 Um I I was actually kind of curious um you've got any comments on the um >> I guess this is another constructive result, but the elliptic curve of rank 30 that just came out yesterday. So if that >> That's right. >> Yeah. Yeah. >> Not Well, I mean, we don't have any details about it yet. >> I know. There's rumors. >> Yeah, we don't know how it how it I mean, so it was it's due to uh I guess Claude Fable um prompted by Levent Poge and um right >> a collaborator whose name I unfortunately forget.
49:47 Maybe you can >> add it in post. Um >> I just know the Twitter handle. >> Yeah, we can. >> So yeah, I mean uh so a lot of these nice recent results have have come from Levent with unclear amounts of autonomy. Um so my sense is that some of them are semi-autonomous rather than fully autonomous. >> Um yeah with this we don't know anything about the methods.
50:07 Uh so yeah this is a fun construction. Um without knowing about the methods it's very hard to say how significant it is. >> Yeah. >> What I would say is that if you want to understand um kind of historically such results have been proved or sorry have been have been understood by the community they're like cool but I wouldn't say they're like a big deal.
50:27 Um so like the typical place a result like this might go is like someone's website of records >> like not >> it's not an anals level result. >> Yeah it's not an but that said it's cool and it's like a you know it definitely there were a few very very are a few very very talented mathematicians who like these kinds of questions and no males being maybe the most famous example.
50:47 So Alis and Clagun are the ones who kind of have been pushing this record for a while and they recently found a rank 29 >> example which was you know the previous record. >> Yeah. you know, so those are mathematicians. Um, and I think, you know, people like it and it's cool now that the models can do this sort of thing. Um, but yeah, it's very, you know, one thing I always say about a model result is like you cannot evaluate it except in retrospect.
51:10 And like this is also true of human mathematics. Like >> sometimes a problem we thought was really important or would require really deep new ideas does not >> and sometimes it does and you can't really know in advance. So it's always exciting when a problem like that gets solved, but then like to figure out how significant it is, you still have to look and like well Levent and Claude uh and I think Ava Howell maybe is the third collaborator.
51:31 >> Okay. >> Have not yet told us how they did it. So uh >> have they been mostly more secretive? I I think they've released some traces for stuff. But >> yeah, so for this one, I think they haven't yet, unless I missed it. Yeah, they they they you know, Levent likes to tweet out his uh his results, but yeah, he has been, you know, slowly releasing some kind of PDF writeups, too.
51:49 I think it's, you know, he's just having fun on on uh fun on the internet. >> Yeah. Yeah. >> Um it's it reminds me of when you're saying, you know, you have to evaluate how the results came. Um so um I had um Mark Zel Zelki and um uh Metab Swani on from OpenAI recently and um they were saying how you know what's kind of been the most charming or delightful is just that the proofs have been relatively short.
52:17 They're not like 200page and you know maybe corresponding to your grinding point as well. Um, but maybe this is kind of optimized for in retrospect it was picked that it was short or maybe do you find that on average uh stuff that you throw at um GPT, you know, Saul or Fable tends to be shorter and more of legible to the human uh or versus it might just go haywire and just grind it out.
52:45 >> Yeah. So, so I mean I I think it is nice that they'll sometimes produce short clever proofs. Yeah. Um so that's of course everyone likes a short clutter proof. Um I think my sense is that the uh reason they're not producing long complicated proofs is that they cannot >> um like that just like that the ability to check correctness is not yet there.
53:07 Um so uh you know I mean even actually you know if you ask the models to produce uh >> short proof you you can then often ask them to you just ask is that correct and they will often say no. >> So like they're much more reliable than they were six months ago for sure but like it's still uh you know they will still sometimes produce things that are just wrong and they know they're wrong.
53:29 >> Yes. >> Um I think the problem with producing a very long thing is they might not know they're wrong. Um, and so what I wonder if uh, you know, presumably internally OpenAI and and Anthropic have probably solved a lot more problems than they've released. >> And I imagine quite a few of them are they're just not sure if they're true. >> Um, so for example, uh, you know, with these this recent list of 10 problems uh, released by OpenAI, those were all formalized in lean, >> which is of course a very good good evidence
53:56 that they're true. I have no doubt that there were a lot more that they could not formalize in late like because the prerequisite options have not been put into mass yet for example. >> Um there were probably more that were longer which also makes it challenging to check. >> Y >> um so yeah this is my guess and we do actually see very long generated um proofs on the archive.
54:15 Um so for example someone recently posted a claimed proof of resolution of singularities and positive characteristic which was 800 AI generated pages. Um it's like definitely I mean I'm sorry I haven't read it I haven't got an error but there's no way it's correct like this will be a >> major result. It's just not within the capacities of the current models if you're reasonably well calibrated.
54:36 >> Yeah. Yeah. >> Um and it's definitely no human has read it. Definitely the models are not able to check this kind of thing. Yeah. >> Um so well I mean um yeah. So this is my my expectation is that the reason it's producing short clever things is just like that's what we can check. >> Yep. and you can get it to produce long things that are grindy or hard to check.
54:55 But then >> yeah, yeah, we'll have to get there. This is actually more what it says about capabilities, frontier capabilities, rather than in in a perhaps more negative way rather than, hey, it's just, you know, so good at these clever. No, I I think that's totally fair. I mean, especially if you look at an adjacent domain like code, right? Um remember reading something um that cursor put out about testing their long horizon um you know, harness.
55:16 Um in this case they're trying to reproduce SQL light in Rust and it was just so telling like how far we are from you know and that and that seems like a very you know comparable task of like it just it's very long you have to make sure there there's something to verify that it's a correct implementation by testing a suite of you know cases um but you know you actually need a harness there in this case it's not just like the raw models and uh it takes a while and you can actually compare with different frontier versus
55:46 um you know non-frontier models, who's the planner, etc. like differences in in their capabilities as well. >> Yeah. I also think in practice like the uh to elicit a long proof, you kind of need a harness. >> Yeah. >> And um when you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output.
56:09 >> Um so like TEDGPD 5.6 6 Pro is like very it like really tries not to say wrong stuff for example and although you know it happens and then you'll ask it like oh was that correct and it'll say no. Um but uh you know when you're trying to get it to ideulate like you try to get it to be creative. You try to get it out of this very like um you know uh rigorous rut in order to like you know actually get it somewhere.
56:33 And I think if your interest is in I mean there are a lot of people who are uh you know trying to arbitrage the prestige mechanics of academic mathematics and trying to elicit a lot of proofs >> which are not necessarily being checked and in order to do that I think you just decrease the reliability in order to get a lot of stuff. Um, so you know, I think any harness that can elicit a 250 page paper is probably not being very careful about what it's producing.
56:56 >> And and you're saying that it's decreasing or it's not reliable just because it's the capability isn't there yet and so it's just kind of forcing a long longer horizon task on it. >> Yeah. >> Yeah. I mean in practice like how does like a human checking it like can't check a 250 page paper either like like you you you know you can't reliably check it line by line.
57:13 What you try to do to understand it is you try to like understand the overall global structure of the argument and you like stress test it in various ways like like would this argument apply something else that I know to be false like oh like does it work in this special case blah blah blah and also the models seem not to be able to do that kind of like more >> I don't know fuzzy like unit testing a proof very well yet.
57:33 So actually like one of my favorite tests for the models which they haven't haven't succeeded at yet is there's a paper um I won't name it that came out a couple years ago uh that was wrong um and it was like very hard to find the like in close enough area to myself that I like immediately like was claiming some big result like I immediately downloaded and like started reading it and it was like very hard to find the actual specific error >> but it was also very clear from the structure of the argument that it couldn't
58:01 work. So it was like, you know, me and a bunch of other arguing something that was too strong to be true if you kind of took that a little bit further. >> Yeah. >> Um so me and a bunch of other experts like immediately like realized it was wrong and like we you know we emailed the the author and like we kind of went back and forth until someone um figured out what the precise specific error was.
58:21 >> Okay. >> Um and so far the models seem not to have been able to do this. like the specific error is quite subtle and but this they're also not able to do this kind of overall uh kind of big picture tax checking which is kind of like how people in practice check papers. >> Yeah. Yeah. Which is kind of it mirrors you know why in code it's like it's it's so clear it's good at the syntax but you know higher level architectural stuff still very weak.
58:45 Maybe it'll ascend there but probably need some harness help. Who knows? I mean, people have, you know, evolving opinions on how much the harness and the model have to co-evolve and which one is necessary, but the next model requires less. So, I I feel like in in math, it it'd be very interesting to see how you um uh if you do any experiments with the harness there and how how that improves um because it is a mark of general reasoning.
59:07 Yeah. Yeah. >> Yeah. I mean I I do you know I have my own sort of bad little harness and codeex and codecs but yeah it's um but you know I personally I do not enjoy uh autonomous mathematics very much so I mostly mostly do not use the harness I mostly try to use it to help me understand stuff. >> Okay. No it's totally fair. Yeah. Exactly. You don't want to automate your job away because that does involve you being in the loop to understand it which is you know necessary to participate.
59:32 Um maybe to to finish off, I'd love to and this this could be just something you haven't thought about or actually thought a lot about. Um I think you also have a toddler, >> right? Yeah. >> Yeah. Yeah. And so how have you a three-year-old? Yeah. >> Yeah. Three-y old. Great. Right. So you're one year more advanced and probably you know have more thoughts on this.
59:51 Like how are you thinking about um is it her education or his education? >> Her. Yeah. How are you thinking about her education in in math and >> you know not to grind but you know really just like how you know pass on the love of it and and how to how to react to AI. >> Yeah. So she's she's three. She's she's never used uh AI. Um she she is starting to add that's that's about as far as we are in math.
01:00:14 >> That's that's more than she can add single single digit numbers like you know by counting on her fingers and >> count up to like maybe 30 reliably and 50 semi-reliably. So I'm very proud of that. >> That's good. Yeah. >> Yeah, I definitely encourage that. We talk about shapes and stuff. Um, uh, la actually a couple days ago I woke her up and she was like hiding under the blankets >> and I was like, "Oh, what are you up to under there, Sophia?"
01:00:37 And she was, "Oh, I'm doing some math." >> Oh, I did. I think you tweeted about that. That was horrible. >> Yeah, it was great. So, you know, I think she has like some sense that I like math and she's into it because of that. >> Um, yeah. I mean, I I would say I uh don't know, you know, I think the world is probably going to look pretty different in, you know, 20 years or whenever she's kind of uh you know, fully adult and and doing her own thing.
01:01:02 Um but you know, I think a lot of what we educate people for is like pretty robust changes in the nature of the world. like uh I think the reason to learn math has always been like to think clearly and like better understand the world and like presumably that's something you want to do even if uh there are sort of extremely capable AIs and this is also true you know I I personally like math a lot but also the reason to like read a lot of books and do the humanities and so on.
01:01:24 So I I mean I think like the actual values of uh the math profession and like education more broadly are like things we definitely try to want to try to preserve like that I I hope to instill in in in my daughter. Um you know how much institutions have to change to make sure that happens is is maybe an open question I think a lot. Um but yeah I mean at least at a personal level I'm I'm definitely trying to you know uh convince my my three-year-old that math is super cool.
01:01:54 Oh yeah, 100%. >> One of her first words was icosahedron. So >> was was what? >> She she has a she my my parents gave her a little iicosahedron toy >> uh when she was one. She she >> Yeah. Not literally a first word, but she learned the platonic solids quite early. >> Oh, very good. Very good. Very fun. >> Next up, group theory. I mean, it's like very natural.
01:02:14 That's right. Yeah. I mean, it's >> Yeah. I'm actually when I when I teach her about uh uh addition and subtraction further, we'll we'll do it in the context of a general group. >> Oh, very good. Well, at least you give the motivation. I think a lot of probably skip that part. Um, yeah. Yeah. And you know, maybe this is how uh math grad students can also, you know, focus like training the next much younger generation to use AI in in service of actually getting better at math rather than >> just lacking understanding.
01:02:44 >> Wonderful. Okay. Well, thank you so much, Daniel. Um, >> thank you. It was a lot of fun. >> Yeah. Yeah, it was a lot of fun. And yeah, I mean, I think there's going to be a lot more progress very soon. I'd love to maybe catch up and chat again. >> Sounds great.