Build a digital experience that combines pre-match forecasts, live point-by-point probabilities, key-moment explanations, conversational match questions, personalized player information, and biomechanics-based performance metrics to help fans understand live sports.
Confluent is an event-streaming platform used in the described system to carry live event messages from a win-probability system to downstream processing. That processing determines whether an event is a key moment and generates explanatory text.
Lean 4 is a programming language and theorem prover used to formalize mathematical proofs, including proofs generated by AI systems. Its project repository provides the language implementation along with theorem-proving and functional-programming tutorials, reference documentation, examples, installation guidance, and instructions for building from source.
Mixture of Experts is a weekly news podcast produced by IBM and hosted on the IBM Think podcast pages that recaps trends and innovations in the artificial intelligence industry. It publishes audio episodes that discuss recent AI research and developments at the frontiers of the field.
Searchable transcript of OpenAI talks GPT-6 Astra and Millenium Prize, researchers create WeWorm exploit & IBM’s US Open app — IBM Technology (33:07). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 I have yet to see a use case where my customers would have needed a model of this magnitude. All that and more on today's Mixture of Experts. I'm Tim Hwang and welcome to Mixture of Experts. Each week MoE brings together panel of brilliant brains working at the frontiers of artificial intelligence to lead you through the week's news. On this week's episode, we've got Olivia Buzek, staff AI engineer Aaron Baughman, IBM fellow and Bri Kopecki, AI customer Success engineer.
00:35 And today my co-host is David Zax, staff writer for IBM. Think we're going to cover three big stories today. We're going to talk a little bit about IBM at the US Open. We'll talk about some worrying news about WeWorm. But I think I wanted to start today by talking a little bit about all of this news we're seeing coming out of OpenAI. There's a new model, some major advancements in mathematics, apparently.
01:01 David, do you want to talk a little bit about The model is it's the the big new model from OpenAI. And if you've at a minimum seen the demo. These demos are becoming a staple of product announcements now. And often people go into a room and they begin talking with the model. And mainly I walk away thinking, I really want a big kind of AI cave to kind of, you know, talk to a model in this way.
01:24 But we've seen in this video incredible developments in terms of how they can turn 2D renderings into 3D renderings and simultaneously order UberEats and get that delivered, apparently instantaneously. So just before we turn to the mathematical side of things, I would just be curious to hear from our esteemed panel on what they are excited about in Astra.
01:47 Whether this, as the NVIDIA CEO has suggested, is truly the AGI that we've been waiting for. Whether it's a step change, what do you think? Olivia, for instance. So every time I look at things like this, like, you know, big new model coming out, I think I think a lot about the computational power that's going into it. So there is yes, maybe it can do those things, but I'm wondering just how many essentially tokens are being used when it's doing something like controlling your computer.
02:22 Like, is this really like the way that it makes sense for agent models to act in the world? And I think there's a lot of kind of low hanging fruit that we have to do in terms of making agents actually accessible to the world. Sorry. Making the world accessible to two agents. Before we get overly excited about, well, it can solve an arbitrary problem.
02:46 I can also order DoorDash. I could probably write a harness to convince a lower tier model to order DoorDash. For me, like I'm in some ways it's impressive, and in some ways I just wonder if we're solving problems the right way. And so I think what worries me the most is like, are we approaching problems or the right way as humans, or are we simply throwing bigger and bigger models at problems without thinking through what the actual implications are?
03:19 So it's a cool demonstration. I just wonder whether it has the right real world implications. Aaron, Bri, have you ever tried the models or seen tried Astra, or seen a use case that you think is really novel or can can be performed more effectively than by this model than other models? So first I want to scratch that. It's a bit about is this really AGI or not.
03:43 Right. Because I mean, first of all, I think that we need to understand in the community there's no converged definition of what AGI really and truly is. Right. If we define AGI as a system that can take a novel goal, it can understand the problem. It can even go out and get missing knowledge formally. That said plan, use tools, software, execute work, and so on.
04:04 Maybe that's AGI, but we may not even recognize when AGI has arrived. It just might look like a really good digital employee. You know, just like one of us doing work. But I do hope we're at the beginning of AGI because the reason is at what cost, right? I mean, if this number is right, you know, it's it took 100,000 Blackwell NVLink 72 systems just to train this model.
04:28 Right. That's that's on the order of enough power for an entire region and the continental United States, right? It's it's just unbelievable. Right. How how many resources it took. Now, the other point I wanted to make is that, you know, the real AGI test. Should it be a benchmark? Right. That's known. It should be a system or a goal that it hasn't seen before to see what happens.
04:51 And what we're all looking at here is to bench CAD benchmark. Right. Astra astounded 95.9% accuracy level on it. And the next highest one of course is the GPT 5.6 SOTA, which is what in the lower what 80s and Fable roughly around this same place, you know. So it did really good on that benchmark, right. And that benchmark is allowing the model. So so I think circling back to your question, you know, one task that is really interesting is this one, right, is given the model multiple 2D views, right.
05:27 And having it generate executable CAD code and understanding that environment, building a 3D object, comparing it with real geometry. It's pretty neat that it can do those types of things, you know, and we're moving from being able to ask AI for answers to now asking AI to do work. Yeah. And let me add my my perspective, which is very customer centric.
05:49 I have yet to see a use case or as of today or as of this, this or last week where my customers would have needed a model We need to solve the Navier-Stokes right now. The things that we're solving on a day to day basis for, you know, for enterprises, it doesn't really require a model of that nature. I mean, could it make it somewhat faster? Yes. And speed is very critical and an enterprise setting.
06:22 But I think I like what Olivia said, which we have to ask ourselves, what are we using this for? What problems are we really solving? But it's great that we have the capacity. So I think it's great. We're pushing for for those technological advances, be it computation. How many tokens are we burning? But then of course, we have to ask ourselves, what are we solving?
06:49 What can we what is the most efficient way to do it? But I have yet to throw this into a use case. I'm sure it's probably going to be requested, and I'm going to play around with it at some point. But it's it's exciting to see what's happening. I'm a huge NVIDIA fan. I think they have they have the power and they're they keep pushing their own boundaries, just like most other companies these days.
07:15 yeah. What are the interesting aspects of the Navier-Stokes area? Isn't that AI solved Navier-Stokes. But it's, you know, and created the proof. Right. But it's also that it used 10,000 AI agents to explore the problem. Right. So it's this community of agents that work together with the coordinator, almost like sub agents that go out and respond. And they can take an MCP server, get a tool right, and run all these different explorations with chain of thought, right.
07:42 But it's but it's pretty cool, right, to see all of that work together, almost like a swarm of agents, right, that are now being unleashed on Maybe let's pause there just to, if you don't mind, Aaron, to zoom out and just make sure viewers are up to speed on what we're talking about. This Navier-Stokes, Navier-Stokes. I've seen it pronounced any number of ways, but it is this so-called millennium problem in mathematics.
08:06 It's in 2000, the Clay Institute said these seven problems in mathematics are like the new Mount Olympus. In mathematics, we should all focus on them. And if you win, if you solve them, we will pay you $1 million. Only one had been solved in the past 25 years, and the second one was was solved. It's called Navier-Stokes. It has to do with fluid dynamics.
08:32 I don't understand it. I call it the exploding water problem because apparently they've proven that water can spontaneously explode, which sounds awesome, but basically Astra solved it. There is some soap involved. There's kind of a soap opera because actually some humans at NYU Anthropic had maybe simultaneously solved it, and it was only because of a rumor that they were about to publish that, that OpenAI began to throw these 10,000 agents at the problem.
09:02 But yeah, basically, Aaron, since you were kind of on this tack, do you want to tell us a little bit more about what the problem is, how it was solved, the amount of I gather it was sort of tens of millions of dollars worth of compute was thrown at this problem to win a $1 million prize. But what else should we know about Navier-Stokes? Before we get into the weeds on sort of how it was solved?
09:24 And maybe if we have time, the soap Yeah. So I mean, like you said, this is one of the seven Millennium Prize problems, and they have been open for 90 years, you know, so it's 90 years, almost a century. This has not been solved. Right. And what it's looking for, you know, it's it's looking at when a fluid velocity becomes unbounded, even though the system starts very smoothly.
09:48 So like exploding water. Right. So it's very smooth. And then you change the constraints and you figure out when does it become unbounded. And you know, the fluid velocity just change and moves and it's almost like it becomes sublime and it just changes. Right. But now, now you have a proof that has been solved by Fable so that you can go look at and ask the question, is it safe to take a shower?
10:24 Right. So, so it's pretty neat, right? But I do think we need to be careful, right? I mean, I mean, there is that soap opera, but even so, the mathematical community, it still needs to scrutinize the solution that came out. Right. Of of Astra here. Right. Because we need to make sure that it that it's universally accepted as the solution. Right. And it's correct.
10:45 Right. So so let's just wait and see what that ground truth actually is. You know, but I think the indicators, the litmus test, you know, it is looking pretty good especially, you know, because other folks in the mathematics area are kind of up in arms, say, wait a minute, we solved it first. Right? So it seems as though the signals are there, that it's pretty close.
11:07 Bri, you have great expertise in terms of agent harnesses and how humans interact with with agents. And I think a huge part of the Stokes story is about it's not really that AI solved this problem, it's that humans managed AI to solve this problem. First of all, humans gave us the hunch of where to focus our energies. I gather that as as Astra was turning on it at a certain point, it solved in intermediate, but they were cranking on all seven millennium problems at a certain point.
11:38 Astra kind of solved an intermediate problem. Then they decided the humans decided, let's throw all the energy at solving Navier-Stokes specifically. So there are a lot of human decisions in in the loop as this was cranking. So as you read through this story, what were you thinking about this as sort of a collaborative human AI effort? Yeah, I think that I love how you're framing this because it shows to the world that how regardless of how powerful AI has become and is becoming, it still requires human intellect,
12:06 human creativity and human collaboration to get those answers to these huge problems. So I think that's a good thing. So for everyone who's watching, it is scary. You're going to lose your job next week because of AI. I don't think that's imminent, but we can have a different conversation about the future due to AI. But I think this is this is one thing we can learn from that humans are human brains are still incredibly relevant, and also what we deem as important, because the fact that this was thrown at
12:44 Navier-Stokes and not something else that tells us we as humans say this is important and we need to throw compute at it. The other thing that this elicited excitement within me is think about the advancements we're going to make in health and medicine. If we can use this, and I am convinced we will be solving, we will be finding cures to to probably all cancers.
13:15 I think I was listening to this podcast. I forgot the expert who was on, but he said, we will not be dying from diseases because AI and quantum physics and everything that's going to help solve it. So, I mean, I think the initial solution took 88 hours and then they verified. And I'm leaning on to what Aaron said, because I do believe they used another process where they translated the proof into Lean, which is a, I believe, a mathematical I'm not very familiar with with the concept of Lean, but that only took 17 hours
13:50 to verify whether the results of those 10,000 agents is, in fact, correct, and whether we can start to believe it. I'm sure they're going to do additional verification around that. But I think this is incredibly exciting, and this means we are a lot closer to cures and finding solutions to. I mean, think about global warming and, you know, things that were problems we're facing on a very major scale.
14:20 So those are my thoughts on this subject. Yeah, that's a good note and an optimistic one, which is good. We don't always hear that about AI. going to move us on to our next topic. Long time listeners to the show will know that when Aaron appears we usually talk about sports. And so we're going to move very briefly from curing cancer and global warming to what IBM has been doing at the US Open.
14:45 Aaron, I believe you want to talk a little bit about the demo that you've been working on, right? Yeah. So for more than 30 years. So this has almost been since like 1992, IBM and the USTA, we've partnered together to transform these digital experiences of the US Open. Right. So it's bringing it's about bringing the tournaments data, insights to the more than 14 million fans around the world.
15:09 Right. And just a fraction of them can actually attend. Right. But this year, what we did is we took the partnership to another level by using AI. Right? So so some of these techniques that that we are talking about, and I'll show you what we did, is we predicted matches before they even started. Right. We would understand what's happening like point by point identifying the moments that change the match as the match is being played.
15:35 And then we also created a conversational and personalized experience for every fan. So you could ask the burning questions that you had as tennis was unfolding. And then for the first time, which is really, really neat. It was one of those grand challenges that we took up, but we went all the way down and we traced and tracked limbs, right? For athletes by biomechanics and particularly on their serve.
15:58 Right. And so what I wanted to do was to just attempt to show you some of these demos that I have. So what you're looking at on my screen, if you go to and it's live right now, the tournament is going on. It'll end this upcoming Sunday. But if you go to US Open source and then if you go to like the scores here, you know, you can begin to see who's winning, who's who's doing what.
16:26 And what I did is I picked the the match that just ended against a Navarro. Right. And if I start at the top. So let's look at the preview. So here is our pre-match likelihood to win right. So what happens is is we have the system that ranks all of the players. We have three match forecasters that attempt to say what the volume of the media is being written about them, what the sentiment of different factors are.
16:56 And then we also look at their head to head stats if they have any. And then we also take into account different ratios. How are the players playing on each of the different surfaces, whether it's hard clay, grass, but but in essence there's two classical machine learning models, right. So that should be boosted trees there and and logistic regression as the other.
17:19 Right. And they help predict who's supposed to win pre-match. Right. And here we got it right where Badosa was supposed to win by 77%. And she did. Right. It turns out, you know, that that her odds were pretty good. I mean, we end up, you know, in the I would say in the 60s, you know, with respect to predicting the right pre-match. And then as I scroll down, I can see all the different stats that did happen.
17:43 Right. That just helps to give us the explainability. Now if I want to look here, right, this is the AR replay. I can go in 3D right. And and then visualize and see how they're playing. But before I do that I want to show you the summary of the match. So match chat is a place where I can go and ask questions, you know, during or after the match, so I can better understand what's what's happened.
18:07 And then we have our live likelihood to win, right? So it's bootstrapped and started out with our pre-match likelihood to win. And then as the match goes on, you know we have these momentum equation with different decays that change it. And then we also have different boosters based on the situation at hand. But it moves, you know, the probability that somebody is going to win, you know, as every single point is made.
18:34 So this is a real time system. So as every single point is hit won or lost, we get the message and then we compute, you know, the odds. And then at the same time after we do that we then send other messages over Confluent. And we have a, a marketplace that then takes the data and it measures whether or not this was a key moment. And if it was then we want to produce text.
19:00 And and here this one is the very end key moment. But you can see the, you know, closes out a brilliant 2 to 1 victory over Navarro. But I can scroll back. Right. And just view all the key moments of the match so I can better understand it. And it's pretty cool. It's like a second screen experience. You know, you can watch this right as the match is actually going on.
19:23 Now going back to match chat. If I want to know what happened to the match, let's just push this button here accept the terms. Right. And then what's happening is this is exploring match chat. So it opens up another experience. And this comes and it queries through tools the data sources and feeds that we have. Right. And we create a prompt that has a context.
19:45 Right. And then it goes through right. And it goes through this myriad of graph layering type agents that three of them race, right. One of them is a pure gen AI solution. Another has like a summarization, and then another uses a data synthesizer just because of the load scale that we have. If I want to see. Match details. Then I can do that as well.
20:09 And one question for you is, I'm curious if you guys have information on what types of fans are using this kind of experience? Because I feel like in my family I've got kind of like two kinds of fans represented. One of them just likes to tune in and watch the ball go back and forth. I have another one who's just really would love something like this really is intense about the data.
20:29 Just kind of curious about what you're learning about usage. So we do. Right. We have different user personas that we create these experiences for what. You know what I showed you like match chat is, is for sort of the middle of the road. It's it's not not like a, you know like a hardcore fan or the novice, but it's still sort of in the middle. Right.
20:50 And maybe that fan wants to become or see the data to, to transition them into a hardcore fan. Or alternatively, maybe they're just too busy and they just want to see surface level and they can ask any questions that they want. But we do collect analytics and people do sign in and favorite players, right. And we're able to take that information and, and actually trace it back into types of questions of match chat to generate the types of personas that are used in the system.
21:19 So what you're looking at here this is limb tracking. So we introduced serve quality. And we used IBM's BOD to build it. But what we do is we track 21 points across a player's body at 50 frames per second. Right. So so that turns out that over the course of the tournament, we're processing 1.2 billion data points and apply this to all 254 singles matches.
21:43 And by doing so we're capturing the biomechanics. So if I hit play, you can see the server hitting the ball. And we're measuring effectiveness, whether or not the ball goes into this polygon and the efficiency or how they're swinging the racket. And we've broken up the serve into phases. So we have, you know, the the start phase, a loading phase cocking, you know, whenever you get the racket up and your off to hit the ball acceleration when the racket's going down to contact to deceleration to finish and follow through.
22:11 And we broken out those phases and we apply rules to help us, help us measure that. And in fact, in the match chat experience, you can ask questions about the serve quality. And then on your mobile device, when you go to the home page, we have this live experience that's personalized to you, and we show the serve quality and some insights around it.
22:34 So it's pretty neat. And it's and it's a first of a kind. And I'm really excited about expanding this to other types of shots. Nice. This is great. I love having you on the show because it feels like we every time you come in, like the platform is advancing as we go and we get to watch it together here on MoE, which is, which is really cool. Well, great.
22:56 I'm going to move this on to our sort of final topic of today. You know, I think we've covered sports, we've covered curing cancer. And I think we're going to talk a little bit about the dark side for our final segment. Really interesting story, David, today that just broke over the last few days about the use of AI to generate viruses and effect a thing called WeWorm that a lot of people have been worrying about.
23:20 So I mentioned now that AI has me worried about water and whether I am safe to set foot in the shower. Now, I'm really terrified of my phone because essentially this security startup called Calypso used AI to discover a bug in WeChat, which is the very popular Chinese messaging slash everything app billing. Over a billion people use it in China. And this.
23:54 But this US based firm found this bug using their AI. And it's it's worse than a bug. It's what they were able to develop against. This sort of is a so-called worm. And not only that, but a zero click exploit where essentially this virus can worm its way into one person's WeChat app autonomously, I believe kind of place calls to other contacts. The the recipients of that call don't need to do anything, and yet they are now infected and it just can, in theory, grow exponentially.
24:29 Now, this exploit was discovered in July. Of course, a Calypso did the right thing. They they went to the folks at Tencent, which makes WeChat. They alerted them to the vulnerability. The vulnerability was patched. Only now, months later, is, you know, is the news story breaking. But nevertheless, I am. Personally, I'm not a technical person, but I'm terrified because this would seem like an exploit on the level of what we are more familiar with from like really sophisticated state actors.
25:03 But it would seem that the the rate at which AI is accelerating, the discovery of bugs, the generation of exploits of these bugs, Where should I p(doom) be at this point? I think you have a reason to be concerned, I think. A private consumer of AI, you know, as the end user, you know, we have we interact with these tools. I personally have used WeChat during the two years I lived in Shanghai.
25:31 So I'm very familiar with the app. I paid my rent using WeChat. That's how advanced it was back then, in 2015 to 16, in China. So and whenever I talked with with our customers, security and implications of this, what we're discussing is always part of the discussion. So I mean, I wouldn't lose sleep over it, but it is definitely something we're seeing more and more every single month.
26:00 I believe there's at least one major story about something like this and and several smaller ones. I think the key takeaway here is that AI didn't make the hacking more powerful. I mean, hacking case is always significant event. It made it more accessible. So this is going to continue to happen. And I look to our security experts, which I am not one, to solve these solve these issues around it.
26:30 It's definitely a A question I want to put to Olivia is it's a little technical. I gather that we don't know exactly what the exploit was. Calypso wrote that the bug is a memory corruption issue in WeChat VoIP stack. And they also wrote something that this. This bug is an instance of the many unconventional attacks surfaces that are present across many messaging apps.
26:59 The other thing I noticed in their timeline is they said something like, our AI discovered the bug, and it's unclear to what extent they were just sort of letting the AI roam through various apps to find vulnerabilities. But it would seem to me that is sort of pointing to this, that the advent of AI, the AI era that we're living in, is, is exposing vulnerabilities in places we didn't quite suspect them.
27:23 Right. These so-called unconventional attack surfaces. Can you help explicate that a little bit to us here? Like what is why are new corners of technology stacks suddenly discovered to be vulnerable in the AI era? Sure. So and I've talked a little bit about this the last couple of times we've had like big math problems getting solved by AI. So I think the, the way that I look at this, it's, it's important to remember that we live in a dynamic system and an emergent system.
27:54 There have always been rough edges in everything that humans have ever created. It's actually one of the core beauties of being human, is that the things that we create are not generally perfect objects. There's like entire lines of art basically about that fact. Right? There's things like, you know, kintsugi where we put pottery back together and try to create a whole and things like that.
28:22 There's a lot of imperfections. I know that sounds really philosophical, but what I want to bring that back to is all software does definitely have vulnerabilities of some kind. Some of them require more advanced things to to take advantage of them, and others of them have very simple holes on them. But it's actually near impossible to build perfect software.
28:49 So we can very, very much say that that's probably what's happening. But then I want us to take a look at these models. There has not been a case yet that I have seen where models are truly acting on their own. It is always in response to a prompt, and it's really important to remember that basically models are do not have the capacity to act on their own.
29:18 They are responding to text of some sort of some sort of prompt. Granted in as these models increase in their capacity, they are responding. They're able to do more creative behavior in response to simpler and simpler prompts. So what may have been a while ago, hey, please explore the surface of this code base. Look for this type of vulnerability and that type of vulnerability.
29:46 And this other type of vulnerability may be as simple now as go off and look for vulnerabilities, but that lack of specificity that people are putting into their prompt now is coming back to them in the form of just the sheer amount of resources used. And so, yes, we you know, we can see things like this, but unlike with the. The Astra bit that we saw here, OpenAI's Astra solving the Navier-Stokes problem, where they've actually reported the number of tokens that it required.
30:25 And I think, as we've said, it took some $15 million worth of compute in order to solve a $1 million problem. Obviously, the economics of that just get weird at that point. But anyways, the point is, it took a lot, right? We don't necessarily have that same data, or at least I haven't seen any for this. What does it take to break into WeChat? Right, to to take advantage of these more complex worms.
30:55 And my guess is the number is similarly quite high. So I think there's these real world economics that there are, I think essentially, and something I think we all have to kind of confront is OpenAI and Anthropic are behaving with so many resources available to them that they are able to apply these kinds of things relatively casually. And you can tell that it's casual, right?
31:19 The way that they talk about it, it's like, oh, yeah, we just decided to go solve the Millennium Stokes problems. I don't have access to a $15 million budget like I work in AI. That, regrettably. I couldn't do that. Like, even though I have access to these models, I can't go create worms. Right. And so I think what we have to think about in this space is, yes, things have changed, right?
31:48 We are in a new world, in new types of systems, but they are still very complex systems that are ultimately constrained by real world resources. So personally, my fear is a little bit proportional to who holds those resources and what their priorities are. Right? Right. they have $15 million worth of compute to burn. Okay. But like in the era before, like cybersecurity was still an issue.
32:21 If you had $15 million to pay a hacker, you also could have done a lot of things. Yeah, exactly. Well, this is always a great panel to have on MoE. We get a dose of hope, a dose of terror, a dose of grounded technical information. So, Bri, Aaron, Olivia, great to have you on the show. And David was great co-hosting you. And thanks for joining all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere, and we'll see you all next week on Mixture of Experts.