Claude is an AI assistant developed by Anthropic, positioned as 'The AI for Problem Solvers'. It is a general-purpose AI system used for tasks such as generating code prompts, refining requirements, creating advertising strategy and copy, and processing creative content like storyboarding and video prompts.
Jev is TypeSafe AI's System One model for making fast, structured decisions within software. It accepts unstructured data or program state and returns predefined, type-safe structured values with calibrated probabilities, confidence, and uncertainty rather than generated text. Jev uses a model architecture and parallel sampler that produces outputs in a single query, together with TypeSafe's Reinforcement Learning for Calibrated Decisions (RLCD) training method. It is intended for classification, routing, scoring, extraction, moderation, verification, guardrails, and other AI-powered workflow decisions. TypeSafe describes it as an early-access service for integrating probabilistic decision functions into software, with reported end-to-end response times in the tens to hundreds of milliseconds.
Mixture of Experts is a weekly news podcast produced by IBM and hosted on the IBM Think podcast pages that recaps trends and innovations in the artificial intelligence industry. It publishes audio episodes that discuss recent AI research and developments at the frontiers of the field.
Searchable transcript of New frontier AI models, TypeSafe’s Jev AI, & NASA’s IBM collab — IBM Technology (39:24). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:01 This is becoming a [music] big efficiency problem. And now because you know we're reaching the point where the model names matter less. A new frontier model used to feel feel like a major event. Now we're getting them every few weeks. What I really care about is what's changed. >> All that and more on this week's Mixture of Experts. [music] I'm Tim Hang and welcome to Mixture of Experts.
00:26 Each week, Moe brings together a panel of the sharpest and most good-looking minds working at the frontiers of artificial intelligence to lead [music] you through the week's news. On this week's episode, we've got Couter El McGrowi, principal research scientist AI Native Systems, Gabe Goodart, chief architect AI foundations, and Martin Keane, master inventor, and as always, I'm joined by my co-host, David Zach, who [music] is a staff writer for IBM Think.
00:48 We've got three big stories today. We're going to talk about the widely hyped Jev model. We're going to talk about a really interesting collaboration between IBM and NASA. But first, I really want to start by talking about just the welter of model releases in the last few weeks. [music] Once again, we are in like sweep season for model releases. Gro 4.7 is out.
01:14 Claude Opus 5.5 is out. GPT6 Sol Luna is out. Um, everybody is releasing models and the prices are getting uh lower and lower and lower. Um, and I guess maybe Martin I'll start with you. I mean the question I have always, this is the perennial question with these model launches at this point is um does this matter? What's different here? Are these just incrementally better models?
01:34 Uh, what do you see happening, you know, in the meta across all of these model releases? Maybe that's the more interesting question. >> Yeah, the model releases this week have all been around a theme of efficiency. So we've seen the the big model releases of like Astra and Fable and these are you know Frontier models super super clever models the best really we have but they also eat tokens like crazy and anybody with a subscription to one of these or if you're paying API costs uh very hyper aware of that that uh it
02:04 just takes so much compute for these to run. So the models that came out this week were all addressing that efficiency in that they are a bit smaller I believe in parameter count but they can get a lot more done with hopefully not too much less intelligence and in fact in the case of the Claude model Claude uh 5.5 for Opus that has come out and is maybe as good as the model the the big model of of Fable 5.1 um which goes to show there's probably a fable 5.5 coming uh down the road at some point.
02:41 But it can do basically everything that those big models could do uh at a much much reduced token cost. And uh I've the machines behind me right now are running that model doing a bunch of video editing tasks and I'm just kind of blown away about how efficient they are. And the same thing could be said with uh the new release of Soul as well for CH GPT uh 6.0.
03:03 That is basically a much more efficient version of the previous model which was 5.6. So what we're really seeing today this week is just much more efficient models that are really as good as the biggest models that these labs have. >> Yeah, for sure. And I guess CO the question I had for you because that was what I picked up as well as kind of like everybody's about efficiency now.
03:23 C I'm curious about how you read this as someone who works a lot in hardware and you think about the hardware side of this a lot is it sure seems to me that you know in some ways it's like these companies have to do this because otherwise they're also just maxing out their compute as well and so I'm kind of curious about like how you read this sort of efficiency push um you know weighed against you know the the massive cycles that the companies themselves are burning on their own compute.
03:48 Yeah, definitely. This is becoming a big efficiency problem and now because you know we're reaching the point where the model names matter less. A new frontier model used to feel feel like a major event. Now we're getting them every few weeks. What I really care about is what's changed. Can it work longer? Can it use tools reliably? Can it help me reduce my cost?
04:10 Is it cheaper? Is it efficient? What's the impact on the infrastructure? Or what do I need? Can I reduce basically costs? Does it have maybe new algorithms or better hardware software design techniques and uh things that will help it basically drive that efficiency and I agree with Martin here efficiency is becoming really so important and and now the competition is moving from model intelligence to system intelligence.
04:40 the model is becoming one component inside an agent that has memory tools, a browser, code execution, access to enterprise systems. So that's a much bigger shift than just another few points on a benchmark. So I think thei the system as a whole and how do we optimize all of these things together is becoming really important. >> Yeah, for sure. I feel like the little model selector in all of the desktop versions of these apps is becoming like increasingly confusing.
05:07 It's like here are the 10 models you could use some legacy and you can also choose effort. Um, which is definitely becoming very complicated. Um, Gabe, if your game I mean I think just to kind of exercise some editorial here, I thought 5.5 Opus was maybe the most interesting model that released out of the set. Um, and I don't know if you agree with that, but if you want to take a walk with me, I'd love to talk a little bit about Opus 5.5.
05:32 And you know, the first thing I want to talk a little bit about is like people were kind of actually getting annoyed at Claude in my world where it was like, "Oh man, there's like almost like too much Claude, right? There's like too much personality on these models." And and in comparison, a lot of people have been commenting like for 5.5 uh basically anthropic seems to have buttoned up the model a little bit.
05:50 It's a lot more tur. It's a lot more to the point. Um and that's kind of interesting because I think there's always been a debate over how much personality people actually want out of these models. And I guess Gabe, I'm curious about how you take that. Is it just that they overindexed or is does it turn out that people really don't want a a high personality model?
06:08 >> I think a little bit of both. I do think um there's a bit of a sort of corporate branding delta here. So I think Claude very much attempts to be the get work done corporate branding and I think uh OpenAI especially powered through the chat interface or chat GPT um aims to be a bit more interpersonal. So I think it is in keeping with corporate branding for Claude to try to button it up a little bit and keep it straight and to the point.
06:38 The other important thing to the earlier topic is efficiency, right? If you use fewer words to sound human, you can still get the job done and you cost fewer tokens. So um it's actually probably also in service of that. And that was that the the big the big claim that I saw about all of these models was that they were more token efficient which is an interesting difference uh not necessarily correlated with compute efficient.
07:04 Um so uh I do want to quibble a little bit with Martin here. Those machines behind you are absolutely not running the model. They are running an agent that is talking to the model. So you know what's cheaper than a cheap token rate? a model that actually runs on the machines back behind you. Um, so that's an interesting thing that is not disclosed about all of these new models is actually how big they are and how much compute is required at inference time.
07:32 So this is the biggest question that I had in this wave of models is they're clearly aiming for consumer efficiency by keeping the price low and the number of tokens that a given consumer needs to accomplish a task as low as possible. And that is genuinely goodness because that also correlates with the overall token efficiency on the uh inference side because they don't have to run as many tokens to serve the users that they need.
07:56 But what we don't know is what is the multiplier there? How much does each one of those tokens actually cost them to run? Um and we have some intuitions about how much it cost to train these models. And I think there's rumors floating around about how big they are and what their architectures are, but a lot of those details are still hidden. So the part that I am still really curious about, and I imagine we'll get more details in the coming days and weeks, is how much this race to the bottom in terms of token cost is
08:24 correlated with a race to the bottom in terms of the actual compute costs to run these things or whether we're still waiting for that final shoe to drop of the actual difference between what it costs the companies to run the models versus what they're charging us. Um, I think there's been a long-standing worry that all of these models are subsidized by the venture capital that's backing these firms and once they are no longer uh privately held firms and they're out there in the in the public that that is going to have
08:53 a cost reckoning. Um, and so I don't think we have an answer to that with this round. I think what we have an answer to is that from a consumer standpoint there is clearly downward pressure on the pricing. So, uh, I don't know whether the technology they have is actually backing that up or whether they're just leaning on the extra time by pushing out their IPOs and saying now we've got plenty of time, uh, to continue subsidizing and get more market capture.
09:17 So, we're all going to try to get that lowcost market on our models. So, we we'll see what happens with that, but uh, I'm really curious what happens in that gap between the two. Yeah, the efficiency of the token I think is really it almost reminds me of back in the day, you know, Google always used to say, well, we actually want minimal time on site because it means that you found what you wanted and you left.
09:37 Um, and I guess I'm kind of curious if those kind of same KPIs will start to play out in the AI space where it's like look the minimal number of tokens actually is best because it means that you solve the problem. >> Yeah, I think also from a hardware perspective, this is I think why I think the next phase is going to be very interesting. things like model architecture, inference software and accelerator design.
09:56 They're also converging around efficiency. Things like quantization, sparities, the KV cache management, the speculative decoding, the better batching and also the data flow accelerators, the specialized accelerators. They're no longer just, you know, these things are not just optimizations after the model is built. They increasingly also determine whether the product is economically deployable at all.
10:19 And if we see the trends that we've been been seeing like between 2023 and 2024, it was mostly about bigger and bigger models. In 2025, I thought the better reasoning was kind of the headlights there. And this year it's more focused on better reasoning per unit of compute cost. So I I echo what everybody's saying. Efficiency is becoming very important.
10:44 And there is some conversions in the uh uh hardware accelerator design space. There's still a lot of things that we still need to work out in terms of really reaching that really efficiency targets. It's still very challenging. >> Yeah, absolutely. Yeah, I feel like I'm paying a eye watering amount for the the top tier, but I'm still getting a good deal out of the companies.
11:04 I'm still being subsidized. So, [music] I'm going to move us on to our next topic. We're going to stop talking about the latest generation of model releases to talk about a model release. Um the next topic I really wanted to cover was uh Jev which has been getting a lot of press and I think a lot of hype recently. Uh the background is a company uh called Typesafe which has announced that they have built a what they call a new class of models called system one models and Jev is really their kind of first public release
11:34 of this for coding. Um and David there's a bunch of really interesting attributes for uh what they appear to be doing here. First of all it comes from a former open AI engineer or researcher named Dooo uh Almeida. who was sort of early at chat GPT or at OpenAI rather. He wrote the co-wrote the kind of foundational instruct GPT paper. Um that said, he was kind of I hadn't heard of him.
11:56 He was kind of under the radar uh from from my point of view. And then he he released this model and it just sort of blew up, right? He had a X account with 1500 followers. His post got 40 million views, I think. So what is so exciting about Jev to people? I think it does speak to this theme of efficiency that we're talking about, right? I think um there's sort of one or two value props that I want to talk about with respect to Jev.
12:20 But the first is to do with kind of fit for a fitfor-purpose model, you could say, right? So I think the contention of of Jev is hey we are using LLMs which are basically like novelists, right? They can fluently generate word after word of pros for human readability. And then increasingly in this era of the agent and the enterprise use case, we're then kind of, you know, retrofitting that and forcing that into making these simple structured decisions for sort of often the multiple choice questions that that kind of
12:52 underly software and processes. Um, so Jev contends that it can if if I understand correctly like be about as smart as a frontier model, but can be much more efficient by instead of outputting um words that then get rendered into, you know, decisions within a kind of set of of constrained values of typed values as the term in software, right? Um, it's just going to output those typed values directly and it's going to be much more efficient in the process.
13:19 Karpar, did I get that right? Have I understood this correctly? And secondly, are you are you buying the hype here? When you looked at the Jeff release, what you know, are you as excited as everyone else is? >> Yeah, I think it's a very interesting direction that they're taking. And of course, the claim that they have is we don't need every time to generate strings and sentences and so on, certain tasks, you just need maybe decisions or scoring or yes and no answer.
13:43 And those especially will be very useful when the software is using these LLMs for in the agentic world close etc where I don't need to read a whole sentence then try to interpret what this is saying and then turn it into decisions or some score etc. And if I look at the hardware angle here what I find interesting about Jeff from a hardware perspective is that they're changing the workload itself with a normal language model.
14:10 Once you start generating an answer, you generate one token, then the next token, then the next etc. And that creates that sequential decode loop. So every step needs to go back to the accelerator, read the model state, the KV cache, does the computation and write things back and then repeats. And that is not a particularly efficient way to use expensive hardware, especially when all you really wanted was a decision like yes or no.
14:39 which option should I choose or what is the probability of each outcome. So Jeff gives us this free form text generation for those workloads and instead it kind of uh produces the structured decisions and these probabilities in parallel. So from a hardware side I think that is attractive because parallel computation is exactly what GPUs and AI accelerators are really good at.
15:05 So you have less sequential decoding here, much less output generation and potentially much less KV cache traffic. Uh so I think part of their efficiency gain is not necessarily that they found a magical way to make matrix multiplication faster, but they change what they're asking the hardware to do. But I still I have some reservations about you know some of their claims uh the their their in terms especially of the correctness because they still rely on models to generate some of their scores and some of their claims.
15:42 So I think that part you know we need to be careful about especially the calibration that they mention if a models gives me a confidence score I need to know what that score whether that score stays reliable uh when the data changes so for automation I think that it becomes very important and maybe even more important than maybe the latency numbers that they're claiming.
16:05 I want to chime in there and uh I'm always a little interested in catching up the people at the back of the class like me who are non techchnical and I want to talk a little bit about what I learned this week about this theme of calibration. you mentioned confidence scores and calibration and broadly again my understanding is that um a model uh can make a guess right it can say um I I think that's a stop sign and it can be right or wrong but almost just as important it can be overconfident or under underconfident in
16:34 its prediction right so it can say that's a stop sign and I'm 90% sure but if lo and behold when it makes that that guess and that and that confidence score and and it's only right 10% of the time, you have a major overconfidence problem. And so that is what we call I I believe calibration, right? It is is sort of um adjusting it so that the confidence scores actually reflect how good the model actually is.
16:58 And for years now, we've been so frustrated with hallucinating models that very confidently spout utter nonsense, right? And so um this isn't this they claim that this model can't hallucinate which maybe is true or not but more to the point I think if I understand correctly what they've optimized for is really getting that calibration right so um they their model kind of lives or dies by getting the confidence score right and it may not be if I understand correctly again the best at guessing whether that's a stop sign
17:29 or not but if it tells you it's 60% sure it it better darn well be 60% sure not more and not less. So first of all Gabe did I get that right? My understanding of calibration and secondly if I did why is this so important like what becomes unlocked especially in enterprise use cases once you can be confident of the model's confidence once you can be once it builds the trust of its confidence score.
17:54 Um, okay. Yeah. So, uh, on the confidence score, uh, you did get that right. And I want to, one of the things I loved about this announcement was it, I think, basically hit on some critical moment in every single aspect of my historical journey as a software engineer. So I have so many reactions to this model. But on the confidence score, it reminds me very very directly of problems we had in the pregenerative AI days building classifiers and creating REST APIs for said classifiers and deciding what was the name of the
18:28 field that we used to indicate the quality of the output. Was it probability or was it confidence? um we always chose confidence and not probability because probability has a mathematical definition. It is a distributional statement. Confidence is just a number. Uh confidence is just we made it up. Uh we slapped a sigmoid on there and it looks like a nice bell curve when we look at it over the distribution of inputs.
18:57 And sure, this one's better than that one and this one's got a higher number than that one. So trust it. Um so this you're exactly right to pick on this point because uh this gets into some real challenges of using this type of opaque approximate math to do things that actually have rigorous mathematical definitions like probability. Um, the end goal of all of this, however, is useful decisionmaking, which is where a model like Jev or even our pre-generative AI um, classifiers typically tended to say, well, actually
19:36 confidence is fine because the only thing you really need to be able to do is take the scores relative to one another and make a decision based on those scores. So, the downstream consumer of a Jev output is simply going to take that confidence distribution and say okay where's my threshold that I say left or right where's my threshold that I say red yellow or green where like where I just have to pick a number and and I have to do that empirically that's that's left as an exercise to the user um so uh the calibration
20:06 is important and I'm not sure whether they are claiming that it is a proper probability distribution that would be a hard mathematical claim to back up but um they are probably at least claiming that the confidence scores are are true relative to one another. Um, I want to pick on a very different aspect of this model. Uh, I'm just going to steal the floor while I have it because I think it's really important to couch this model relative to all the other models and where it fits.
20:34 Um, so we've talked about this as essentially a distribution predictor. Um, and uh, in other words, it's a classification model. uh classification is a problem that has been it's it's in many ways the oldest problem in NLP or even not even NLP in in AI in general. You get a feature vector in uh and you try to classify that feature vector based on some uh set of labels that could apply to it.
21:04 Um the reason that many many many people have moved away from training their own classifier models which can be orders of magnitude smaller than LLMs to be clear smaller and more efficient both in terms of you know runtime and compute power needed. The reason everyone has moved away from that is that in the bad old days you had to train your own classifier in a supervised learning manner.
21:28 You had to have examples of the inputs and the outputs. And LLMs because they have this massive ground truth knowledge behind their weights are really good at taking what's called zeroot classification problem. So you give it a few examples which essentially tickle the the weights area that is relevant to your classification problem and you say use those examples to form your model's intuition and then splat me out a classification for this thing please.
21:53 However, uh as Kowar accurately pointed out, generating tokens by tokens that says I believe this is a stop sign period. I think so because it has and then you just have to write a reg x on the back end that says okay I got a big pile of gobbledygook. I I got the word stop sign in there like you know do I have to check for in case insensitivity? Do I have to look for weird white space in between?
22:19 Like you know it's just a a software engineering nightmare. So getting out a vector of numbers that says you know entry one is stop sign entry two is you know yield sign etc like that's way way way more efficient both in terms of inference compute and in terms of the engineering effort to extract the output. So in that sense this is great. However LLMs are used for so so so so much more than classification.
22:44 So um what it comes down to is that this model is targeting a very very specific and important problem that is a subset of the problems that LLM can tackle. Now uh I want to point out that this area of making deterministic decisions uh or at least close to deterministic decisions uh is one that we at IBM have been tackling with a program uh programming model not a AI model but a programming model called Malaya.
23:13 So this is essentially a way of inverting what many people are doing which is describing a program in plain text and hoping that the model will actually execute the steps that you say. Those steps in the plain text description are perfect for a model like Jev. So uh a programming model like Malaya can take a model like Jev use it where there is a clear decision point to be made in a workflow and actually create a deterministic program that has small non-deterministic elements where you have to make a decision.
23:43 So that's the real power of a model like this is being able to take something that you know is a well- definfined path through actions and make good decisions with non-deterministic inputs at the important junctures of that thing so that you can walk the right path. >> Yeah. The the thing that really interested me two things. One the name this is a system one model.
24:03 What on earth does that mean? Well, apparently it's called a system one model as reference to the system one and system two thinking from Danny Canaman's thinking fast and slow book, which I was delighted to see that because I have the book with me. I like I read this thing >> quite frequently. I come back to this. So, I love that idea. >> For every purchase made off of this episode.
24:25 >> Oh. Oh, absolutely. Click click on my affiliate link in the description. What that was saying is, you know, system one thinking is kind of the automatic thinking. It just kind of happens. We don't have to consciously think about it. And system two thinking is like thinking more through step by step. And if we think about that in terms of AI models, large language models today, thinking models in particular, they do this chain of thought reasoning, this kind of system two thinking.
24:47 And one of the the arguments for Jev is to say, look, for a lot of these decisions, we don't need the system to step-by-step thinking. We've been using system two when we really only needed system one for things that are just a classification problem. as as Gabe was mentioning. So if the answer is yes or no or the answer is pick one of 15 options then a large language model is a huge overhead to to go through um to to come up to that answer to drive to that answer.
25:17 So it was interesting to hear how Jev was trained. That was the other thing that really got me interested because large language models they're typically trained to be primarily chat bots certainly initially and that uses reinforcement learning from human feedback. So we're trying to get an answer that is most agreeable to a human. And then these models as they've become thinking models and are now used more for AI agentic purposes as well.
25:43 There's reinforcement learning from verifiable rewards which is to say if I ask for you to do this thing, did you actually do this thing? And the model's rewarded when it does. But Jev and these system one models are built on a different reinforcement learning which is reinforcement learning for calibrated decisions. So this is not trying to come up with a answer that is pleasing to a person, you know, it's structured in that way and it's not trying to come up with an answer that that comes up to a reward.
26:11 It's just trying to say is this answer right or not. Now David, you talked about calibration and you know this is a typed model. It gives you back a number basically as to say you know what the answer is and of course the question is is the answer right? And it's not coming back with any supporting information. So if I ask a large language model a question, it will give me an answer, but it might also actually tell me why it thinks that is the answer.
26:36 Now that reasoning chain, how it got to that answer may be entirely or at least partially hallucinated, but at least it gives you an answer. With this model, it's just going to give you a number. So how do you know it's right? Well, the benchmarks that the team behind Jev have have offered up show that it is able to get to the same decision as the top frontier models.
26:59 So, if you ask the same question to Astra or to Fable, most often Jev would also come up with the same answer, but much quicker and with fewer token costs. So it's able to reach a an answer that a large language model will come up with but it hasn't been tested against you know actual sources of truth. I mean it could be that the large language model is also wrong.
27:21 So it will be interesting to see more benchmarks where it's actually tested against things in real life. Um that said some of the demos they provided. So, there's a demo of of Jev basically playing Doom where it figures out how to move around in space and to shoot the bad guys and so forth. Um, and you know that is making constant decisions, these these yes or no decisions or where should I move to?
27:48 Should I move left? Should I move right? Yes or no? And so forth. And it kind of works in that that perspective. So, it was pretty accurate at at, you know, hunting out the bad guys. And then I saw another demo where they kind of played the Wikipedia game which u we've all played this right where you you start at one page like I don't know um the English medieval period and then you've got to get to another completely unrelated page like flying cars >> like Kevin Bacon.
28:14 >> How many >> Yeah, exactly. How many clicks does it take to get there? Right. And it turned out that um the best LLMs could often get to from topic A to topic B like within normally five clicks which is pretty impressive. But then Jeff could do it in five clicks as well but it could do it in subsecond periods of time whereas a large language model would have to go through it all of its reasoning chain of thought to get there and would take you multiple seconds to do it.
28:39 So there are some early indications that not only is it coming up with the same answers as large language models, but also that those answers are generally going to be right. Uh but there's there's so much more to see with this. So I think with with Jeb when it first came out, you know, everyone's first thought or at least my first thought was, "Oh, another classification model.
28:59 That's wonderful." You know, we've had these for a long time. But but seeing the training behind it and this idea of reinforcement learning for calibrated decisions and seeing how well it stacks up to the very top frontier models with large language models, I think that's why it's this kind of such an exciting technology to keep an eye on. >> Well, this is all great.
29:17 Awesome discussion and I'm glad we got the explainer here because I think um you know again I think this is why for in a lot of ways is to be able to cut through some of the hype and get uh behind the scenes on what's going on. >> [music] >> We're going to stop talking about a model to talk about another model. I think the last one I really wanted to cover today um is a one that caught my eye.
29:39 It's a nice collaboration between IBM and NASA. Um specifically the release of a set of models that are kind of specific for lunar exploration. Um, and so these are effectively kind of computer vision models that are really good at kind of picking up on sort of lunar surface uh images and being able to identify, you know, features for researchers. And um, you know, David, I like this one a lot because I think like in some ways this is sort of a really interesting story of essentially the way that machine learning can
30:07 be used to extract more knowledge out of the data we already have, which I think is a thing that we don't talk about so much, but I think is really interesting as part of the story here. >> Yeah. At first, I mean, sort of like Martin, my first reaction, I'm kind of like a moon and like Apollo fanboy. Uh, but um, at first I read about this, I kind of wanted to cry tears of boredom because it was like, it helps you count craters quicker or something like that.
30:35 And I was like, why does this matter? Are there are humans counting craters at all? That sounds like the saddest job in the world. But then I learned that actually the these this is very important that if you can count the number of craters in a certain area of the moon, you can infer certain things about the the age of the solar system and how often because actually those craters are sort of like a clock for the universe, right?
30:55 Because they represent the rate at which kind of external debris kind of hits the moon and then you can infer things about the age of Mars and all sorts of things. So I became very excited uh learning about how this model can accelerate the science. Uh, beginning with you, Gabe, what got you excited about this model? Um, I saw there was also it can help identify ice, which is apparently a very difficult kind of multi-parameter problem to to solve uh to and could have implications for lunar exploration and colonization.
31:25 So, I think what got me excited about this model was uh it's AI for nerds again, which is awesome. I mean, AI used to be the the realm of of of nerds and scientists and geeks and people that like to think about esoteric problems that most people would look at and say, "Why the hell are we counting craters on the moon?" And then, as you accurately pointed out, or at least I hope is accurate because I am not a moon nerd.
31:49 Uh, apparently craters on the moon means a whole lot to the age of the universe and and we can learn tons from that. And that's where geeks live. That's where nerds live. Like so um you know we we live in an age where AI is starting to become not just a fascinating technology but a hugely important sociotechnical problem. And those of us that have lived in this from the interesting math and the interesting geeking out about what we could get our terminals to do are starting to find ourselves wrestling with very human
32:21 problems. And that's important. Um but it's also uncomfortable. Um, so it's really cool to see a good classic like AI for the sake of the science uh problem here. So I think that got me excited. Um, you know, I also love seeing AI applied in a domain that is not a chatbot or a coding agent. Um, simply because it's a good reminder that this foundational technology applies much more broadly than the consumer use cases that we're all consuming.
32:51 Um, and I know that that conversation is often lost when we get into the latest releases from the Frontier Labs. Um, but there's some real benefits to the overall knowledge and state of humanity that we can glean out of the core technology behind all of this that probably have nothing to do with uh how effic how token efficient we can uh generate, you know, comics from our favorite blog posts.
33:22 >> Yeah. I think I agree with you Gabe and I what I really like a lot about this project is that really it shows that a foundation model doesn't have to be a chatbot and the value here is reuse and so you train a large body of lunar data once and then you can adapt the representations to really several scientific questions instead of just starting from zero every time and that is of huge value and and also what I Think the most important work here is is also the data engineering not just the model size.
33:56 The sensors here have different resolutions, different observation conditions etc. And aligning all of those measurements and giving the model uh information about how the data was captured can also be just as important as adding more parameters. There's a lot of interesting engineering data problems here that we can learn tons from and um and of course just you know ask being able to answer these big scientific questions you know has huge value to you know to to humanity of course >> I think the other thing that makes
34:31 us super interesting is you know we when we think of foundation models uh you the first most obvious thing is to think of a textbased foundation model I put some text in it generates text out. But of course, we have multimodal models now that work natively with other modalities other than text, right? So, we're all familiar with image generation models where an image can go in and image can come out.
34:55 And the way that works is just basically as long as the information can be represented as vectors as a series of numbers, then it can be used in a foundation model. So, it's nice to see that expanded to things like, you know, pictures taken in space and to be able to represent that information again as vectors and then to be able to use a foundation model, really the same sort of foundation model that's dealing with text, but now it's able to do clever things like look at a particular image that was taken where the sun
35:22 is at a particular angle that's kind of obscuring a little bit how the craters are visible and to be able to adjust to that. all just using the same sort of math in these models, the same sort of reasoning that that is also being used to like write us a poem in a in a textbased large language model as well. So, a really cool use of a foundation model just applied to something completely different.
35:45 >> That's great and a great note to end on. Um, well, that's all the time that we have for today. Martin, Gabe Kowar, always a pleasure to have you on the show. [music] Um, and thanks for joining all you listeners. Uh this is actually going to be my last episode hosting Mixture of Experts, but you are all in good hands [music] uh with David here on the show.
36:01 And uh if you joined what you heard, you can get us on Apple Podcast, Spotify, and podcast platforms everywhere. And we'll see you all next week on Mixture of Experts. I whether whether this could edit it out or not, want to do a little bit of a bigger send off for Tim and and we can, you know, see what uh whether this passes the sensors. But uh I will just say, you know, I've been a fan of the show for a long time.
36:24 I think Tim, you've done a splendid job, uh making technical topics really approachable. And um uh my daughter, my 2-year-old daughter is obsessed with uh wearing her parents' shoes and stumbles around in my apartment, and I'm stepping into some very large shoes here. May I fall on my face less often than uh my 2-year-old daughter, but um yeah, you we uh this is quite a responsibility to step into.
36:55 So, thank you so much for what you've done. >> Yeah, thank you. It's very generous and and Godspeed. Uh echoing that, Tim, I feel like one of my favorite things actually is to always listen to the episode myself after we have recorded it because I feel like when we're recording, I'm in the mode of like, okay, I got to think of what my insight is going to be about this topic.
37:13 And then when I listen to it again, I realize that like your transitions between topics always have this like nugget of leading insight that clearly like cues it up so nicely. So I you know I know your job is basically passing the ball back and forth but you're the way you have done that it just orchestrates so nicely this flow. So uh I've always appreciated that.
37:35 >> Thank you. >> Totally because um it would be very easy I think when you're asking the question to the next person that you kind of put could easily put the person on the spot. It's like oh I wasn't expecting that one because we don't have like a script knowing what's coming in. and we just have the topics, but you always manage to sort of move the conversation to the next person in a way that makes it just like blindingly obvious to me what I need to say, [laughter] which is like that's a real >> Yeah.
38:01 >> Yeah. Yeah. I echo what everything was said. You've done a tremendous job, an amazing job here, you know, to keeping the show so exciting. Even for us, you know, as we participate, we're always at the edge. What's going to happen next? You know, what's going to be the the next question? how do I get prepared for this? So, and that really makes it interesting and so naturally flowing and I think we can see the the engagement also from the listeners.
38:26 So, thank you so much, Jim. >> Thank you. Yeah, I it's a great compliment of like two years of just keeping us on our toes. So, I really appreciate that. Thank you, Kat. My testament to this is that I got a a cold email the other day asking me to be an angel investor in a startup because they had heard my quotes from like the last year on mixture of experts and thought I was somebody that would be ready to be an angel investor on so and I was like well >> yes how much do you need?
38:53 >> Well why yes I have I have a $5 bill in my wallet. Would you like that? >> [laughter] >> But but you you've all you've brought a level of uh sort of panache to this uh this show that really has has made it stand out as a you know a a rare light of consumerf facing value out of IBM. Let's put it that way. [music]