← All transcripts

Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1 Transcript, AI Summary & Key Points

IBM Technology · 20 days ago · Education · 39:02 · EN

📄 Transcript

Searchable transcript of Thinking Machines Lab drops Inkling & Meta’s Muse Spark 1.1 — IBM Technology (39:02). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 Again, the debate question isn't a leaderboard anymore. It's better open base plus a fine tuning platform can beat the closed model or other model strategy, right? All that and more on today's Mixture of Experts. I'm Tim Hwang and welcome to Mixture of Experts. Each week, MoE brings together the best banter in artificial intelligence to walk you through the welter of this week's news.

00:23 On this week's episode, we've got Aaron Baughman, IBM Fellow, Merve Unuvar, Director, Agentic Middleware and Applications. And Chris Hay, Distinguished Engineer. Welcome to you all. Four big stories today. As always, we're going to cover Muse Spark, GPT-5.6 Sol's performance against ARC-AGI-3 and this really odd paper out of Anthropic about the J- space.

00:45 But first I really want to start by talking about Thinking Machines. So this is the kind of hotly talked about anticipated lab from former OpenAI CTO Mira Murati, and this is a lab that's kind of floated around for a while. They've done some blog posts, some research paper releases, but just this week they have launched their first model, a model they call Inkling.

01:09 It's a 975 B total parameter model, but it's a mixture of experts. So 41 B active. And this is really interesting. So it's the first real release from the lab kind of manifesting their approach to some of the contemporary problems around model training and model pre-training and their sort of philosophy for going about doing this. Interestingly, it's an open source model, and I guess maybe Chris, I'll kick it over to you first.

01:33 I know you're the most excited to talk about this. What's your take? I mean, one of the most interesting things about this in my mind is, you know, they're quite straightforwardly saying that this is not the top model from a benchmarking standpoint. What do you what do you read into that, if anything? I love it. It's like the Inkling model. It's like, ah, it's okay.

01:52 You know what I mean? you know what I mean? Like it's cool. Yeah, exactly. I don't think in this market you need to be the best in the world. I actually think they needed to get a model out. Well, that's the biggest thing. They were lacking, and they did it. And I love it. Actually, it's it's it's not as you say, it's not the greatest model in the world.

02:14 You know, the frontier labs are all better than it. The, you know, all of the Chinese models, GLM, etc. they're all better than that. So but here we have a great and it is a great open weight model from a US lab. Awesome. And there's actually I, there's some really nice architectural things in there which I really like the direction they they're heavily influenced by the DeepSeek models, which I think is great.

02:38 But their approach to multimodal is a little bit different as well. So I definitely appreciate that. So if you think of the other multimodal models that we've seen, it is like a what's the best way of saying it. It's like melding a bunch of different models together. Right. Here's a vision encoder. Here's an audio decoder. I'm going to go mix it together and then you go, yeah, no, it's one model.

03:01 Honestly it's one model. Don't don't look under the hood. Please don't look under the hood. But actually they've they've actually went truly multimodal. And they're really sort of taking it down to token level. And they're taking a small set of pixels I think it's 40 by 40 projected on and then it's turned into tokens. So it's actually true multimodal.

03:20 And I appreciate that. And then I really I like their approach to speed as well. So there's sort of taking a bit of a fast path. So it is quite a fast model and that since I'm not going to go into the technical details of it, but, they've made interesting architectural choices. It's not the same model as every other model, so I applaud them. I've played with it.

03:43 It's pretty good. I'm waiting for the great version and yeah, I'm excited. Yeah. Merve, I think, you know, to what Chris is saying, it does kind of feel like Thinking Machines is almost trying to change the competition a little ways, right? Like that. Like usually what we'd say is, is it the top benchmark? Okay, then it's the best model. We should all use it.

04:02 Here they they seem to like just say, yeah, yeah, yeah, there's a benchmark. But here's all this other stuff you should care about. And you know, it certainly sounds like Chris has kind of bought that story, right. Which is natively multimodal and kind of all these interesting architectural things that we're doing. Do you think that's going to shift the market?

04:18 Like do you think people do you feel like this is where things are going, where, you know, benchmarks really are just not going to be the thing that people care about as much anymore? Yeah, I believe so. And I think like they are, different approach here. Like more than of course, like whatever little the advancements they've done and differences in the architecture and the model is Chris said, they what I'm impressed by the, this like a fine tuning platform that they built, right?

04:44 Like, you basically like you don't give a prompt, but you say like, I don't know, the example they I read was like writing using letter E, and it drafted the plan, generate its own eval and synthetic data. Ran the post training through their Tinker API, loaded the new weights back into its coding harness, and answered correctly. So this is like a full closed loop, right?

05:07 Like that. You know the MoE that is you make. But the orchestration is the point here, right? Like they can do this fully, automatically. And I think going back to the competition here, like, as you said, like they stayed in like Inkling is not the strongest model available, open or closed. I think the whole pitch here or the differentiator. Best open base to customize for your needs, right?

05:31 Like the multi-modal efficient thinking. Fine. Tunable on Tinker. So again, the debate question isn't the leaderboard anymore. It's better open base. Plus a fine tuning platform can beat the closed model or other model strategy. Right? I think it's worth the try. And we'll see how it plays out in the market. Or maybe we can kind of zoom out a little bit in this discussion.

05:54 You know, I think this is a caricature, right. But I think the discourse for some time has been like, America is really good at the closed source proprietary models, OpenAI Anthropic. And when it comes to open source, we're okay. But like we really look to is like places like DeepSeek and, you know, these kind of Chinese labs that are really advancing the state of the open source.

06:13 You know, I think this is notable in some sense because it's kind of like this, you know, kind of widely watched US frontier lab using open source first, you know, as part of its release. And I guess maybe a question for you is like, you know, maybe do you think we've been thinking about this all wrong? Like, that actually just turns out that, like, actually America can do open source, super, super well as well.

06:34 And we can do fundamental things there, you know, same way as a DeepSeek. You know, I guess kind of how you weigh that up. Maybe we've just been mistaken in all this time. Yeah. I mean, I mean, I think with Inkling, the real headline here is customizable intelligence that's been open sourced, right? Because really, the future of AI, it might not be about building the biggest model.

06:54 That in turn is a frontier model that could be closed source. But it's about giving people the ability to customize, adapt to control these types of models for their own needs. Right. And and I noticed that, you know, within this chain of, you know, what is the Inkling, what's the difference? Why does it matter? And then does it really matter that it's not the most biggest, most powerful model?

07:14 I don't think it does at this point. You know, because, you know, there's there's there's many layers right to this. You know, you have this general intelligence layer that's a pre-trained model on broad capabilities. But then you have this customization through fine tuning. Right. So these different ways of, you know, the parameter efficient fine tuning capabilities of LoRA, for example.

07:34 Right. Then you go down to being able to use RAG to, you know, get the data to fine tune right. Then you can also customize, you know, behavior control and the way of which it's going to think through the prompt itself. You can customize agents and customize through these different types of feedbacks. Right. But this way of having Inkling with this self improving method so that it's really customizable towards your own needs, I think is key and even better, right, is that it is open weight, because this is a bit of a

08:08 different spin about, you know what all of the, you know, the main competitors, right, are doing, you know, such as, you know, Sols, you've got Fables, you've got, you know, Gemini Pro for example. And then I think underneath that is where Inkling sits. But I don't think Inkling right now is trying to compete right, with those big models that are very deep.

08:26 Right. But but it's a smaller type model when I say small, from, from what I could find. Right. It was about 975 billion parameters. So it's still a pretty big. Yeah. These are all models still. So, you know, the measuring stick has changed. You know so much, you know, because I remember many years ago, I, you know, a hundred thousand models and parameters is big.

08:51 But now but now we're at 975 that are now open source, open weight with 1 million token context window that can handle multimodal input. But again, let me circle back to what the headline here. It's customizable intelligence that's now been available made through these open weights. So so I'm excited about it. And and I do think Inkling has a place I think one of the things if we go back what was about a year ago, maybe, maybe just less than a year ago, where they did their first blog posting and everybody's like,

09:23 what's Thinking Machines gonna say? And it was like, oh, here is how to do determinism, right? We have managed to control the batch. So, you know, so this whole reproducible and we're like, boom, we want to see a big model from your research. We don't want to see this. And then but actually if we if we fast forward now. Actually the very things and principles that they put in from the beginning, the fact that they do have determinism across all sizes and reproducibility, I suspect that's what's allowed that going to

09:55 allow them to move from, we're not quite at the frontier level yet, but oh my goodness. I think they've got that control experimentation wise to be able to move past that very, very quickly. And again, you can kind of see it in the architectural choices. They've dropped RoPE. They've got, you know, the fast lane for the sliding context windows, etc..

10:10 So they're taking an awful lot of optimizations. And the only way you can take those optimizations is through measurement and reproducibility. So they're they're not just sitting there going, hey, I'm the Swedish chef. I was like hurdy gurdy, gurdy. And then put the ingredients in and we go, we got the model there. They're doing proper controlled experiments.

10:32 And and the fact that they're at this point here, I, I, you know, I'm, I'm an optimist. I think maybe a few months time we're going to start to see a really competitive model because I think they're taking the right approach. Well, that's actually a great transition to our next segment. We've talked a little bit about Thinking Labs, which is maybe a lab that's kind of competing different.

10:52 We should now talk about a lab that I think is like kind of competing. Same in some sense. Meta. After being quiet for some time, but creating a lot of noise in terms of its recruiting and the money it's pouring in. Is out with kind of its kind of first statements of the kinds of models it wants to release. Muse Spark 1.1, which is the latest model from the sort of so-called Meta superintelligence lab.

11:22 And, you know, this is in contrast, in some ways to the the Inkling launch, a very different kind of model launch. Right? I guess, Merve, to your point, earlier, you know, it shows all the charts and the end result of all the charts is that they're killing it on the benchmarks, and they are better than, you know, all these other proprietary models, OpenAI and Anthropic.

11:44 And, I guess the question for you is, Merve, do you read these results as saying that Meta is now sort of once again a serious contender in this space, because it feels like we have not talked about them for quite a while. You know, you know, I feel like Llama was the last time they were really capturing the narrative, and this might be a chance for them to maybe get back into the game.

12:03 Do you think there's kind of a credible foot forward? No, I think it is like after that rough Llama stretch. I think coming back with this model especially, I think, aiming for the market is going right. Like, the agents, not chat, multi-agent orchestration. I think they emphasize the planning and then the delegating, you know, the parallel sub agents to cut latency.

12:26 Computer use, for example, is a popular use case across multiple apps. That that that is gonna, I think, gonna make them stand out. And I think the interesting fact is like the million token context it actively manages and compacts like that's also a strong, feature that they packed into this. And I think they also have a strong coding, harnesses in this release.

12:51 So this is like a, as you said, like after Llama open weights, you know, leading the open, you know, model, era for, for a long time right now we have a coherent opinionated that like they want to be there. Also, the other highlight is like the cost efficiency, right? Like, they're cheap and fast long context engine. So you can run agent workloads on scale.

13:08 So they're also going after the enterprise use cases here. So it's it may not be the smartest I know the benchmarks leading, but I think the most economical one to build agents on. And I think this could be their, catch and to build success on, Aaron, is it it's a little bit odd, in light of kind of what Merve is saying, is that, like, it does seem like where you would go with a model like this is to start offering kind of enterprise AI services.

13:36 And I guess we don't normally think about Meta in these terms, right? The creator of Facebook, the owner of Instagram, WhatsApp, you know, these are these around large consumer tools, right. And so it's kind of interesting that like the suggestion here that's kind of reading between the lines is does Meta believe that they're going to become an enterprise kind of B2B SaaS type of operation.

13:58 Is that is that even something that you think is possible? Right? I mean, you know, you know, if you start to organize all the different players you know, about, what have they contributed to the AI field? You know, the way I look at it is that OpenAI's somewhat proved AI could think. Let's put that in quotes, right. We'll talk about that in a bit. And Google proved that AI could scale, right.

14:20 And now I think what Meta is doing is they're betting and showing that the winner will be the company that puts this intelligence everywhere where the people are. Right. And and so that connects the dots to all of their social media platforms that they already have, you know. So so it's pretty exciting to to see you know, what they're going to do because this is their biggest strengths.

14:41 They have a large user reach, you know, hundreds of millions of people on this platform. You know, you already mentioned WhatsApp, Facebook and so on. But they also have, you know, somewhat of a nice hardware ecosystem. You know, I was just just playing around with their Ray-Ban smart glasses. It's I mean, it's fun, you know, but, but but if you could get this model linked up into it, it'd be very powerful.

15:01 They have the research talent right there. A long leader in AI. You know, we've we talked a bit about the Llama models, but don't forget about PyTorch, right? Even back in the day with deep learning libraries, you know, if people are still which they are, you know, building their own neural networks, which in turn are building blocks for these AI pieces.

15:16 They contribute to that. And then they also have an open plus, a proprietary strategy which we're seeing emerge. Right? I mean, open, you know, they've been very, very good at that. And now this proprietary with Spark I think is where they're starting to, to attack. This model does have a little ways to go. I looked a little bit at the benchmarks and noticed that on the ARC, the AGI you know, piece, it's a bit weaker right, than the competitors right now at least.

15:43 So it's not quite up there at the frontier capability. You know, for that particular type of test, it is great at these knowledge base kind of tests, right. Where if you look at the artificial analysis index, for example, GP diamond. Right. So there's a whole suite and it does very, very well and it puts it up in the upper echelon. But I think where the field is going, it needs to jump up a little bit.

16:11 Right. And we'll talk a bit about Sol and what all this means. Right. Coming up. Chris, the other thing I want to kind of investigate on this story was, you know, we've been talking a little bit about how, like the the Meta strategy of launching models is changing. You know, particularly with Fable, Mythos and kind of how OpenAI approached their model where they said, okay, well, only some certain people are going to get access to it, and then we're going to more broadly distribute it.

16:33 In some ways, like, I feel like the the the Spark release here is like kind of a throwback. And by throwback I mean like two a few months ago where, no problem, we're going to just make the model kind of like openly available. And part of it is the kind of like push distribution versus being really, really kind of constrained about who accesses it early on, I guess, I don't know.

16:53 Do you think like Meta is eventually going to go the way of the other big kind of frontier model companies in this respect? Or is it kind of still open like that, that they're going to keep kind of leading the way in terms of like shipping and launching without a whole lot of restrictions on initial launch? I, I mean, in order to get restrictions, your model needs to be good enough, first of all.

17:09 So I so, you know, once you get there, we can have that conversation. And and and to Aaron's point, it's not there. And I have to say and I kind of feel this is kind of like a meh release and it's and I and I love Meta, right. I love the Llama models. I thought there were a lot of fun. They were releasing things that were open way they sparked the whole community to get started.

17:33 Right. And and you could argue that that was really the Mistral guys and they, you know, went off and created Mistral. But but but I you know, here's a closed model that powers social media. And you can use the APIs and and look it's better than 4.6 Opus. Yeah. We all know it's not better than 4.6 Opus. I mean, you should be comparing it at least to 4.8, if not Fable also.

17:57 But but you're comparing it to 4.6 and I just don't buy that it's better than 4.6 because everybody goes, oh, it's better than holding this. And and then never is. It never is. So am I going to go and spend money on it? No. You know what they should have done. They should have released it as an open weight model. And then we'd have been like, woohoo!

18:18 Well done. Meta. Yeah. You know, and but no, now I'm gone. Oh, you know, Thinking Machines etc. that's what they needed to do. I understand why they're not doing that, but I'm it's like, do I care about a closed model that's worse than the great models? No, I just I just don't care. You know what I mean? It's like. So there we go. And, Chris, it's funny.

18:44 Merve Murphy. I'm sorry. Go ahead. I just want to say, like I noticed throughout the release note, they kept using, like, competitive with leading alternatives or like, competitive with, like, today's leading frontier models. Like, nobody writes competitive with when they beat everyone, right? Like you're right. The number is competitive. You're like, I know.

18:58 Yeah. All right. Well, Zuck, if you're listening to this, you should work harder to impress Chris Hay, he'll be he'll be on it for the next round. I want to like their model, I like Meta. I know I'm one of the few people that like Meta, I like Llama models, etc.. I, I just want you to open it up or, you know, but I mean, if you turn, I mean, Merve to your point, if they were like, this is 20 times better than Fable, I'd be like, let me add this model.

19:26 I want to play with it, but it's not. It's like, you know. This is this is like the, you know, the Honda of models, and you're like, great. You know, I don't want to drive a Honda. I want to drive a Ferrari. Do you know what I mean? Yeah. Well, that's probably going to be the last word on this one. Move us on to our next topic here. So the next story we want to cover is, pertaining to a benchmark called ARC-AGI that we've talked about before here on MoE.

19:56 When it was originally released, it was sort of slated to be like the hardest challenge ever for AI to solve. And really the kind of core idea the benchmark is, can models learn entirely new skills? You know, without further kind of like prompting or frameworks or kind of like being guided along to solve the solution, and we've seen more and more progress on it.

20:19 But this week, ah, there was once again in the news, because OpenAI's GPT-5.6 Sol one of their latest releases really showed this big jump in progress where, you know, you look at sort of Luna and Terra and they're scoring low like many of the other kind of like models. But then with Sol, you actually see this jump right where it goes from 0% basically to 8%, which, given the difficulty of the benchmark, has people sort of chattering.

20:49 And I guess maybe I'll kick it over to you first. You know, how should we read these results? You know, 0 to 8% maybe. Doesn't sound like a whole lot. But in contrast, if you look at the benchmark, it does seem something that's really difficult for AIs to do. And so I know, I guess. How do you handicap this? Is this an impressive result or, you know, I guess the, you know, if you want to be the Chris of this segment, maybe, maybe this is like a boring result.

21:09 I don't think it's boring, but I don't think we're nowhere near AGI yet. Right? Like, so like every time the field saturates, you know, on one of these, the ARC releases a harder one targeting like what the models still can't do, right? Like and it's 12 year old kid can do. Right. So and then the score collapses again like 8% versus you know, any human that can do like 100% is not true AGI.

21:39 And like I don't think the score is what's going to convince us when we get to an AGI. Like the day when someone releases a new benchmark full of tasks that are easy for people and none of these models just, you know, or frontier models just handles it out of the gate, right? No specialized personal, you know, work at front. And then I think this AGI will be here.

21:56 So but this is a I think it's a good progress. Like if, if we look at the ARC-AGI-2, it was, you know, up until, you know, 92% I think like very high percent. So it can come up there too. But, we're nowhere near, I think, for AGI yet. Yeah. I mean, this is one of the hilarious parts of that. I buried the lead a little bit, which is that this is performance against ARC-AGI-3.

22:22 And, you know, in some ways when this benchmark was first released, they were like, once you pass this, this will definitely be AGI like saturated. And now they've just been this game where they keep having to launch new benchmarks to be like, it's not it's not AGI just yet. And and I guess Aaron, I mean, do you, do you think we're going to see the same thing with AGI, AGI three, which is, you know, this is like, you know, it's like a boat with a hole in it, like the water keeps coming in and you cannot just you just

22:52 can't keep ahead. Like anything you think is too difficult for a computer, you know, it ends up appearing to be able to do in much less time than you think. Yeah. I mean, I mean, throw me some more buckets so I can keep getting that water out of the canoe. Right. I'm getting more and more holes. But yeah, I mean, I mean, Sol is not is not AGI right.

23:05 It's but but it is one of the clear signals that the gap I think between today's frontier models and something we would recognize as AGI, AGI is beginning to shrink from like the fundamental limitations that we have now to towards you know not not yet but towards, you know and engineering and scaling problem. And I think that's indicative of the ARC-AGI-3 results.

23:28 You know popping up to roughly 8%. You know. So that's I mean that is pretty impressive. You know, because AGI is all about you know, I have an AI system that can perform on a whole range of intellectual tasks that humans can perform, rather than just being specialized just in one particular domain. And I think what this is showing, right, is that Sol is beginning to be able to reason within a world of uncertainty that maybe it hadn't seen.

23:56 But it's this unfamiliar environment, you know, that it's had. And and I mean, you know, it was it was released not that long ago. Right? So, I mean, I know that they've used, you know, megawatts of power, you know, you know, maybe 10,000 houses worth of power to a train, this type of model. But but, I mean, you know, I, I think it's only going to get better.

24:21 Right. And, it it is great. You know, whenever you look at the benchmarks at AGI one and two. You know, the ARC set said it did really, really well. And then on three it you know, it's beginning to sort of look underneath the water at that iceberg, you know, and begin to explore that space of it ordinarily couldn't have seen. So. So yeah, I mean, I mean, I think over time, you know, that that ARC.

24:46 Right. It's it does. First of all, we need to make sure ARC-AGI-3 is the right benchmark, that we want to march all of our models to. Right, you know, and then look at other benchmarks and sort of decorate that so that we can define what AGI is as a as a group in a community. And then began to get these models marching towards where we ultimately hope that they can go, at least some of us hope they can go and a responsible way.

25:11 But let's also think about, you know, it's a trade, the amount of power. Right? Because I was stunned. I was reading about this that even on inference, typically, you know, whenever you're starting to get into class and in the neighborhood of of AGI. But inference cost about 60 homes worth of power, right, just to run that type of job. So it's pretty incredible.

25:37 You know, just the power that all of these use so that we can begin to have a model that's somewhat going upwards on the trend of being able to solve 8 percent, solving the ARC-AGI-3. We got to do better than that. I mean, I get your point. We maybe moved up to 8%, but I think $19,000 to solve that particular task. I think there is some architectural changes that will need to happen to hit that, to hit that number in my opinion.

26:01 So it's great to see the number going up. But but is the dollar value going up at the same time? That's that's a bigger question. Maybe a final thought on this and we'd be curious if anyone has any thoughts on on this question is, you know, I guess, like, are we just already at AGI? There's like, another position here, which is that we keep trying to come up with tasks that are harder and harder and harder.

26:26 But it is true that, like, if you just, you know, thought back to where we were three years ago and you talk about what the models can do today, you would have said, yeah, okay. That's that's AGI, we're there. Right. Like we have a model that can just do these incredible things for most of the tasks that you ask them to do. They can do it. And in fact, now we're in this situation where we kind of have to like keep coming up with more things that it can't do to kind of hold this line.

26:52 And do people agree with that? It's like, maybe we're just actually already at AGI. We've actually been at AGI for some time. I think I already said it at the beginning. I don't believe so. I think the day when, you know, we released these, new benchmarks and the model just is out of the box. Great. That's the day I think we can claim we're getting to AGI because, like, every time, you know, when you release a more difficult or a little bit different benchmark, it's very embarrassing to Chris like 8%.

27:17 I mean something, but like, come on, me for 12 year olds can do 100%. And if a model does 8%, this is definitely not a right. At least 12 year olds. Exactly. Yeah. Yeah. I mean, I mean, I always, always look at, you know, humans and and their growth, you know, into it into an adult. And so I think maybe we're at like a infant level AGI. Right, right.

27:36 Maybe we're trying to get to a one year old AGI, you know. But we have a long ways to go to get these models into adulthood. AGI you know, but you know, because fundamentally we're still trying to define what does AGI actually mean. And and I just always say that it's the full range of intellectual tasks that humans can, can perform. But then you ask yourself, so what age group, how are we going to stratify a human's capability and measure it?

28:03 Because, because yeah, we could fulfill that definition. If we're looking at, you know, a one week year old baby that is born, right, because that is a human being. Right. So, you know. So it's just a matter of where you're measuring stick is going to be. Right. And then I think all of these benchmarks that we're creating is then tagged, you know, to what what a human capability is.

28:25 But on the other hand, I think we're also trying to surpass human capability. You know, because humans can't look at hyperspectral bands. You know, you know, we but but we have all of these different interpretation of tools to go out and view the world and then pull in data that humans can't sense. Right. And and so so it also, I think there's an AGI that goes beyond, you know, what the humans can do.

28:46 So. So I think bottom line, it just depends on how you define what what AGI is. And how does AGI map to your organization's business goals of which they want to track to. All right. Last topic for today. This was an interesting Anthropic study that I wanted to quickly touch on before we closed the episode today. And I guess I'll read the tweet where they announced it, because this created a bunch of controversy when it came out.

29:16 Anthropic writes of everything happening in your brain right now, only a tiny fraction is consciously accessible. Thoughts you can describe, hold in mind and reason with. We found a strikingly similar divide inside Claude, and the paper goes on to describe what they dub kind of the Jacobian space or the J-space. And so the claim of the paper is basically that there are sort of regions within the model that do reasoning and processing that aren't sort of explicit outputs.

29:45 And so Anthropic looks at this and says, well, there's actually really this kind of interesting structure between sort of the conscious, you know, outputs of the model and all of the processing that it does that is kind of like sub lingual or subconscious maybe. And so this created a bunch of controversy. Right? I think, you know, some people took a look at this and they said, you guys are reading way too much into what you're discovering here.

30:10 It's interesting, but it's definitely not. You know, the AI has a subconscious. And I kind of want to bring this to the panel because I read the paper and I was kind of like, I don't know, I can kind of see it both ways. Chris, maybe I'll start with you. I mean, do you buy kind of the, the maybe the frame that Anthropic is bringing to some of the results that they have here?

30:29 And how much of it is kind of it is kind of, you know, kind of poetry collaboration and how much of it is like a real, true way of understanding what's going on in these systems? I think all of the above, I hate to say it. So to answer. So, I mean, the first thing is I think the J-space is a good breakthrough, right? Because if we if we think about how we would look under the hood, we'll ignore chain of thought for a second.

30:55 Right. But the techniques you would tend to use is things like logit lens. So you would look at the various layers, and then you would be able to look in and then see the top tokens. And you can kind of project the activations against the tokens. And you could see roughly this aligns to this text. And and what they're doing in this particular case is they've got a new technique which is kind of similar to that, where they look at the activation to do some averaging, etc..

31:18 And then you can say this is roughly within this space and you can they can roughly say, well, if, you know, if we average this out here, it's thinking about this, is thinking about whatever. Right. And I and I think if you I so I think they have a better technique there. And they're absolutely spot on. Right. Which is there is the final output which is your token output that you see.

31:40 But there's a whole lot of I'm going to avoid the word thinking, but there's, there's, there's a whole lot of calculations that happen underneath there. And what they're doing is lighting that up. But the big thing here they're doing is not projecting that into token space, right. They're using this Jacobian technique. So and that gives you an idea of what's thinking about there.

32:00 So I think that's really interesting. And and we kind of see that. So everything that's going on before but before it sort of projects down into a token. So would I call it consciousness, though? I mean, I, I mean, they avoided saying that themselves. I would say they're strongly implying it, I think. I mean, they're smarter than me. So, you know, they've got the Claude model.

32:24 Maybe, maybe if you Anthropic, if you want to send the Claude model across to me, I'll take a look at it. I'll even agree with you if it's conscious. But you need to give me the weights, my friend, to play with. So I think, yeah, I the answer is, I don't know. I, I'm definitely not smart enough to do that, but I think the technique is great. And I think they're shining lights on the darker space that is below the, projections to be able to see what's going on.

32:53 So I think all of that is, is goodness. And, and I think the outputs that they're getting from, from this is great. I think probably the most interesting part of the paper, you know, and I've done a lot of experiments in this area myself, but I think the most interesting part of the paper for me is the fact that they. They see the chain of thought. Whenever you know.

33:14 The chain of thought is used as a scratchpad, etc. to, you know, to basically say, you know, this plus why, blah, blah, blah, blah. But, but actually one of the things you're seeing is when you kind of take away the chain of thought elements, or rather, there was thinking that was going on in the J-space where you introduce chain of thought. It then ends up on on on your chain of thought, and it's not no longer sitting in the J-space.

33:34 And I think that's interesting because it's kind of showing the modal scratchpad is like it's internal thought versus, you know, writing it down onto a piece of paper and then offloading that. Now we kind of know that that sort of exists beforehand. Because if if you play with the model there and you think about next, and this is really the crux of things like distillation, when you do a token and Anthropic some previous papers on this before, when you do your next token prediction, there is enough hints on the

34:04 previous token, and fingerprints in that sense, that leads you into what the next token is as well. So so they they all exist, I think is great. I'm just not in the whole conscious thing, you know? But I mean, good luck to them. But I mean, awesome piece of work really is. Yeah. Consciousness is a is a big word. Aaron, what do you think? Are you in all of the above kind of guy like Chris here, or are you more skeptical?

34:25 Yeah, I mean, I'm yeah, I think so, right? I think there's a lot of overhype here, right? I mean, just just the headlines of the paper, right? I mean, you know, the overhype elements seem to be, you know, Claude has subconsciousness or the model's thinking, or now we understand how it work. But I think what this paper is on par for is that we can see inside the model, you know, you know it.

34:48 It appears that really what this work is about is almost like building. You could think of like an fMRI, you know, to see the pathways, what's really happening within these deep neural networks that have been constructed and trained. But it's not really about that. This is AGI. This is conscious. Right. Right. Right. I think that's very misleading. Right.

35:13 And and if you can cut through all that fluff and really get into what the contribution of this work is, then it's really these things that they call attribution graphs, right? You can identify all of these meaningful features that are happening. You can test different hypotheses like some of the experiments that they showed. And then one interesting aspect about this was about AI safety.

35:35 Right. Right. So so if they can do all of this right. And, and figure out what parts are being lit up, then maybe you can inspect that internal representation and be able to know, hey, maybe this maybe we can detect hallucinations better. We can find, deceptive reasoning easier, right? You know. So, so it helps us to have a little more patterns that we could use Within the security aspects.

36:03 But what this paper again going back to. Hype. Right. This paper I do not think claims that it's conscious. It has this subjective experience. This it has emotions. I don't think this has anything to do with that. I think it's that that that is just trying to get eyeballs. Right. But I would say check out the paper, you know, and the work and look at the kind of experiments they run.

36:27 It is very interesting what they're doing. And it does open up, I think, some new avenues of interpreting how these, you know, gen AI models really do do work, and it can help us with things like jailbreak or hallucination and so on and so forth. Merve, you've got the last word. Hot takes on this one as we close out. I know, so. Well, you asked if this is a self hype.

36:51 I'm like, there's a ton of writing controversial parts in this paper. And I think this is by design. Like they wanted to create this debate. And I first thought this was a great marketing material for Anthropic. Right there. We're just talking about it. But then I realized that Neel Nanda, who runs Interpretability at Google DeepMind, and he was able to replicate some of these results on open source models.

37:12 So there is an external validation. So as Aaron was saying, there is something into it, but it's hidden too much under consciousness. And this notion is both interesting and very scary. Right. But I think I want to look at the practical side, as Aaron was mentioning. Like I work with Agent Builders, right? And I don't want this to be lost until consciousness headlines that the paper has.

37:34 Like, if you strip all these words and look at what is there is like this localized, readable, editable internal state, right. So for anyone building agents that take actions on behalf of the users, that's exactly like the handle that you want, right? Like today we mostly monitor agents by their outputs. Like basically judging someone's thinking only by what they say out loud.

37:56 Right. We don't see inside. But this is a tool that sees like, what a model represents internally before it acts or, or like, it's like a real lever, I think, for safety and control, as Aaron was saying. So I think this is huge for real world applications that we're building. So I see a good benefit. I'm looking forward to seeing like what the, you know, neuroscience community is going to take this to from consciousness.

38:23 But yeah, it's quite interesting. Well, whether it's consciousness, AGI or, you know, just cost efficiencies and models, you get everything here at MoE. Chris, Merve, Aaron, thanks for joining, as always. And that's all the time that we have for today. Thanks for joining all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere. And we'll see you all next week on Mixture of Experts.

💡 Answer

Inkling prioritizes an open, customizable base model and fine-tuning platform over benchmark leadership, while Muse Spark 1.1 emphasizes efficient enterprise agent workloads but is viewed as a less compelling closed-model release.

🧠 AI Summary

Thinking Machines Lab released Inkling, a 975 billion total-parameter mixture-of-experts model with 41 billion active parameters, emphasizing open weights, native multimodality, speed, determinism, reproducibility, and customization through the Tinker fine-tuning platform rather than benchmark leadership. Meta's Muse Spark 1.1 targets agent workloads with multi-agent orchestration, million-token context management, computer use, coding capabilities, and cost efficiency, but is criticized as a closed model that does not clearly surpass leading alternatives. GPT-5.6 Sol improved ARC-AGI-3 performance from 0% to roughly 8%, showing progress but not AGI; high inference costs remain a concern. Anthropic's J-space research provides a method for inspecting internal model representations and may improve interpretability, hallucination detection, jailbreak prevention, and agent safety, but does not establish model consciousness or subjective experience.

🔑 Key Points

  • Inkling is a 975 billion total-parameter mixture-of-experts model with 41 billion active parameters.
  • Inkling uses an open-weight, natively multimodal architecture and is designed for customization through fine-tuning.
  • Thinking Machines Lab's Tinker API supports an automated loop for generating evaluations and synthetic data, fine-tuning, and reloading model weights.
  • Muse Spark 1.1 focuses on planning, delegation to parallel sub-agents, computer use, million-token context management, coding, speed, and cost efficiency.
  • GPT-5.6 Sol achieved roughly 8% on ARC-AGI-3, compared with approximately 0% for earlier referenced models, but this does not demonstrate AGI.
  • ARC-AGI performance must be considered alongside power and monetary inference costs.
  • Anthropic's J-space and attribution graphs expose internal model representations without proving consciousness or subjective experience.
  • Interpretable internal states could improve monitoring and safety for agents that act on behalf of users.

✅ Actionable items

  • Use an open base model with a fine-tuning platform to customize behavior for specific needs.
  • Apply an automated customization loop that drafts a plan, generates evaluations and synthetic data, fine-tunes through the Tinker API, reloads weights, and tests the result.
  • Evaluate AI progress on new tasks rather than relying on a single saturated benchmark.
  • Assess model capability together with inference power consumption and monetary cost.
  • Inspect internal model representations and attribution graphs to investigate hallucinations, deceptive reasoning, jailbreaks, and agent safety.

📣 Marketing

Branding

  • Thinking Machines Lab positioned Inkling around customizable intelligence rather than being the top benchmark model.
  • Meta positioned Muse Spark 1.1 around agents, long context, speed, and cost efficiency.

Distribution

  • Inkling was released as an open-source/open-weight model.
  • Muse Spark 1.1 was described as openly available at launch while also being accessible through APIs.

Customer acquisition

  • Anthropic's controversial framing of its J-space research generated attention and discussion.

🧭 Frameworks

Open base plus fine-tuning platform00:31
  1. Start with an open base model.
  2. Customize it for specific needs through fine-tuning.
  3. Evaluate the customized model.
Inkling customization loop04:44
  1. Draft a plan.
  2. Generate evaluations and synthetic data.
  3. Run post-training through the Tinker API.
  4. Load the new weights into a coding harness.
  5. Answer the evaluation task.
Model customization layers07:17
  1. Use a general pretrained model.
  2. Apply parameter-efficient fine-tuning such as LoRA.
  3. Use RAG to provide relevant data.
  4. Customize behavior through prompts, agents, and feedback.

🧰 Tools & AI usage

  • Tinker API — Run post-training and fine-tuning for Inkling.08:17
  • RAG — Provide data for model customization and fine-tuning.07:40
  • LoRA — Provide parameter-efficient fine-tuning.07:34
  • PyTorch — Serve as a deep learning library and building block for AI systems.11:11
  • Logit lens — Inspect model layers by projecting activations against tokens.30:58
  • Attribution graphs — Identify meaningful internal features and test hypotheses about model behavior.35:23

AI is used for

  • Fine-tuning and evaluation generation — Automatically customize Inkling by creating a plan, generating evaluations and synthetic data, running post-training, and testing the new weights.00:44
  • Agent orchestration — Plan tasks and delegate them to parallel sub-agents to reduce latency.02:00
  • Computer use — Operate across multiple applications as an agent use case.02:06
  • Internal representation analysis — Inspect model processes and support safety work such as hallucination and deceptive-reasoning detection.04:51

📊 Numbers mentioned

Costs

  • Sol's training was described as using megawatts of power, possibly equivalent to 10,000 houses' worth of power.
  • Inference for a job in the neighborhood of AGI was described as costing about 60 homes' worth of power.

Growth

  • Sol's ARC-AGI-3 performance rose from 0% to roughly 8%.
  • Earlier referenced performance on ARC-AGI-2 reached 92%.

Pricing

  • A particular ARC-AGI-3 task was described as costing $19,000 to solve.

Traffic

  • Meta was described as having hundreds of millions of people across its platforms.

⚖️ Advantages, risks & lessons

Advantages

  • Inkling combines open weights, native multimodality, a 1 million token context window, speed, and customization.
  • Thinking Machines Lab's determinism and reproducibility support controlled experimentation and architectural optimization.
  • Muse Spark 1.1 is positioned as economical and fast for large-scale agent workloads.
  • Meta has broad user reach, social platforms, hardware, research talent, PyTorch, and both open and proprietary AI strategies.
  • J-space analysis offers a way to inspect internal model states before they become explicit outputs.

Risks

  • Inkling is not currently the strongest open or closed model by benchmark performance.
  • Muse Spark 1.1 may not be sufficiently capable to justify choosing it over leading alternatives.
  • ARC-AGI-3 performance remains low relative to human performance, so it does not establish AGI.
  • High training and inference power consumption may limit the practicality of advanced reasoning systems.
  • Consciousness claims about J-space can overstate what the research demonstrates.
  • Monitoring only agent outputs can miss internal representations that precede actions.

Lessons

  • Benchmark leadership is not the only basis for evaluating an AI model.
  • Open, customizable intelligence may compete with larger closed models for specialized needs.
  • Reproducibility and measurement enable more reliable model experimentation.
  • AI progress should be evaluated across changing benchmarks, generalization, capability, and cost.
  • Interpretability research is most useful when connected to concrete safety and control applications.

💬 Quotes

Best open base to customize for your needs

Concise description of Inkling's central positioning.05:31

Customizable intelligence that's been open sourced

Captures the main significance attributed to Inkling.06:48

It's not AGI

Direct conclusion about Sol's roughly 8% ARC-AGI-3 result.23:05

This paper I do not think claims that it's conscious.

Clarifies the limits of Anthropic's J-space research.36:09

👤 People & companies

Tim Hwang

Host of Mixture of Experts.

00:18
Aaron Baughman

IBM Fellow and panelist.

00:23
Merve Unuvar

Director of Agentic Middleware and Applications and panelist.

00:27
Chris Hay

Distinguished Engineer and panelist.

00:27
Mira Murati

Former OpenAI CTO associated with Thinking Machines Lab.

00:53
Neel Nanda

Leader of Interpretability at Google DeepMind who replicated some J-space results on open-source models.

06:11
Thinking Machines Lab

AI lab associated with Mira Murati that released the Inkling model and built the Tinker fine-tuning platform.

00:53
OpenAI

AI company referenced as a closed-source model provider and as the developer of GPT-5.6 Sol.

00:59
IBM

Company associated with Aaron Baughman's IBM Fellow role.

00:23
DeepSeek

AI lab and model family cited as an influence on Inkling's architecture.

00:38
Meta

Company behind Muse Spark 1.1, its Meta superintelligence lab, Llama models, social platforms, PyTorch, and Ray-Ban smart glasses ecosystem.

01:50
Anthropic

AI company whose research on Claude's J-space and attribution graphs was discussed.

04:51
Google DeepMind

Research organization associated with Neel Nanda and referenced in the J-space discussion.

06:11
Mistral

AI company referenced as an open-model competitor and as a company formed by former contributors to the open-model community.

02:43