← All transcripts

OpenAI's always-on agents, 700+ math manuscripts & HackerRank's AI interviewer Transcript, AI Summary & Key Points

IBM Technology · 7 hours ago · Education · 34:43 · EN

Watch on YouTube

AI Summary

OpenAI's DevDay launch of Dots brings always-on, persistent agents to the mainstream, with the panel questioning how much work can truly be handed over without constant permission-seeking. OpenAI released 722 AI-generated math manuscripts on GitHub covering 372 families of math results, including a claimed solution to the four-dimensional Canelo conjecture, raising the question of whether humans can keep pace with AI-generated mathematics. HackerRank's Chakra runs AI coding interviews — half a million beta interviews over six months, used by Capgemini and other large firms — and the panel debates bias, expanded access, and whether skillful AI use should itself be a screened competency. Reflection AI's Beam prompts an explainer of mixture-of-experts architecture and the difference between open-weight and open-source models: open weights give you the matrices but not the training data or the compute.

Key Points

  • OpenAI introduced Dots at DevDay: an always-on, persistent agent with a cloud environment, computer use, and definable personas that drew cheers from the crowd.
  • Dots handles delegated tasks like keeping released code healthy, but it keeps returning to users for permissions and clarifications — so full delegation is not yet real.
  • OpenAI released 722 AI-generated math manuscripts on GitHub covering 372 families of math results, including a claimed solution to the four-dimensional Canelo conjecture.
  • Mathematicians at Meta produced six research papers through the meta.ai chat interface, using Meta's Muse Spark flagship model from Meta Superintelligence Labs run by Alexander Wang.
  • Terence Tao's position is that mathematical contributions must be understood by humans; one open question is when AI will generate new math questions rather than only answering decades-old ones.
  • The 722 manuscripts may start a treadmill: human attention is finite, and more AI progress can be made while humans sort through the last dump of AI-generated results.
  • A Quanta Magazine article by Jordana Cepelewicz argues the profession of mathematics must transition into interpreters who explain AI's work to the rest of humanity.
  • HackerRank's Chakra runs AI coding interviews where candidates are allowed to use AI, and the AI pauses them to probe why they made choices — half a million beta interviews over six months, with Capgemini and other large firms using it.

AI in practice

Used for

Agents

  • Dots — Persistent, always-on delegated tasks such as monitoring released code and keeping it healthy, running multiple tasks in parallel without the user coordinating. 2 held 01:00

Tools & resources

6 items

CNo. 5465
AIAINotes.us AI product

Chakra

In the AINotes directory

Chakra is HackerRank’s AI interviewer for coding-skills interviews. Candidates can use AI during an interview while Chakra observes how they use it and asks probing follow-up questions. The video describes it as publicly available after a beta in which it conducted approximately half a million interviews over six months, and names Capgemini among its users.

Mentioned in
1 video
Kind
AI
HNo. 5467
AIAINotes.us AI product

HackerRank Chakra

In the AINotes directory

HackerRank Chakra is an AI interviewer for coding skills assessments. It allows candidates to use AI during a task, pauses the assessment to probe their reasoning, and evaluates how they use AI as part of the interview.

Mentioned in
1 video
Kind
AI
MNo. 5468
AIAINotes.us AI product

Meta AI

In the AINotes directory

Meta AI is a chat interface used by Meta mathematicians in work that produced six AI-assisted research papers.

Mentioned in
1 video
Kind
AI
MNo. 2316
AIAINotes.us Tool

Mixture of Experts

ibm.biz/~OIYjPLWCH

Mixture of Experts is a weekly news podcast produced by IBM and hosted on the IBM Think podcast pages that recaps trends and innovations in the artificial intelligence industry. It publishes audio episodes that discuss recent AI research and developments at the frontiers of the field.

Mentioned in
8 videos
Kind
Other
ONo. 5466
AIAINotes.us AI product

OpenAI Dots

In the AINotes directory

OpenAI Dots is described as an always-on, persistent AI agent introduced by OpenAI at DevDay. It maintains conversations, manages tasks, and accepts delegated work such as monitoring released code.

Mentioned in
1 video
Kind
AI
ONo. 0097
AIAINotes.us AI product

OpenClaw

Open source · openclaw/openclaw

OpenClaw is a self-hosted AI assistant and agent platform developed by the OpenClaw Foundation. It runs on macOS, Linux, Windows, or WSL2 devices and connects hosted or local model providers, tools, skills, plugins, messaging channels, and optional companion apps through a Gateway. The same Gateway architecture supports a personal assistant on one device or a trusted shared-team deployment. The Gateway is the local control plane for sessions, tools, events, and channel connections; the Control UI, CLI, and terminal UI connect to it. Channels include WhatsApp, Telegram, Slack, Discord, Google Chat, Signal, and iMessage, while companion apps and nodes can provide voice, Canvas, camera, screen, and device-local actions on supported platforms. OpenClaw’s security model treats inbound messages as untrusted input, pairs unknown senders by default on direct-message-capable channels, and runs tools on the host for the main session unless sandboxing is configured. The project is distributed under the MIT license.

Mentioned in
25 videos
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of OpenAI's always-on agents, 700+ math manuscripts & HackerRank's AI interviewer — IBM Technology (34:43). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 How much can you reliably [music] just hand over and forget about? There are instances wherein it does come to you for permission. So if it keeps uh coming back to you asking for permissions or asking for more clarifications, then you're not truly relegated. All that and more on this week's Mixture of Experts. [music] Hello, I'm David Zach and welcome to Mixture of Experts.

00:24 Every week we sit down with a panel of IBM experts to talk through the week's news in AI. My guests today are Utsavv Doakia forward deployed engineer, Skyler Speakman, a senior research scientist, and Kush Varsni, IBM fellow. Welcome to the three of you. We're going to talk about a couple of things today. We're going to talk about how AI is doing math at scale, about a robot that does job interviews.

00:46 We'll even explain the name of this show. But first, let's talk about how Open AAI wants to give you an agent that never logs off. [music] So, OpenAI had its dev day uh about a week ago and it introduced something called DOTS. So, in last week's show, we talked about Meta's Muse, which is this sort of relentless always on agent persistent thing, sort of almost evocative of the open claw uh thing that everyone was panicked about about a year or two ago, and now this relentlessness has come to the the larger labs.

01:23 Utav, you were there at Devday and you you got to go hands-on, I think, a little bit with Dots and certainly to see Sam Alton and others present it. What were your impressions about dots? >> Yeah, thanks David. So, uh I was there and I could truly feel the uh energy in the crowd. Uh that was their main launch. Uh of course there were engineers who were also going to talk about the durable nature of the agents and we have seen other uh such products in the market as well.

01:53 The interesting angle uh OpenAI will have at its disposal is going to be the cloud environment uh CEX on cloud the personal computer use that you get with it which may be available but they have got a strong ecosystem where engineers developers are all embedded into it. Um everybody knows OpenAI as a household name with Chan GPT. So that that comes with its own reputation that they bring and um uh dots as much as I've played with it they are able to maintain the conversation well they're able to manage tasks you can

02:26 delegate to them but what I'm more curious to test now is how much of the work can you actually delegate and not have to coordinate later on. um whenever we have talked about and one of the demos they showed is you can hand over a responsibility to dots saying that hey I just released my code here and now can you go and uh make sure that it stays healthy.

02:48 So that kind of work dots can do persistently well. But what if you had so many other tasks running in parallel and you were to say that hey can you just stop doing that then does it mean that it's going to stop pursuing all the tasks or whether it stops pursuing one single task and how much can you reliably just uh hand over and forget about? Um there are instances wherein it does come to you for permissions.

03:14 So if it keeps uh coming back to you asking for permissions or asking for more clarifications then you're not truly delegated. So that's the angle I want to test it with. But super exciting uh super personal take they took on it by giving it persona and characters that people can define and there was a there was cheer in the crowd when they launched it.

03:34 So that says something. I think open AI and chat GPT have been synonymous and the I think outside the technical crowds they've related AI just to a chatbot right something you see in these in this little square on your screen it'll be really interesting to see how both Muse andor dots take AI to the broader public not in the form of a chat interface so that's there's not that's not really a question to anyone but I want to come back to that question a month from now and look back to see if that really has taken AI out

04:07 of the chat box to the broader population. I think that's something that both Muse and Dots uh might be judged on in the next couple months. >> It is sort of funny to me that like a week ago we were criticizing OP open AI for agents escaping boxes and and now in this other context it's like this so cool that it's breaking out of its box. But I I do think that like this was kind of one of the things that felt weird about OpenClaw um a year or two ago now, right?

04:34 Is that it showed up in your text messages, right? It was just like the AI broke out of where it was supposed to be and it it just sort of was an ambient her like assistant for you, which a lot of people thought was cool. not least Sam Alman who then hired the the OpenClaw guy and one imagines that that probably he had something to do with the development of dots.

04:58 Kush, what do you think about this? I I totally agree with Sky that the mere fact of an open AI agent turning up in your text messages feels more important than it should and I'm curious what you think about that. >> Yeah. Um and actually OpenClaw it's not a year or two, it's only 9 months. I mean time is flying by, right? So um Everyone knows that time is dilating as though we are near a black hole in the AI universe that we're living in.

05:24 [laughter] >> For sure. For sure. Yeah. No, I mean um it's dots, it's muse, um instinct came before it. Um there was even one that came out yesterday, I think, called Tab. And um yeah, like all of them like this cartoonish identity, I think um is the like just [clears throat] like it's a different way of approaching it, right? And I think um the fact that all of these things are coming at the same time like everyone's just copying each other.

05:52 It's the same ideas. So um I think one angle that I wanted to kind of make a point about is that it's like this whole field is kind of over capitalized in a sense like everyone has invested so much money um that they're just like there's no imagination left. I mean, everyone's doing the same thing. Like, there's no freedom to to do anything different because they need to show that, oh, if one person did it, the the other person needs to do the the same thing.

06:17 And I think that'll kind of permeate in what I'm going to say throughout the episode because it's like everyone's copying everyone else. And I think it's a symptom of uh just the fact that you have to invest so much money to to get these models up and running. Yes, it does seem almost [laughter] almost as though the firms are coordinated in terms of their decisions to to release a cuddly agent at the same time.

06:40 Um I'm sort of struck, Sky, by the um disconnect between the cuddliness of the agents on the one hand, right? So Muse, if you've tried it out in the meta app, it's adorable. It's fuzzy. It looks like a Disney character. dots, the the, you know, these adorable little gumdrops uh on the one hand versus the absolute relentlessness in terms of what again it's the open claw thing of 9 months ago which which everyone thought was crazy and dystopian.

07:10 And so is it precisely that they've you know that that that many firms are trying to wrap something that is kind of terrifying in a in a cute package to make it palatable. Um, or is it that hopefully these firms have really um designed a version of Open Claw that is safer, that is uh defanged, and that there's some justification to the the cute and cuddly package?

07:34 What's what's your take on that, Sky? >> I'll I'll give the companies a bit of credit and argue they are probably going after a less technical crowd with this messaging and therefore the the the cartoonish uh design is intentional for that. Um, but I then that's probably where my um my credit to the company stops. And yes, they they are trying to repackage this.

07:58 It's not a coincidence that both companies released this within a week of each other um with a kind of cartoon themed um uh caricature. It is um the the perance uh that you mentioned it and the the concerns of open claw from not quite a year ago I think still remain but are less talked about when it's cute. Are you saying lobsters aren't cute? [laughter] >> They they've defanged the the claws.

08:25 Um I mean I I will say just uh from my point of view, I was struck by the positivity of the coverage from outlets that are often extremely skeptical of a meta. For instance, you know, we discussed on a previous episode that uh meta that uh the Ver's reviewer of MetaMuse uh was kind of blown away. The New York Times review also kind of blown away. And it does seem like a moment where um the capability of te technology is at least in sort of so wowing people that people are letting their guard down and um we'll see what

09:01 happens in a couple of weeks or months. I suppose [music] we were going to speak about Met uh Meta's Muse Spark which is Meta's flagship AI model. It's the first from Meta Super Intelligence Labs which is run by Alexander Wang who's also that's the the the lab that has put out MetaMuse. Um and they announced this week that uh mathematicians I think working for Meta went into just the meta.ai chat interface and produced six uh research papers.

09:35 This seemed like a huge deal to us on Tuesday. Then a day later um uh on I think Wednesday the 7th, OpenAI uh announced it released 722 math manuscripts on GitHub covering 372 so-called families of math results. So there's some on number theory and some on other things that I don't understand. Um there's a solution a claimed solution to the fourdimensional CA conjecture which again if you're a mathematician that means something to you.

10:07 Um now these uh we sort of knew this was coming because we covered a few weeks ago this thing called Navier Stokes and this announcement that uh an internal model that they were testing uh had solved this century open mathematics problem one of the millennium problems. So we knew that this model that they were testing internally was a big deal. We didn't expect however for 722 supposed solutions to drop uh in GitHub all of a sudden.

10:36 There is a technical story here, but there is a a story that is just as interesting, I think, which is kind of a social story, a story about a crisis sort of presented to the mathematics field in general. Kush, I'm going to start you with a really easy question. What is math? [laughter] >> It's so easy. Kush, please. Um, yeah, >> I think I mean, it's a way of thinking more than I mean, I've had this conversation with my kids as they're now growing up, right?

11:06 um what they learn as math isn't what math is, right? Um so it's not about uh fractions and multiplication and like all of that stuff. It's the thought process that goes into um like going into um some world that's defined by some axioms and postulates and being able to kind of um uh definitively show with no doubt that um some property holds or something is like true in in that space.

11:34 And um I think yeah I mean people like I didn't even realize this. I mean I'm supposedly a mathematician of some sort, right? And um until I got to like grad school like I didn't even realize that like what math research is or could be. So um uh I think that's the the thing. It's a journey. It's um like a journey of understanding. >> Uh Sky from your point of view if Kush says it's a journey of understanding must that understanding be human?

12:02 I believe some of the mathematicians I think Terrence Ta is leading that cause would say yes contributions to the math world need to be understood by humans um by definition. I don't know yet if I agree with that take but I think that's something that they they are trying to state that um some of these proofs come up are by contradiction and so they can say nope here's a counter example and perhaps the AI found that counter example by brute force whether it was by brute force by intuition we now know a counter example

12:34 exists and therefore that um solves the the previous conjecture um so in those cases it probably doesn't matter whether how the counter example was found um but it does exist and the math can move on afterwards. Uh but I think in some other cases I think I think one of the claims Terrence had has had hostage had had um cast onto AI was mathematicians have been asking all these questions and there seems to be good progress about you guys answering them but when will the when will AI generate the next math question and I

13:07 don't know if that happened in any of these 722 papers uh from last night or not but I thought that was a really interesting twist on on this approach is when will AI be able to ask the new math question, whereas previously it's solving questions that have been asked perhaps for decades. Um, and is and is answering them. Sorry, it's answering questions that have been asked for decades.

13:27 Uh, but I don't know yet if these last 722 examples are where AI are generating new mathematical questions. >> I'm really disappointed you did not read all 722 papers before Mixture of Experts. I was really on you >> to Can I Can I jump on that that joke there? the amount of work that has now been generated for kind of, you know, human bounded attention, right?

13:50 You know, we're kind of finite beings here. Um, how long will it take for mathematicians to sort through this? And what more progress can be made by AI during the time humans are trying to sort out this last last dump? I think I think that's this I think that's probably what I take away from this more necessarily than whether what is math uh no offense Kush, but rather just is this the beginning of a treadmill where humans sorting out AI generated content?

14:24 We've seen that play out in other places, but is this where this really takes hold in the mathematics field? um humans aren't going to be able to make sense of the amount of material uh coming coming out of the AI uh coming out of specifically mathematics of AI. Um yeah, maybe maybe it happened yesterday to Wednesday, October 7th, >> right? Exactly.

14:47 Yeah. No, and um actually even last week um there was an article that came out. It was um in quant magazine by um this person named Jordana Seploitz and um what she was writing is basically like the profession of ma of mathematics needs to transition almost into interpreters right I mean it's not the humans that are doing the math but they're interpreting it they're understanding it for the rest of humanity that what the AI did >> isn't that the notion behind any any science including including math right isn't that

15:20 there to actually make sense of nature and as you said like math is fundamentally there to to hold a particular property in a given space. So um is it really so bad if there is another tool another sort of virtual intelligent being that may help us uncover some fundamental aspects of the world around us and maybe even worlds that are so far away that we may not even have the reach right now because of the limited thinking.

15:47 Maybe our brains are not even devised to think beyond three or four dimensions and they can actually even compute across these different dimensions that are just very hard for our brains to even compute and even make sense of. [music] Let's move on to our third story. This one is about Hacker Ranks Chakra. It's an AI that runs a job interview. So, Hacker Rank is uh in my understanding a company that's been around almost 15 years.

16:18 It launched in 2012 and it's been a few different things over the years, but it's pivoted to AI like a a lot of people. And they've created this product called Chakra, and it runs coding interviews. And I feel like a lot of us know, even if we're not coders, that if you're applying for a coding job, um one of the things you're going to do at a minimum is a kind of take-home assignment.

16:41 There's going to be a skills test, right? and it's sort of like the hardest thing and if you're applying for a job at Google or IBM you're you're really drilled on this. So my understanding is that Shakra really focuses on that skills test essentially for coding and it has a product that has um you know they they describe it as an AI evaluating the the candidate which is true but I think one thing that seems clever about it to me from my poking around online is that so the the coder goes in the the candidate goes in

17:09 they do a coding task they're actually allowed to use AI within this window and that sort of helps avoid them cheating in a in a in another tab. Um, so Shakra is observing how they use the AI to do the coding task. Shakra can sort of pause them in the middle and say, "Well, why did you choose that instead of this?" And sort of it does these sort of probing interview questions that are kind of almost like you would get in a behavioral interview.

17:35 So on the face of it, it seems clever. Um, they over 6 months have had half a million uh beta interviews. It's uh Capge Gemini and other huge firms have used it. It's now publicly available. Let's start with you. What what uh impressed you or worried you about chakra? >> One of the things uh I'm I'm interested about is how many of those half a million interviews are truly unique interviews.

18:00 Maybe they are repeated sessions. And the reason I say that is because here one of the one of the things that AI may bring in is the um incomparable quality of of how it assesses the candidate. What if the candidates are given the AI to generate code and uh think through with AI? Um are they going to be judged by the quality of what AI how AI performs in that iteration because they may not have a complete control over what the AI output may look like and at the same time if AI decides to then probe further and get in a

18:34 different direction would we truly have a comparable matrix that we can judge each candidate's performance against? Now I'll grant that even human interviewers may give different interview questions and it's always going to be subjective evolution but um there's an AI element to it where you are also introducing another random factor into a candidates's ability to take a tackle any task um and so I I would want to see how it performs and how it can actually introduce uh a layer that humans don't have to go to and that

19:11 can truly delegate the responsibility away from the humans and give them a decision power by giving them the right rubric that they can judge a candidate against. Kush, running with this for a moment, um uh you head up a lot of our efforts uh to make technology be responsible and ethical. What are you thinking about in terms of the ethics of this, the risk of bias?

19:35 I'll just note a quote from one of the founders, the the CEO Vivec Ravi Shankar. He he says AI is way less biased than humans if you tune it properly and I'm curious what your reaction is to that. >> Yeah, I think these days um a lot of the vendors in this space um have taken bias seriously. Of course it can be done wrong. Um but uh uh I mean we've like even IBM has been evaluating different uh sort of hiring uh tool tooling and and so forth and uh I sit on the IBM uh responsible technology board and we've always um

20:07 asked to see um like additional reports on fairness metrics and these sort of things from different vendors. I mean not them specifically but um uh this does happen right and so I think this is a good thing. And then uh when you talk about bias, it's kind of like um looking at the situation from a very privileged sort of point of view. Um so the other sort of completely um like 180 turn of looking at this is it's actually expanding opportunities, right?

20:37 So um when you have humans are finite beings as Kai was saying, right? We can only interview so many people at a time. Um but like opening this up um allowing more people to interview for for jobs um is actually like a a better thing. Like it just can uh can be some person in some place um that's not well connected but still they're able to to apply for a job.

21:02 So kind of expanding opportunities is uh is an ethical it's it's very ethical to to actually do that. And um with the AI in there um again if the bias mitigations and other things have been implemented then uh it's it's generally I think a nice thing for the world. >> Kush and I have been talking about bias probably for a decade now. Um and I think I would just summarize it with a quip that uh machine bias is much easier to detect and correct than human bias.

21:34 So the all this talk about bias it I do believe it can be taken into account with enough time fine-tuning um unfortunately human bias is much harder to recognize and even harder to correct. So I I don't want to just always point the the the blame on on machine bias. Um but back to this particular technology I would say it almost gives off a bit dystopian vibes when used in a job interview sense.

21:58 Uh, but if you take the exact same technology and put it in a college or high school classroom now, I think you have a much better way at evaluating someone whether or not they've learned the content by how they interact through the setting uh than perhaps writing an essay. We know, you know, writing an essay is now uh pretty much, you know, no longer the standard for that uh with AI where it is.

22:20 So, it would be really interesting to see if Hacker Rank would be able to turn this into some sort of kind of pedagogical tool um alongside just the the interview approach um to see whether or not students are learning what they're supposed to be uh because they are demonstrating their knowledge by interacting with an AI chatbot rather than using an AI chatbot off to the side uh and then acing the test um follow-up.

22:46 So, I would really like to see it go in that type of direction. Um, if a company's goal is to find the best candidate purely by volume alone, then yes, this this tool might help. Um, although I I would like to think that, um, HR has a job to find slots and fill fill them rather than just interviewing everyone on the planet. The last question, I'll just put it to the group and anyone can chime in if they're eager, but I I do find interesting this notion that one of the skills we may be screening for and screening in in

23:15 the future is the ability to use AI well cleverly on the fly. And so the fact that chakra can detect uh AI use and not penalize it but maybe evaluate oh that is a good use of AI or that is a use here we see in this minute long clip that the person has revealed they have a really deep knowledge of data structures but they had this one gap and they leveraged AI in this crucial moment to plug the gap that that to me is compelling and I'm curious what you guys think about this possibility that one of the very things we may

23:46 be looking for in candidates increasingly is is good use of AI. >> This is exactly what we should be doing, right? I mean, using AI to uh to fill in our gaps, to expand our capabilities to um uh to to be able to do more and uh I think it is a skill that uh that needs to be both learned as as Sky was saying and evaluated as employers need because um we do need to to realize the the benefits of of this technology.

24:15 [snorts] I'll move us to our final story, a kind of bonus story. We're going to treat this as a little bit of an explainer segment. Um, where obviously we are a technical podcast. We are technical people talking to technical people, except there's a little bit of a footnote. I'm actually not a technical person. Uh, I just think technical people are cool.

24:39 And I do think that a lot of our listeners and hopefully increasingly more will tune into the podcast not only to nerd out on technical stuff but maybe to learn um and and kind of bring up their their understanding of technical things. So we're going to talk about a new model that came out. It is called Beam. It came out from a uh a firm called Reflection AI.

25:00 It's a startup founded by two Google former Google deep mind researchers. Um it is Nvidia has a lot of money in it. Um they let its two billion round uh about a year ago and um they released this thing beam and it's a model and it has a few uh jargony terms to describe it and we're going to run through a few of these things and what they mean. The first is that it is a a mixture of experts model.

25:28 Um and guess what? But that's the name of this podcast. And it's been a while since we uh when when this podcast was founded two years ago. This was a novel architecture. And I see my producers reminding me that we're that uh we're we're going to say from now on this is a show for both nerds and nerd groupies, people who sort of aspire to be nerds and hang out with them.

25:50 So for the nerd groupies out there, Kush and I collaborated on an article a couple months ago called what does AI look like? And I understood I began to understand for the first time in my life what a transformer was and and sort of a little bit of the history of neural nets. And I'm going to give my basic understanding of what a mixture of experts architecture is and you can set me straight.

26:11 But um both during the training when the model is sort of born and during so-called inference when the model is kind of processed the the architecture of the model is is uh concocted in such a way that unlike earlier models where every single number every single parameter every single weight uh in the model is sort of activated and used with each query you push through.

26:36 I believe when they when they train and then deploy a mixture of experts module they've kind of uh model uh uh for slip there they've kind of carved it into little brain lobes right and they've sort of created specialized regions that are experts on a given topic or or sort of style of cognition and then then I don't really know what happens does a router come through and decide which ones to activate for a given query but basically kush could you take my halfformed understanding and make it a little bit more fully

27:04 firmed for for for ner nerd groupies like me and Sky. >> Sure. Sure. Yeah. Um no, he's as much of a nerd as as all of us, but um yeah, I mean, David, when you wrote the article, right, I mean, you described the the parameters as lots of spreadsheets, right? Um and uh uh when we use spreadsheets um in Excel and these sort of things, there's also tabs, right?

27:28 Different sheets. Um so I think what we can think about is um the different experts the um are those different uh sheets or those different tabs within u like one given layer of that uh that overall spreadsheet. And so uh as you said right I mean um they're trained uh so that there's kind of differences in in them. Uh one of those tabs might be better at one thing another at another thing and so forth.

27:53 And as you exactly said, I mean there's a gating function or um kind of an orchestrator router sort of thing that um is also learned along the way, right? I mean nothing is set in advance. You don't say that this is an expert for um math and this is an expert for um being cute and and so forth, right? Um what you're kind of um like all of this is happening naturally.

28:16 It's not like you're you're trying to differentiate them. It's just it is learning in in that way. And then um because uh the number of activated parameters during inference is just one tab at a time rather than all of them. Uh the uh the inference speed is uh is much faster. >> In my effort to to move along the spectrum from mere nerd groupy to actual bonafide nerd.

28:38 I'm curious. So okay so you have you have this brain that is the model and then often when we talk about transformer architectures we talk about layers or little kind of slices of the brain that um as a a query sort of progresses through the model or even just one you know one word is is being processed through the model it's kind of like layers of cognition happening.

28:57 Now in the mixture of experts model are the experts are there is there an expert per layer or is this upstream or at a higher level than at the layer? >> Typically right now um there's like they actually are separate spreadsheets in a sense right so um all the way down um so you'll have uh like one subdivision of all of the layers another subdivision of all of the layers.

29:20 Um but uh and because of that you only need one router one gating function at the top. Um but uh it's not impossible. You could imagine an architecture in which um you would kind of switch between experts as you go down as well. So um uh yeah, I mean right now the general situation is that uh each of them are like kind of like siloed um like they're separate but um there's nothing preventing us from uh having other kinds of mechanisms and people have experimented with those.

29:52 Um I think uh Jeff Hinton had something called capsule networks um a few years ago and um even at here at IBM research we had kind of these routing neural networks that would um kind of uh do this as well. So yeah I mean I think there's um plenty of options. It's just what's popular these days um just as complete separation. >> Very helpful. So we've explained mixture of experts.

30:17 Let's try uh you know in the spirit of this uh explainer segment that we're experimenting with also getting a sense of what is meant by open weight. Why do we speak in terms of open weights in in model architecture rather than say open source this this term that we're more familiar with from the past. So I would say for the weights specifically when you do this big long pass you were talking about earlier um you know information propagating throughout the network at every step in there is a matrix multiplication and

30:47 you're taking input data you're multiplying it by a nearly fixed matrix and every one of those entries inside that matrix is referred to as a weight. So it's what allows you to take an input, multiply it by something, it becomes an output. That output becomes the input for the next layer. It's multiplied by another set of weights, it becomes the output, which is the input to another layer, and so forth.

31:10 So when someone says they're releasing the weights of the model, it allows you to go out and download, you can imagine, a huge number of very large matrices. It does not give you the ability to actually do the compute that separate. Just because you have the weights does not necessarily mean you have enough GPUs to actually run the model. It also does not tell you what data the model was trained on.

31:33 And I think that's another big caveat between the difference of an open weight model versus more kind of an open source setting where you'll be able to say here is the exact data that we used to arrive at these weights we're providing you. If you just say open weight, you're not necessarily releasing the data it was trained on. you're not providing the compute to actually execute all of those matrix multiplications that those weights are informing.

31:59 So I think those are probably the ones that really are the largest caveat. Um the they don't necessarily provide the data or the compute um that let them arrive at the weights. Uh but they hand off the weights and then somebody who is has access access to compute would be able to then pick up and run from there. Uh so it's it's not a terrible thing but it's not necessarily fully open source either.

32:24 I think with the largest distinction is um the uh people who are providing the weights don't also provide the data that the model was trained on. Okay. So this is something that I'm just learning kind of this week actually even as someone who prided myself on trying to understand this. I I used to think that we called open- source software open source because it was source code.

32:44 it was code and we called models open weight because they don't they're not made up of made of code but they're rather made up of these numbers that are learned through this heavy computational process but actually this week I feel like I'm glimmering something more complex which is that in the AI model world there are degrees of openness if you just release the numbers the results of your training process but nothing else that may be said to be an open weight model but if in addition you are transparent about the

33:14 recipe that you used to learn the weights about the the training data if you're if you dump on GitHub this is what we trained it on etc. If you're if you're more transparent not only about like the the the actual neurons of your model but the process by which the model is is grown, we begin to call that an open-source model which is something even more expansive than open weight.

33:35 Did I summarize that correctly, Kush? >> Yeah, absolutely. Yep, that's exactly it. And um yeah, so there's the the recipe as you said, the the training code itself, the the source code of what you use to to get to those weights um would also be part of an open- source fully open source release according to the definition that uh some people are very like kind of sticklers about.

33:58 So So yeah, >> this is clarifying to me. I feel like on on the sort of meter, I've I've ticked just like a little bit closer to to nerds. So thank you so much to my guests. Uh that's all the time we have for today. Um, so we started by talking about an AI that works while you sleep [music] and we ended up talking about one that only wakes up the parts that it needs.

34:19 Uh, thank you to Utsov to Skyler Tush and thank you all for listening. You can find us on Apple Podcasts, on Spotify, on YouTube, wherever you get your podcasts and [music] we'll see you next week on Mixture of Experts. [music]