Grok 4.6 has caught up with frontier models on some benchmarks and agentic capabilities, but it is not clearly worth adopting over Grok 4.5 because it is slower and more expensive; its main value is as a preview of a potentially much stronger Grok 4.7.
🔒 12 more in the full analysis
🔒 14 more in the full analysis
Searchable transcript of xAI just caught up (Grok 4.6 is here) — Theo - t3․gg (25:31). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Seems like Xi is on a bit of a tear lately since they acquired Cursor. They've been moving so much faster. We went from a drought where they just kind of didn't put out models that mattered or really at all for months at a time to getting a bunch back to back. We just had Croc 45 not long ago and it really impressed me. It's not the model I default to for my everyday work, but it's the first model that's good enough from a lab that isn't anthropic or open AI that I could see myself actually daily driving it.
00:23 And now Grock 46 is here and they've went even further with it. With everything I've seen and all the playing I've been doing so far, it seems incredibly promising. It's fast, it's cheap, it's reliable, and it's good at doing long running tasks with lots of sub agents and not losing track of what it's supposed to be getting done. This is important because as the way we build with agents changes, the ability for agents to go off for a long time and solve hard problems and come back with working results is more important
00:49 than ever. Sadly, that doesn't mean this model's perfect, though. There are rough edges as there are with everything and I've already encountered a handful of those with my own testing. There's also some big bright sides to this as well that I've been experiencing the more I've been using the official Grock Build CLI which is awesome by the way. I actually really like the CLI.
01:06 There's layers to this one. There are some really crazy benchmarks which uh spoiler Gro 46 is now neck andneck with GPT56 Soul according to artificial analysis. We need to see if this holds up because if it's actually getting that close at a fifth the price, things are changing quickly. While Intelligence seems to be getting cheaper, my payroll isn't.
01:28 So, I hope you pardon me for a quick break for today's sponsor. I have a real quick question for you. If you're building an app and you want to add video to it, how are you going to do that? Whether it's a video on the homepage that you want to use to market your site or userfacing uploads where they can upload whatever. How do you know that these videos are being served correctly, that they're being processed properly, that they don't have inappropriate content?
01:48 And what happens when you want to do things like generate captions or get more info about the video? Well, I hope you would stumble on today's sponsor, MX, before you get too far, because they solve all of those problems and more. Whether you're just adding one video to a small homepage on a side project, or you're trying to build a service that competes with YouTube, there's a reason everyone picks MX.
02:07 Whether it's Fox, Band Camp, or Robin Hood, or even companies like Cursor or Dropbox, everyone leans on these guys for a reason. It's cuz they get video. I've built a lot of video services, and if I learned anything in that time, it's that processing and managing user generated content is just a huge rabbit hole you want to avoid to the best of your ability.
02:24 And that's why I'm so thankful for the new MX Robots product where you can do all of the automation you would ever need to trivially. You can set up Muk to answer questions about videos as they're uploaded using AI, which is super powerful by itself. But when you combine that with the captions and the moderation, as well as a summary feature, there is so much stuff you can do without having to build all of the details yourself.
02:44 And once you're processing that video, it gets really expensive really fast. But when you use a service like MX, you know you're getting the best possible deal. I could tell you all about how great the code is and how easy it is to integrate, but honestly, you should ask your agent. it already knows about MX because these guys have been the default for a long long time.
03:01 There's a reason I've been building a MX for eight years and you can figure out why at sidv.link/mucks. Let's start from the official announcement and then go into the real world use cases, the benchmarks and my own experience building with it. Introducing Gro 46. Gro 46 builds on Gro 4.5 with a particular focus on longrunning agents and more ambitious interactive and visual work.
03:20 translated to actual English quickly. Grock 4 6 isn't a new pre-training. It is a new post-training on Grock 45 because part of acquiring cursor is getting all of cursor's crazy RL and post-training stuff. And now they have all of this post- training and RL stuff. They're able to improve the model after the pre-training is done. And this is the technique that makes these longunning tasks and agentic work stuff much more reliable.
03:47 That's usually done in post- training, not pre-training. As they mentioned before, 46 is building on 45 with a particular focus on longrunning agents and more ambitious interactive and visual work. It stays with complex tasks across many steps. Whether it's researching topics, analyzing information, working across code bases, or turning ideas into polished apps or work artifacts.
04:04 As I mentioned before on the artificial analysis intelligence index, it is neck andneck with 56 soul and right behind fable 5 and opus 5. I am obligated to say that this shows that these numbers aren't the best way to measure how smart or useful a model actually is because for my experience, Opus 5 is significantly worse to use than Fable and Saul. It looked and felt really good day one, but the more I started merging the code it wrote, the more problems I ran into, the more cleaning up after the Opus sloppus as I now
04:34 call it. Yeah, Opus 5 seemed like it was good and by every measure seemed to be solid and the more I use it, the more I hate it and I'm just on Fable all the time. So again, this bench isn't the most trustworthy thing, but there are others in here that are good, like deep sui, where it is just behind soul and fable and meaningfully ahead of Grock 45.
04:52 Like this is a huge bump for a post-training improvement to go from a 54% to a 66 is huge. Cursor bench saw a smaller jump, but it did jump Grock 46 past 56 soul, but remember they accidentally trained some of Cursor Bench data into the models. I think they might have cleaned it up with cursor bench 32. Not sure, but you get the idea. And then with Frontier Code, they saw a meaningful jump as well from 56.6 to 61.3, putting it smack dab in the middle between Soul and Fable.
05:21 It's available now in cursor and Grock build. They've offered 2x usage inside of your Grock build and cursor subs for the first week, so you can start trying it immediately as I have been doing. I am just on my subscription on Twitter right now because I already have the like blue check on X because they pay me for it. So, I'm on the the pro tier, the one that kills ads as well.
05:40 X premium pricing. Let me see what the costs are for the different tiers. Yeah, I have the $40 tier because I don't like ads and it pays for itself because I'm paid to use Twitter, as stupid as that is. So, I've just been using Gro 46 through the included data I get with that subscription. And I've used it to generate a handful of games and also do some real work in T3 code.
06:01 Specifically trying to get things to parody with both cursor and gro in T3 code because we do have the Grock bindings built in. Shout out to the Grock build team for working with us on that. Really proud that we got that in as quick as we did. So I was working on tidying it up using Grock and so far pretty solid. Show you guys everything I did with it after we get through the official information.
06:22 Grock 46 underwent a longer supplemental training run than 45 did with curated model generated data for reasoning and advanced technical concepts, highquality engineering data, and an improved optimizer and training recipe produce a stronger foundation for the SFT and RL stages that followed. When they used Gro 45 to regenerate the SFT trajectories across reasoning efforts, agent harnesses in domains such as STEM, software engineering, and knowledge work as well as filtering out problematic traces with modelbased
06:49 checks. The resulting SFT checkpoint shows strong performance and improved behavior. 46 is trained on a wide range of agentic RL tasks including knowledge work, general coding, and domain specific environments for kernel optimization, webdev, computerated design, and more. I'll show you guys some Grock 46 designs, don't worry. They tested Grock 46 on projects designed to stretch its range and ability to sustain work over many steps.
07:12 The model's especially strong at turning a broad product idea into working first versions. It can research unfamiliar domains, structure the app, implement the core interactions, and continue refining the result through several rounds of feedback. They've also seen it doing more self tests and verifications on longer trajectory runs with the model checking its own work before moving on.
07:30 Good. That's been an issue I had with grock models is they would just write the whole thing and it would work sometimes and then fail other times. Sadly, that was not really my experience with the game gen stuff I was doing, which I will show you all momentarily. They added more safeguards due to new safety capabilities and potential to hack and whatnot.
07:48 Safeguard evaluation work reflects 46 expanded capabilities with our widest ever suite of pre-eployment testing for capabilities and safeguard calibration as well as extensive post- deployment and third-party testing. Let's see how this goes. Do a thorough security audit of this project. Are we ready to launch? What should we fix first? We'll see how it does with a thorough security audit in a real codebase.
08:14 I have it pointed at my cloud that I'm still hopefully getting out soon. Lake bed. They show a chunk of eval here which we'll look at while I wait for that to run. Intelligence inex we already covered. GDP val I don't care about. Cursor bench it's still behind fable but it's now ahead of 56 soul which is crazy. Deep suite it's getting up there. It's neck andneck with soul and fable.
08:32 Actually it's a little bit more behind there but still huge improvement for 45. It got bestin-class in AA briefcase and Harvey Lab vals, which are interesting. Also, the best in GDP val, but again, like I don't care. It's GDP val. SpaceXI's Gro 4.6 scores a 61 on the artificial analysis intelligence index, joining the frontier in line with 56 soul with standout agentic performance at lower costs.
08:56 46 gains five points over 45 in the intelligence index just over one month after its release. That's a five-point improvement in a few months and 23 points compared to 43, which is really not that long ago. The pace of releases from these guys has been nuts lately. And it looks like that pace will maintain because Elon's already tweeting saying that 4.7 is significantly better than 4.6 and it should be ready in 3 to 4 weeks.
09:18 The initial training is complete and they're adding a massive amount of SpaceX company data in supplementary training and they're expecting it to be quote something special. Very interesting. This is the jump where they are now in the frontier range. comparable to the other frontier models according to artificial analysis. Much stronger agentic performance especially with things like GDP val frontier level intelligence at lower costs because again it's way cheaper.
09:42 It's $2 per mill in and $6 per mill out which is 60% below Opus 5's price and even more than 56 souls price costs 84 cents per task which is similar to Kimmy K3 with slightly higher intelligence which gives it a really good place on the Pareto Frontier. The cost versus how smart it is. Cost per task is right there with Kimmy K3. Slightly more expensive than Gemini 36 Flash and Terra.
10:08 Slightly less expensive than 56 Soul. Comically cheaper than Opus 5 and Fable 5 for realworld use. That is $3.14 for Fable and 84 cents for Gro 46. This is more than double the price that Grock 45 was though, which is the important detail they seem to be hiding. I'm guessing this price gap is a token difference between 46 and 45, which we can confirm here.
10:28 Yep. I was so pumped that Grock 45 was as efficient as it was. Seeing them lose that with 46 is a little sad. They've increased the number of tokens per run by over 30% here. Previously, they were one of the few major providers that had a model that was more token efficient than 56 Soul. Now, they're not. Now they are still relatively efficient, like comparable or better than Fable efficiency-wise, but it's no longer that magic efficient small fast cheap model that I wanted it to be.
11:00 This is something I've actually felt in my own usage because I noticed that the Gro 45 test when I did them ran almost immediately, like creepily fast. 46 is back to the usual send off the prompt and then go do something else that I'm used to from Frontier. They talk about their AA briefcase bench where it is now Fable 5 tier. Context window is the same at 500k tokens.
11:21 The price is the same as well. They did increase the price of cash reads. Previously it was only 3 cents per million tokens red. Now it's 5 cents per million tokens red for the cash. That might also be affecting price here. Some of these jumps are nuts. Like I know I don't like GDP Val, but this is one of the biggest jumps I've ever seen in that benchmark.
11:42 Like that's crazy. It's clearly much better at these long running things, but it also thinks a lot more, too. And again, as I said, it is now much more expensive, but it's also slightly smarter. It's got a really good score here in the cost versus intelligence, but this does yank them out of this golden area of like the green corner where it is the cheapest and the smartest and forces it out over to be on the right side of the median for cost on the log scale, which is sad.
12:10 They're clearly trying to rush their way to the frontier. So now it's time to see how it does with UI. I have other things I've been working on with it in the background, too. But I am very curious how its design work is because it's a different model, and it is nice to see the different ways different models handle design. As I mentioned in my Muse Spark video, I was actually surprised at how different the Muse models felt for design work than the usual anthropic and OpenAI slop.
12:35 I've experienced less of that with Grock models where they feel more like a dumbed down version design-wise. So, let's see what they cooked. I'll start with none of the skills on and just go through these Grock 46 outputs. Shout out to Dra for making the site. By the way, the witchai.dev is super handy for this type of video. Here's the first design that it made.
12:55 I don't like the noise pattern, and I really don't like the way that the text is so sharp with the blurryish background. It just doesn't feel cohesive. This feels like old era AI Slavic. This is GPT 5.0 is how this feels to me. The purple in particular is This one has a nice little mock app in the corner, which is fine. The coloring is okay. Too many cards.
13:17 Too old school like Tailwind Templaty. This is the usual brutalism that we see from a lot of these. And this is a boring centered page. Yeah, none of these impress me. I did happen to notice that when you use the design skill from Claude, it can and do better. Some of them are sinful and disgusting like whatever the this is. Some of them have okay ideas on them like this with the floating card things when you hover.
13:43 Yeah, could be tidied up. This is sinful. This is sinful. Not digging this. And just for reference, here is Fable doing similar designs. Much much better, like comically. So, I hate that the terminal ones, but other than that, or even like Opus does a much better job here with these, or even 56 Soul, as cringe as it often is, is meaningfully better at this type of thing.
14:12 Like, this is a much nicer version of this style of page. So, not great at design at all. I will show the real work I was doing with it earlier, but I first want to play with these fish slop ports because they're fun and I'm curious how it did. Those who aren't familiar, I made a game at the end of last year with Opus called Fish Slop with Opus 45. Never got very far with it.
14:33 It was meant to be an insane aquarium clone. So, I like throwing models at the old codebase and say, "Rebuild this. You can reuse assets or whatever, but from scratch, make a new version of this game using that codebase as a reference." And it's a fun way to see how well it like understands a lot of end toend work in a complex thing like the relations between systems in a game.
14:51 I'm already noticing that like the controls are uh bad. The sizing of elements is all wrong. Like these fish pellets are too big and the fish are too small. The fact that the thing you had highlighted on the bottom here stays highlighted when you move off it is really bad. Yeah, the movement just feels so not great. The pace of the game seems wrong, too, where things don't like grow fast enough and you don't get to the next stage quick enough.
15:19 Yeah, this just doesn't feel great. This is like Kimmy K3 did a meaningfully better job with this than Gro 46 seems to be. But this is the easy version. The hard version is 3D. This is what I got when I first opened the 3D version of the game. A black screen. This is the first new model I have pointed at the task of a make the game 3D and had it just outright fail.
15:48 All the others might have had bugs and issues with like the movement or the mechanics or the models they created in the 3D space. All of them opened with a 3D environment. Even like GBT4 was able to get that far. This is the first one that just outright failed to do the 3D port. I did tell the model that it failed. I even gave it a picture which showed me how nice the rendering of pictures and things is inside of the Gro CLI, which like the Grock build CLI is actually genuinely really nice.
16:16 Well, I told it to fix it with the screenshot. It appears that it did. So, let's dive in and see. Oh god, this is so cringe. They got the directions wrong. So, up and down is right, but left and right are inverted. I don't know how you get right and left inverted, but they did. The placement of everything at the ground is entirely wrong. Like, comically so.
16:35 God, this is very bad and broken. [clears throat] The models for the fish are This is the worst 3D pass I've seen on this in a long time. Actually, this is like last year I would have expected quality like this. Just for reference, here's the version that Muse12 made for way cheaper. It got some directional issues like up and down tilt it, but like it functions.
17:07 The movement's way better. The core mechanics work a lot better. If we want to go to real frontier, like I don't know, uh, Kimmy, look at how much further along this is. The modeling of the things on the ground. The lighting's a little broken, but like this is a generational gap. This is an open weight model, by the way. Like, you can go download the weights online and have a thing that is this much better at 3D.
17:34 And the fish model is so much better. It's actually the best 3D modeling I've seen any of the models do so far. So yeah, there's your comparison. Grock is not even unimpressive. It's like last generation. It's so far ahead here. But then there is realworld work. And for this I was quite kind to Gro 45. I found it surprisingly capable of doing like real work where it had to touch different things that were complex and stay on task for longer running things.
18:05 Well, take a look at the security audit it just did. It found a previous audit that I did in the git history. So, that's cool. So, it's already strong. The O broker, the worker isolation storage control plane will get us burned. Give us advice for locking endpoints. Put it on the public suffix list. I already was planning on that. like bed chrome on every capsule page.
18:26 Finish the admin trust and safety UI. Stop putting identity tokens in URLs or a wider launch. Refuse sandbox off flags in production. Yeah, did a decent audit here. Didn't get me any errors or problems when it did it. There are other deeper things it didn't find, but that's acceptable. This is the more interesting one I wanted to take a look at with y'all.
18:46 I asked Grock 46 to take a look at how we have T3 code implementing cursor because the current build is using the outofdate ACP adapter for the cursor CLI. They want us to move to the SDK. I believe Julius is work on the new orchestrator includes that. But I wanted to see what it would do. So I asked it to go through and audit things and see how we would migrate from ACP to the Cursor SDK.
19:08 For long lived gooey hosts that need a real model catalog, resume, cancel images, and usage. The SDK is the one that cursor is actually maintaining. Yep. So why is ACP the wrong host contract? Then built in fallbacks. Yeah, these are all real problems that we've had with the current cursor bindings. So it was correct there. Official docs license anywhere to not MIT.
19:27 That is annoying, but that's fine. The SDK doesn't offer guey level approve or deny. That's annoying, but we auto run anyways for most things. Blocking ask questions. Uh the SDK will not allow that. and reusing agent login. That will not work. So, we have to implement our own login, which is annoying. You know what I will do? I don't feel like reading gro slop.
19:50 We'll have 56 soul share its thoughts momentarily. Plan approval could not be completed because the client disconnected. Plan mode remains active. Okay, so it put itself in plan mode. That is obnoxious. Did it put the plan in here anywhere? It didn't. That is very annoying. Yeah, the Grock implementation needs a little bit of work as well. I will put some time in in the near future.
20:10 I have to go turn on the legacy plan mode for this. While those plans are being audited, I'll show off a little bit of it behaving how I wanted it to. Here I asked it, what gaps exist in T3 Code's implementation of Grock build compared to other harnesses. It went through it made a plan. I don't know where it went in the UI. There there's some weirdness in the events that the Grock build like ACP sends out.
20:30 So, we have to put more work into how we clean those events up. But I asked at the end here, I wanted to see it differently. So I just said HTML which triggers my HTML skill and it responded there. The UI collapsed it again because their events are broken. Uh I gave it some feedback on the plan. Had it update then asked it to build the whole plan file a PR and babysit which it did.
20:55 I can go open it on GitHub and we can see what things thought here. We got a bunch of feedback from Macroscope on the changes. It tore things to shreds here, but it seems like it followed my instructions really well. You'll see here that it used my account to reply to things, but I have a little call out in a skill I made for PR comments where I tell every agent to open its comments with note which model is responding on behalf of Theo.
21:18 This is not a small change either. This is a thousandline PR that it made to address a bunch of different gaps in our coverage for Grock Build. And I told it after it makes the PR to babysit it once it was filed. Then I told it to do one of my favorite things, which is to tell the agent to look through my actual history on my actual computer and find gaps in what events occurred and what we actually process.
21:44 Just told it to make a separate PR stacked on the one that we currently have to do all these additional changes. This is a complex issue because it needs to know how to like figure out and stack PRs. It needs to keep track of the gap between the old PR and the new one. All the investigatory work it just did in the context of how it applies with my history on this machine as well as all the other skills and things that it pulled in.
22:07 It's not easy for models to deal with all of this like competing context and stay on track. This is one of the things again I thought Grock 45 did surprisingly well. It was able to take multi-step unrelated work and do it all cohesively and coherently in one thread. So, we shall see how it handles this here. I just put together a really rough like how I would score Grock compared to Fable and Soul for like different categories.
22:33 I think about things like cost where Fable's way too expensive. Soul's significantly better. Probably do a little lower there, but like reasonably priced. And then Grock 45 way better on cost. Intelligence, Fable is the best right now. Soul is surprisingly good, but not quite as good. Grock is not even this. I would put a little lower there. Then you have speed where fable is very slow.
22:53 Soul is meaningfully faster simply because it's so much more efficient. And then Grock 45, especially on the fast mode, flew. It was super fast and really nice. Then with thoroughess, Fable, I find, isn't quite as thorough as it should be. It misses things here and there, but it's it's thoughtful, not thorough. Soul is incredibly thorough. It checks every single thing.
23:13 It touches every single edge, and it writes too much code as a result. 45 just did not have any of that. And then with orchestration capabilities, Fable was the best. Soul is surprisingly close in Grock 45. Better than most, but still not quite there. So, how does this all compare to our new model Grock 46? Sadly, cost is a regression. I'd say this is like a sixish now.
23:34 Intelligence seems like a meaningful bump. I haven't seen too much of it being way smarter, but I can confidently say it's probably meaningfully better. We'll give it a 6.5 there. Speed is where I'm seeing one of the biggest regressions now where I would put it at like a five and a half at best right now. Thoroughess, it is better, but it still misses things.
23:53 I'll bump it slightly there. In orchestration capabilities, I did see it using sub agents decently well. I'll give it a 6.5 here. Why not? The issue is the only things I would consider Grock previously to be a leader in were the speed and the cost. And we saw regressions in both of those categories with this release. The cost went up, not per token, but since the token efficiency went down, it's now more expensive.
24:19 And the speed went down because it's less token efficient. So, it's generating more tokens and it takes longer. These two changes make Grock much, much less interesting to me. But the speed at which the intelligence is growing suggests that Gro 47 will do a lot more in all of these other categories to catch up. The thing I liked Grock 45 for wasn't that it was as good as Frontier at specific things.
24:45 It's that the speed and the cost were far enough away from Frontier that it felt uniquely useful in those ways. And I feel like we are losing that with Grock 46 a little bit. I think that summarizes my thoughts on this release. Is it my favorite model? No. Am I going to use it a whole lot after this video? Probably not. But does it have me excited for Grock 4.7?
25:04 Absolutely yes. It makes a lot of sense why Elon's talking more about 47 than 46 right now. The future seems very clear and it seems like a future where Grock catches up very fast and I'm honestly pretty excited for that. We need more competition. We need more good models and we need more people fighting to make them cheaper. Let me know how y'all feel. Is this an exciting release for you or are you going to just ignore it? Let me know in the comments. And until next time, peace nerds.