🔒 15 more in the full analysis
🔒 7 more in the full analysis
Use an inexpensive, multimodal coding model connected to GitHub to audit large groups of open pull requests, identify overlap and resolved issues, assess merge risk and confidence, rank the easiest actions, and produce an actionable report.
Full plans for 1 idea. Inquire for details →
Searchable transcript of Ox Alpha is INSANE — Theo - t3․gg (43:15). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Last week, there was a really interesting model drop and the more I've been using it, the more impressed I've been. The model came out anonymously under the name Ox Alpha, and it was announced by Open Code and Open Router, which was interesting in and of itself because we don't know who made it, but there was one other detail that was even more interesting.
00:17 We have capacity for 100 trillion tokens per day. That is crazy. That is so much compute, it's hard to even fathom. On my absolute heaviest usage days, I'm doing upwards of like 7 billion tokens. All of Google's Gemini traffic in a given day is estimated in the 150 trillion token range. So for whoever was hosting Ox Alpha to give this to one of a few providers in the trillions, hundred trillions range per day is just hard to fathom.
00:54 So either there is some hidden reserve of compute these guys have or maybe this model's really small and cheap and even then that's still a ton of comput. So then we started to get some numbers. You guys might know Ben Davis. He's been in a bunch of my content. He does the podcast with me also the channel a ton. Good friend. Him and I were playing with it and he started benching it against the subset of deep suite problems that were public.
01:16 It ended up scoring absurdly well. 56 soul got a 52%. People five got a 65%. And whatever the hell Ox was somehow got an 80%. Something is up. Well, guess what? We actually know what the model is now. And it's definitely not what I expected when I first started playing with it. I will also say that the numbers being shared here are misleading, but it doesn't mean the model is bad.
01:42 In fact, Ox Alpha has quickly become one of my most used models, and I really, really like using it. But I bet you wouldn't have guessed it's a flash model. GLM 5.3 Flash. It behaves better than 53. It behaves on par with models like Opus 4.8 at a tenth to a hundth the price. This is a wild model launch and there are so many fun details for us to dive into.
02:07 I can't wait to show you everything that makes this model cool and all the things I've been using it for after a real quick break for today's sponsor. As models get smarter and companies ship more and more code, the risk around security is only getting bigger. Making sure your apps are written in a way where they are secure and safe is more important now than ever.
02:23 But it's also harder than ever. And if you ask models like Fable and Soul to come in and do this audit for you, they're going to throw a bunch of errors because the labs won't let you do security work with them. If only there was a tool that understood your codebase that could actually give you real useful feedback on what was insecure. You know, something like Code Rabbit but for security.
02:40 Oh yeah, Code Rabbit made a security product and it's actually really really cool. Not only does it do security reviews for every pull request as they come in, as well as checking for dependencies and other things that your project relies on that might be insecure, they also introduced a deep audit product that will go through your whole codebase end to end and find anything that might be going wrong.
03:00 I decided to run it across T3 code and I was so impressed with the results, I ran it against everything else that I care about as well. I'm blurring these because these are real security issues we're going to have to go fix. But I can let you guys see the medium and low ones because thankfully they have a little toggle here where I can change what you're seeing.
03:14 Here's a security misconfiguration in my project. I can click here and open it and see exactly what is wrong, what files it is in, how it is reached, and most importantly, I can click the little button that my face is currently covering. Fix with AI. This button will go file a pull request on your real project for you. Since I'm maintaining so many projects, it's great to see the security trends as well, making it easy to know how secure projects are over time.
03:38 I'm really excited to hopefully see these numbers go down as we get to use the product more. Start your free trial and get your first 10 scans free at soyv.link/codrabbit. As I mentioned before, this model came out under the name oxal alpha originally and I and others got to heavily heavily use it in that time and I was beyond impressed. The model has a million token context window.
03:58 It has full multimodality so it can take images, audio and video which is crazy especially considering that 5.3 non flash couldn't even take a screenshot. It is mindboglingly cool. So I obviously had to play with it and since I'm mostly using codeex and cloud code via T3 code I didn't feel like trying to get it working in other things and I also was told by open router that this was worth trying for agent stuff.
04:22 So I started setting up the open router binding inside of codeex through my weird proxy layer for ox alpha and I was immediately really impressed when I started using it. Took a bit of back and forth but I now have 53 flash working inside of T3 code. One of my favorite first tasks when I get access to a new model like this is to ask it to just go over the work I already have going.
04:42 So, let's give that a shot. Take a look at all the poll requests I currently have open in this repo. Help me prioritize them based on ease of merge and the value that they provide to our users. Also, thanks to the cast, I've been voice to texting a hell of a lot more. Thank you, Whisper Flow. Good start using the file system MCP resource read tool and failing from that.
05:03 Now, it's going back to just using commands. This is what happens when your models are rldled on the specific behaviors of a harness. This model clearly isn't RL on the weird stuff in codeex, but it still works totally fine in it, which is what's been really impressive for me here. It's calling the GitHub fetch in order to get all of this info of the PRs I currently have open so that it's able to go through them.
05:23 But everything you're seeing here is possible with various different models that are currently available. So, let's push it a little bit more. Also, love the GitHub error connection issues. As always, GitHub's APIs being garbage. I'm going to ask you to do things a little bit different here. Ignore all the pull requests that were made by Codeex. It should be indicated in the body.
05:43 I only want you to go through the ones that were made by Fable in Cloud Code. This is the type of thing that throws a lot of models off. If you give them a steer in a different direction when they're deep in a a gentic task like this, they'll get lost. But you can go even further with this. You can ask it to do something else with the PRs once it finds them, like leave a comment on the PRs that you care about or what I'll do here, put your findings in HTML so that I can more easily look at them on my phone.
06:14 This is now going to go in as additional steering context and I cannot tell you how many times I sent something like this to another model especially like openweight agentic models and it would respond immediately to that saying okay I'll let you know when the findings are done so I can put an HTML and then send a stop signal and then not keep going until you say hey you should finish the work you were just doing.
06:35 I'll narrow the portfolio to open PRS who subscriptions credit fable and quad code excluding Codex builds and it's already getting to work doing just that. Something is causing the skill file reads to fail again. I've been hacking the hell out of my codeex and I'm suffering for it. This demo would have been a lot cooler if GitHub wasn't rate limiting us and codeex's change to how skills are read wasn't breaking things so badly.
06:56 I could sit here wait forever, but I ended up finding a run I did when I first got access to the model 5 days ago and I was very very impressed. Here, I didn't just ask it to audit my PRs. I asked it to audit all of the PRs in T3 code that have been updated in the last 5 days. I told to use lots of sub aents to break up the work. What PRs does it overlap with?
07:14 What issues does it fix? How risky is it to merge? How confident are you that it's mergeable? Should we close it? Rank them by ease of action. It successfully spun up and ran six sub aents. Three failed because of API call issues, but it was successfully able to recover them and keep going. Yeah, the first worker wave hit a model off failure. Again, I am hacking the hell out of codeex to make all this stuff work.
07:34 But once it was fixed, it was fine. And it realized this though. This is probably the coolest part. It realized that the sub agent off failures were consistently happening. So it fell back to just audit directly rather than waiting on the broken delegation. I'll use the GitHub inspection skill and machine generated summaries, then manually review the highest risk and easiest action PRs.
07:51 This is phenomenal. This is really, really good debugging work that the agent is doing. Not to debug things in my codebase, but to debug its own flows. The agent can unblock itself. Audit complete. The repo is unchanged. Sub agent delegation failed with provider off error. So, I completed it directly from the GitHub metadata, diffs, reviews, CI comments, and local source context.
08:12 It wrote up the report. It gave me a breakdown of what should be, how much should be merged now, how much should be merged after a better review, yada yada. And this was for every PR in the last 5 days. There was hundreds. And it got through all of this in like 20 minutes. It gave me the links to the easiest to merge. And this is a silly little thing.
08:31 But I cannot tell you how many models, even Frontier models, don't make these links. They make them just the PR number, which is obnoxious. And I have to tell it, hey, can you redo that output, but with links for the PRs instead? This model just did it. Small thing, but like these little things you start to notice a lot as you do more and more of this work.
08:49 I ended up going through and merging all eight of these because they were actually very easy. It did a great job of identifying PRs that were easy merges and letting me go do them. And this is all for free. Remember, this model for free, made really good outputs, was able to audit hundreds of PRs, find a pile of things that were easy merges worth doing quickly, and I went through and did it.
09:15 It was awesome. I did not yet know this was a flash model, but god damn, it was very nice to work with. I followed up. Merge pretty much everything you said was an easy merge. Do another pass with those changes in. Let me know what should be closed, what is still worth merging, etc. Make sure all issues that we have addressed are listed as well. Ignore draft PRs, please.
09:33 I don't care about those. Scope is the same 5-day cohort with drafts ignored. After your merges, 352 non-draft PRs remain. It noticed the issues that were fixed by those PRs and closed them. The other merge PRs had no link GitHub issues. No open PR claims 5276, 5289 or 7233. Still worth merging. Another pile of things that it thought I should merge.
09:53 And then a bunch of things I thought I should close. Very genuinely good outputs. Oh, it finally got my skill read to work. We have successfully got it into T3 code now. Hurrah. And that's using my ZAI plan too, by the way, which I somehow managed to get set up in Viroxy. So now the I think it's 80 bucks a month I was paying from the legacy V2 ZI code plan and I believe I should get pretty close to unlimited inference there.
10:19 I promise you guys this model was much faster when I was testing it last week. I'll look into what's making it slow. We could also just look at Open Router because that will reveal other very fun things. That's Quen 38 Flash. No, didn't know Quen 38 Flash also dropped today. That's hilarious. Here we are in Open Router. We can take a look at the speeds.
10:34 They did release it as open weight already and it looks like companies like base 10 are slaughtering with that getting as high as 114 tokens per second on this model. The latency for ZI has spiked a ton. It's 4 minutes and 40 seconds for first token now. Correction 4.4 seconds for the time to first token. My bad. Seems like a lot of people are taking advantage of how cheap this model is.
10:58 So, it's down to 33TPS. It was a lot higher when I was using it before and I'm sure it will recover and other places will host it and make it even faster. I might just point everything at the base 10 version for now. But also the price discount you get on ZI as well as a few other providers is insane. This model with its current discount is 7.5 cents per million tokens in and 25 cents per million tokens out.
11:22 25 cents per million out is insane. This is like Gemini Flash 2 pricing. Ox Alpha was free, which obviously is even better, but to have the free model taken away and replaced with the final real version here, 53 Flash, and have it be this absurdly cheap, this this rounds out to free for a lot of people. And it's tiny. The reason it's so cheap to run is because it's not very big.
11:49 This model is only 320 billion total params and just 18 bill active for the expert that it's working with at any given time. For reference, Kimmy K3 was 3 trillion params with a T. This model is a tenth the size of Kimmy K3 with comparable performance and most importantly in my opinion it just behaves really well. And this is where I want to get into an important separation before we dive into benchmarks.
12:19 I find that people think of model capabilities either in one track or far too many tracks. I don't know how to explain what I mean here, so I'm just going to use examples. Here's some examples of the types of things I hear where I put you into one of the buckets of not really understanding how model capabilities work. Most common I see is people just always using the highest model setting and the highest reasoning option.
12:42 This one sucked for us with T3 chat. When you give a selector like this for low, medium, high, XH high, the worst users always pick the biggest, most expensive option and then they get really mad when their usage runs out because their IQ level is low. Their choice to always pick X high is dumb. Their understanding that that is why they burn through all their usage is low.
13:06 So, their likelihood of complaining is high. This happens a lot, like way more than I ever would have expected. People just default to the most expensive, most high-end thing, having no idea what that even means, and then end up burning themselves a bunch. Most of the people who feel that way would be totally fine with a model like Luna. Then you have these types.
13:28 I want to make sure I'm using the best model for the task. I don't want to have to learn which models are good at math and which ones are good at front end. Just pick the right one for me. Nothing is that simple. It's not that GLM is better at math and Grock is better at making funny jokes and Fable is better at front end and and I don't know, Soul is better at configuring machines.
13:49 I think it's much easier to categorize these into general buckets. And I hate that I'm doing this because it's going to force me to do something I really don't feel like. I have to defend Gemini a little bit. I would roughly say you could bucket these behaviors as intelligence and honestly behavior. I think these are the two categories that we can break things down with that make it much easier to to understand what makes a model like this special and what makes other models suck really hard.
14:23 In an ideal world, you would have models that were both incredibly intelligent and behave incredibly well. You also can have a model that's so intelligent that it behaves better. And you can have a model that behaves so well it seems intelligent. A great example of a super intelligent model GPT 4.5. This model had so much knowledge baked in. It's been rumored to be as big as 10 trillion params.
14:47 It was huge. It was slow. It had an absurd amount of knowledge baked into it. But it wasn't great to talk to. It wasn't a reasoning model. So it couldn't think through hard problems. it could just spit out answers to simple ones or knowledge information. And it wasn't able to do anything resembling agentic dev work because it was really just trying to be a super big super smart model, not a super capable model.
15:09 And ideally, capability comes from a combination of these things. But there's also a floor that you have to hit with behavior before you can be useful at all. An example of a model that behaves incredibly well without being super smart is right above GLM 5.3 Flash. I think that GBT 4.5 and GLM 5.3 are perfect opposites. Or to be a little more blunt, Gemini 3.1 Pro.
15:40 That model is unbelievably intelligent still to this day. 31 Pro has the best score on Skate Bench, my benchmark for measuring how well models can name a skate trick based on the description of how the trick is occurring. Like where's the board spinning? How's the purse spinning? What's the relationship between those? And then can you turn that into the unique grammar of skateboarding?
15:59 The highest score any model has gotten on my closed source extended version of skatebench is an 88%. Which was 5.4. I believe 55 and 56 from OpenAI have been regressions on Skatebench. However, Gemini 31 Pro is an exception because it gets between a 95 and a 97%. It's the only model in the '90s and it's nearing a perfect score because it has more knowledge.
16:25 So why is it so bad at coding? If you gave it code and said, "What should I change to fix it?" it would probably do that really well. If you say, "Okay, fix it." It's going to reason a bunch of nonsense and then say, "Okay, I'm going to rebuild your codebase because it lost track of what it's doing." You could see this as, to be blunt, intelligent models are the ones that have the most knowledge.
16:49 Behavior models are the ones that apply what they're told to the best. In an ideal world, you get both. And the combo of the two is when you get things like 56 soul or fable. I would argue both behaviorally well and are super intelligent. But fable is more leaning the intelligence direction and it has worse behavior as a result. And soul is rldled to Helen back to get the task done no matter what.
17:12 So it stays behaving for really long windows really aggressively. I have found this to be really hard to talk about because the models that matter have historically been relatively high in both of these except for models that are so unusable they're hard to talk about like Gemini 31 Pro. What has never happened for me at least before this model is something that is both genuinely stupid but also genuinely awesome to work with.
17:36 And GLM 53 Flash is a perfect model to highlight this. It's not incredible at front end. It's not super smart at some weird math thing. If we buck it that way, we're not going to be able to have a real conversation. When I tell it to do a thing, it does it. And halfway through, if I tell it to do something else as well, it remembers and does that, too.
17:57 The reason that this exists is because of how hard they RLED this model through the fancy stuff that they've been working on at ZAI to improve post training with provisioned environments that allow the models to do real work to learn how to work better. You could argue that what I mean by behavior here is capability in agentic work. How well does it stay on task?
18:17 How well can it have new things inserted? How well can it unblock itself? How likely is it to go off the rails versus do what it's supposed to? And GM 53 Flash, it just behaves well. So, let's hop over to the terminal I had running here where it took all of the PRs that I had open in T3 Code and then filtered out the ones that weren't made with Fable and then made a HTML page for me to go through here to see it all.
18:40 It even gave a snapshot time. The other things I used for this didn't do that. That's actually super helpful. I might make that a rule in my skill. Oh, here's one. I thought I merged this. Did I not merge? Oh, no. I said all the pull requests in the repo for this one. That is my bad. I thought it was only going to be mine. I then said ignore all the poll requests that were made by Codex.
18:57 should indicate in the body. I only want you to go through ones made by Fable. So, if I scroll here, yep, this one is Fable 5 running in Cloud Code. I thought it might have went rogue and broken and included PRs that weren't mine. I should have second guessed what I sent because OX aka 53 Flash does what you told it. That moment I just had there is one I often have with other models, even OpenAI models that tend to behave really well where I tell it review my PRs.
19:22 Then I make a change and then it suddenly reviews all PRs. I can't tell you how many tokens I've burned asking a model that should be really smart and is really expensive to go review my PRs and then it reviews a bunch of [ __ ] it shouldn't have. I didn't tell it to only review my stuff here. I told it to look at all the pull requests currently open in the repo.
19:43 That was stupid of me, but it did it. It succeeded. It audited hundreds of PRs, filtered out the ones that weren't by Fable, and gave me a set of things that it is confident I can merge. And out of all of that, it only has three that it's high confidence in, and a pile of others that it's medium confidence. Separately, it called it the high value unlocks like I asked.
20:02 Here is one that is useful because apparently busy threads freeze clients sometimes, and there's a PR that fixes it. I am curious. Let me take a look. Longunning threads make clients unresponsive. the thread detail and shell descriptions applied to every stream item individually. Quick question on this one. Are you using the legacy stream token by token setting?
20:17 I have had some pretty brutal long running threads for even days at a time and haven't noticed any performance regressions. I I know this is silly because obviously LMS are capable of like really cool powerful things, but this is a real PR somebody filed six hours ago that I'm probably going to merge tonight that I wouldn't have even noticed if I didn't run this test prompt that cost sense to run.
20:38 This will be an interesting test. I want you to figure out how many tokens I have used in GLM-5.3- Flash in Codeex on this computer today. I want input tokens, output tokens, cash tokens. The price is 7.5 cents per million in and 25 cents per million out. So, get me a cost number as well. I just want to know how much it cost to do that run because again, it audited every PR.
21:07 I'm using my ZI sub. Cool. So, rough math on what it cost me to audit my entire codebase, all thousand plus PRs I had open. 50 cents. That's insane. I've run this on Fable and just my own PRs was like over $100. I I don't think you guys understand just how crazy that is that I just got useful insights on pull requests in my repo that I might have missed otherwise.
21:39 I'm going to merge code tonight that it just found for me and the cost to do that accidentally much wider ranging than intended was still only 50 [ __ ] cents. Insane. So, the reason they could give away so much for free is because the model is hilariously cheap to run. I'm going to try running it on my DJX Sparks soon cuz the weights are out, too. And to the chat are asking a very important question because a lot of the time the answer to this is no.
22:07 Does it have vision? Yes, it does. I'll show you. I am screenshotting that. I am pasting it in the thread we have here. And I'll ask based on the work you just did, do you think these costs are reasonable? I think it didn't calculate cost based on cash. It might have been even cheaper. I'm sorry. It was 12.06. Here I was big spender out here spending 50 cents for useful insights.
22:33 Nope, it was 12. 12 cents to get a useful insight. I'm going to set up a bot that runs this every 3 hours and pings me when there's a pull request that thinks it's ready to merge and it's going to [ __ ] change how fast we get [ __ ] shipped. I don't know if you guys understand how crazy this is. Like 6 billion tokens in, the vast majority of which were cashed, 300k out, useful [ __ ] for 12.
22:58 And to be fair, I'm not even paying this. I'm paying the $80 a month for my ZI account that I forgot about. I probably could just cancel it and move over to open router because these prices are nothing. Like this rounds to zero. The reason they were giving out so many tokens is because they cost [ __ ] nothing. But there is a a cost to that. Like there's no super smart model that is also super good at doing tasks that's also super cheap.
23:20 The closest we have to that is 56 soul cuz it's so token efficient. This model is not as smart as those models. It will not solve hard problems as well as Fable or Solo will, but it will solve easy problems comically cheaper, and it will do real agentic work for you in ways that other things just won't. I've said for a while that I don't get why so many people are using stuff like Deep Seek V4 Flash because, how do I put this?
23:44 There's like different use cases for me. There are models that are useful for like one-off things like going to a web page and seeing if something changed on it or answering one-off questions or auditing my codebase, stuff like that. But if they miss certain capabilities, they stop being useful for that. Deepc V4 Flash has no vision. There's a new private closed source version of V4 Flash on their API that apparently has vision, but it's not open weight.
24:08 It's not part of the ecosystem. It's just a weird thing they're doing. Doesn't count. So V4 Flash is immediately not useful for me for a bunch of stuff. I also found that on longer running things, it would get distracted and lost and need more help and encouragement. Even the new snapshot 0731. All of the things I've seen people saying they're using DCV4 Flash for, I will actually use this new model for OX, aka GLM53 Flash, by having vision, by staying on task better, and potentially being slightly dumber, but way more
24:35 thorough and capable. This model is such a useful tool. It's a thing I see myself using a ton, but I want to see what others have done with it. both the code they've written with it as well as the benchmarks and how it's performing. I've heard really good things and also spoiler, I have a build of fish slop that I made with this model on this machine as well.
24:55 Oh, do we also have some front-end work that it did? We do. Cool. So, we get to go through all of the things that this model did if you did want to use it as a super cheap coding model. For what it's worth, I am still coding with Soul and Fable, but I've been surprised that this model's capable of, especially the 3D stuff I was really impressed with.
25:11 Let's dive in on the official post to start. Frontier Intelligence at Flash cost. If they had just put out this blog post without the early access window and without the anonymous ox drop, I would have called [ __ ] on this and probably not even looked. They made a really good call making this model free as an anonymous drop because it made me interested enough to try it and blown away when I did.
25:36 It obviously became the number one model on open router because again they're giving it away for free and a lot of open router users will abuse things that are free. I know I did and even in open code it became the top model too doing 43 trillion params during this testing window. It incorporates several architectural improvements over GLM5. They introduced a hybrid architecture combining sparse and linear attention which reduces long context serving costs while preserving precise long context capabilities.
26:00 That's another really cool thing about this model. It doesn't increase price when you hit a certain threshold of tokens in your context window. It just costs the same no matter what. They also have a new 30 trillion token multimodal pre-training corpus which is what allows this model to do not just image recognition but also audio and video which is super super cool.
26:16 It is crushing the Pareto frontier which is cost against your intelligence levels where it is so far over here. The next smartest models cost orders of magnitude more. It does cost slightly more than Luna because Luna is so efficient. We'll talk about efficiency in a little bit, but 53 flash is a meaningful jump in intelligence and I've experienced this myself while also still being comically cheap.
26:44 It is 9 cents per task versus the 5 cents per task that 56 Luna had. The cost measure here is logarithmic and it starts off from 0 to 5 cents being really small. In fact, these this indexing is strange, but this gap is nowhere near as big as it seems due to the log scale. So the costs are pretty close. It's a 4 cents difference per task. Intelligence is meaningfully higher.
27:06 Very cool. And again, open weight, so it will hopefully get cheaper and be hostable in places and able to do things that OpenAI would never want you or let you do with 56 Luna. It is only slightly dumber than Kimmy K3 according to the index, too. 60 versus 57 points. Truly unbelievable cost relative to the intelligence. To everybody who's complaining that AI is getting more expensive, I hope you understand that more intelligence is the thing getting more expensive.
27:30 At a given level of intelligence, the costs go down exponentially. It also is outperforming GLM52. I'm surprised they don't mention normal 53 as much in this. They don't even include it in the comparison here, but it's crushing their older models in tests like Automation Bench, which I guess for some reason Gemini does really well, and that's confusing to me.
27:46 Oh, it's 37 flash, but it's getting up there with models like 56 Terra, not Soul Terra, but also Opus 48, which was the last great Opus model, and it is neck andneck with those consistently across a wide set of benchmarks at under a tenth the price. I find this one particularly interesting. This is output tokens on ZI's codebench against 53 versus 53 flash versus Fable and Opus.
28:12 Obviously, Fable crushes. It is the best here by far. It's also not too token inefficient. But Opus 4.8, the gap between high and max for it is relatively big, but it's right up there with GLM 53 Flash on high and max. The big difference being 53 flash on high is close to max in this test, but it uses under half as many tokens. I don't even know what they're benching here, but to do the same work at 70k tokens that takes Opus 48 120K when the tokens are also so much cheaper is insane.
28:43 They have a section here about how the visual intelligence is super powerful, not just for recognizing the contents of an image, but for debugging its own work as it goes. And it was able to use Codeex's computer use surprisingly well when I tested it. Here's an example where it had layout issues on a page where it wrote the code. And then once it saw it, it was able to fix all the issues.
29:01 They had it make slides and they came out surprisingly good. Yeah, it's better at UI and things that look decent than I expected. There's a special word here that we're going to get to in a little bit, but I want to cover more of the model behavior before we get to the secret that made this model so interesting that I currently have highlighted. This chart is admittedly getting a little hard to read due to the number of models.
29:23 I deleted a bunch to make it easier. And they also only tested 53 flash on max. The score it got was a 63%. For reference, Gemini 31 Pro got a 12% on this bench and 53 Max got a 63%. It also ended up cheaper than Luna at its level of intelligence. Luna on X high was dumber but slightly slightly more expensive. Yeah, only high was cheaper but it scored significantly lower.
29:52 It seems like 53 flash in real world code work is an absurd value here as well. It scored higher than 56 soul medium at about a sixth the price. Hopefully I've made my point here. This model is absurdly good for the money. Oh, look. GLM 53 Max the original. a little bit higher score, but like 10x higher price roughly. Yeah. Interesting to see it tie with Muse Spark because there's a lot of things I like about both of these that are similar.
30:18 My issue with Muse Spark is that if you're not on the contributor tier, it is far too expensive for what it is. 53 Flash is very similar. I find it behaves a little bit better and is dumb a little less often. And it's just really nice to work with. Like that's the magic. This is the first small cheap model that I've used that is just a pleasure to use.
30:37 It does what you tell it to. I will say that its token efficiency is not quite where I would like when compared to other models especially recently. Like 53 flash on artificial analysis did around 47k tokens per task on average where Luna on max only did 20k and soul only did 17k. So it is doing meaningfully more tokens. But if you compare to something like Sonnet 5, it ends up being nearly double again.
31:04 So not the worst, but obviously not as efficient as I would like it to be. The only complaint I would have about this model is that it is not quite as token efficient. But now that they have this training run and the model this good, making more efficient versions shouldn't be too hard. Like Muse Spark was 30k tokens to complete similar work and add a similar intelligence level.
31:24 I bet this will get better over time. And for what it's worth, in my own uses, I have not seen it be too egregious with token usage. Somebody asked, "Why does token efficiency matter if it still wins on price?" It matters because it affects how long it takes for the work to be completed. And it also affects how much it bloats its own context during the work.
31:41 If you use more tokens to complete a task, your context window gets filled faster and your likelihood of getting to the solution before you max out your context goes down. Less tokens to complete work at a similar level of quality is always better. That's why OpenAI is working. so hard to win in that world there. But this model's not the worst at it, especially for a flash model.
32:01 Historically, smaller flash models have been really bad about this. Like even 37 Flash from Gemini was an improvement from their previous like Gemini 36 Flash on high. Wait, no, 36 Flash was less. Am I misremembering? I thought that 37 Flash was an improvement here. I guess not. Also worth noting the Gemini answers were the majority of the text, not the reasoning.
32:21 So, it's just [ __ ] yapping. Today, I learned. Huh. Anyways, as always, thank you to Dar for the work he has put in making this bench that shows the different designs the models make. Here's how it looks by default without any skills applied. Not great. Prague isn't on layout, but it's fine. Not my favorite, my least favorite. Oh god, this style. Everything has a template with this style in it.
32:48 Classic Tailwind template. And huh. Okay, hot take. This is the least cringe I've seen the terminal style look. Huh. Interesting. Let's turn on the design scale cuz I noticed it seemed to do well with that. Yeah. Okay. This it screwed up a little bit. I don't think it had vision during its design process to test things and check is my guess with how the bench is constructed.
33:14 But it made this look decent. It is capable of animation stuff I've noticed and I've been relatively impressed with what I've seen. I like where it's going for here where it has things moving in the background with text. It just doesn't look quite right, but it's doing something here. It almost cooked. It's getting close. With a little refinement, you could make this good.
33:38 Animations were broken, but I actually really like what it's going for on this page. Not bad. This is the best I think I've seen the design skill work for something. These are all great over to the taste scale. Yeah, if you give this the right markdown to force it to design better, it does fine. Oh, this one's actually quite good. Oh, it does the fancy little thing when you're scrolling where the cards cover each other.
34:12 This is good. It's also weirdly good at copy, I noticed, where it doesn't read quite like normal LLM text. Understory grows a private network of notes that get smarter the more you think. This doesn't read like usual AI generated slop. There's a lot of little things that make this model very nice. It's just it's pleasant. And it's rare that I say that about things that aren't from the Frontier Labs.
34:32 Usually they're smart, but you have to fight them constantly and different people are more or less willing. This model you don't have to fight. It just does [ __ ] and it does it pretty well. I did promise fish slop. I hit start and it did some things different. First off, it like actually feels solid to go around. It really likes animating though. There's a bug where the fish stop moving after the first time they eat, I noticed, and another bug where they will eat when they're not hungry.
34:57 So again, like it's not the smartest. It's making logical errors, but it's also doing little things really well. Like the animation for the fish moving, like it had the PGs to kind of do it, but it figured out which order to put them in to make the fish movement look good. It also added sound which I don't have routed right now but I hear it on my laptop and it's decent.
35:17 But again like logic bugs the food shows like the food meter shows right at the end instead of showing during. But it like got all these little things right. I'm really impressed with like the animations with the way that like text appears. There's all these little details it got surprisingly right for a [ __ ] flash model. I've had frontier area [ __ ] do a worse job than this 20 cent per million out model did.
35:46 But that's the easy version. I made it port the game to 3D. And yeah, it's not great, but it could do it and it made all the models and everything itself. This is not something that I would turn into a game ever. Like this is far from good enough to like be worth iterating further, but it did it as a 300 billion per model. That's impressive as [ __ ] I also told it to stop using computer use during it.
36:14 It was annoying me when it did. So, it didn't get a lot of feedback in this gen, but it was able to make it. It's more impressive than the Gro run in a lot of ways. While I might not have had the most success with fish slop, apparently this model is really good at Blender. The official ZI team gave GLM53 Flash a Blender scene and then in 12 hours it built this a full like real looking 3D set of kitchen environments in like restaurants.
36:43 This is a lot better than I would do in Blender, that's for sure. Like actually surprisingly good. God damn, I'm impressed. Not bad for a super cheap model you can run on like actually accessible hardware. And now for the final hook at the end here. The other reason they could run this model as heavily as they did and give all of those tokens away for free.
37:02 It's running entirely on Chinese AI chips. Good friend of the channel Barl did some research here. It's running on the Huawei Ascend 910 BC and it definitely has some like issues with memory. So they have optimized it heavily probably at the cost of throughput. This sounds about right to me, but at its price it's insane. And the reason they can serve it so cheap is because they have chips that aren't Nvidia.
37:28 Huawei is a Chinese company that's actually largely banned from the US. They are incredible hardware manufacturers and seeing them dive into chip manufacturing so successfully so quickly has been kind of insane. Running a model like this without using Nvidia in the pipeline is truly crazy. And yeah, I'm impressed. And of course, the most important question, where is this going to fall in my tier list?
37:49 I do want to tell you about one other thing that ranks really high in my tier list though, our sponsor. If you've never had to debug a complex DNS problem, I genuinely envy you because it's some of the most stressful and miserable days of my life. I still run into those to this day and every single time it happens, I'm just I'm not happy about it. But whenever I use DN Simple, that is not the case at all.
38:09 And there's a few big reasons why. First and foremost, they built a platform that's actually for developers, not just another domain reseller, which means they have SDKs for every language you would possibly want to write your projects in. I'm a sucker for Elixir, and they're a big Elixir house, so that always makes me really happy to see. But that's just one of the three reasons why DN Simple makes DNS exponentially simpler.
38:29 The next piece is quickly becoming one of my favorites, the CLI. There are other tools that offer some amount of domain control through the CLI, but almost none of them do it well at all, and all of them have crazy restrictions that result in you having to go into some dashboard and find the right place to make the right change, if you're lucky. And don't get me started on agents trying to debug DNS stuff, they will be confused non-stop.
38:49 Unless you give them a tool that gives them access to everything the service can do, like DNS simple has done here with their CLI. Everything you can do from their client libraries or Terraform provider can be done through the CLI, which makes it incredible for you and your agents to set things up correctly. But this is where the third thing comes in, and this is my favorite part.
39:06 Dian simple are humans. Okay, hear me out, cuz I just talked about AI for a second. These guys aren't some big VCbacked company. It's a small momand pop group that are genuinely just trying to make domain management easier. I've talked to the founders a bunch. They're a married couple that sent me some really cute shirts when I first started talking to them.
39:24 They sent a handwritten letter with it. It's just incredible. I've had problems of DNS on other services and I hit up my friends at D& Simple and they helped me out. Their support is the best in the world. And if you're having a DNS problem, the last thing you want to deal with is some crappy chatbot that responds every 12 hours with nonsense. What you want is real humans who understand the web better than you do.
39:47 And the only place I know where you can get that is sov.link/dnsimple. Get a $10 credit when you sign up. Okay, so time to answer the eternal question. Where does this go in my tier list? This is actually more tough than I expected. On one hand, this model has crossed the level of I actually have use cases for it, which is a challenge when Soul and Fable are as good as they are.
40:09 But on the other hand, there are so many things it just cannot do because it is not smart enough. I almost want to insert more tiers above B and below A. It also one more above A. I almost want to have a small gap between fable and soul then right after soul another tier that is 53 flash then another gap than the other models but in le of having to do that I think it's easiest to communicate what I mean here by just moving 53 flash up next to soul and then doing a much more painful thing which is move V4 flash down a
40:47 tier as well because this list isn't just how capable or smart is the model obviously ly 53 flash isn't smarter than K3 or Opus 5. This list is meant to be ranking what unique value does this model bring? Is it actually useful in ways others aren't? And I honestly believe with these three models with Fable, Soul, and 53 Flash, there is practically zero reason to use anything else in this list.
41:13 The things that made New Spark so strong were the vision capabilities, the speed, and the absurdly cheap price on the contributor tier. But that means you're giving all your code away. And if you use it on the less freely use my data tier, it becomes 10x more expensive. So if you're only using Muse Spark for the price, you now have an option that is the same price, but also slightly better behaved and importantly way [ __ ] more private because you can run it on your own [ __ ] before Flash was really compelling
41:42 because it was cheap and you could run it on your own stuff, but it had no vision. So, I have no reason to use that over 53 Flash like ever. Honestly, now I think about it, I almost want to move V4 Flash down more because there's no reason to use it anymore now that this is out. And no, V4 Flash does not have vision. Stop saying that because V4 Flash is a model, not an API.
41:59 There's a V4 flash vision expi endpoint that you can hit on DeepSeek's infrastructure, but that is not a model capability as far as we could know or not know yet. That is an API feature they implemented. You cannot use vision on any version of DCV4 flash that you can download and run. GLM53 flash again outstanding what it's capable of. And I am comfortable with my list looking like this.
42:26 Now, I think this is where I will leave the list. Let's wrap things up. This model proves out a whole bunch of things that I just would never have believed before it dropped. And if I didn't get to test it not knowing what it was as Ox Alpha, I probably wouldn't have believed the people saying what they were saying about it. But this model's [ __ ] awesome.
42:44 I'm using it a lot. I will continue to use it a lot. And I would honestly recommend you play with it, too. If you want something absurdly cheap, surprisingly capable, and incredibly thorough, and well behaved, you should give a model like this a shot. I have a feeling that you'll be surprised. And everybody who was speculating that Ox Alpha was a Gemini model, you really thought a Gemini model would ever behave that well?
43:04 This is an absolutely wild drop and I've had a great time using GLM 53 Flash. I bet you will, too. Let me know how you feel in the comments.