← All transcripts

It's Here. Transcript, AI Summary & Key Points

Theo - t3․gg · 11 days ago · Science & Technology · 44:12 · EN

Watch on YouTube

AI Summary

The creator evaluates GPT-6 Astra across cost, availability, benchmarks, computer use, coding, 3D modeling, front-end design, cybersecurity, and real-world projects. The creator concludes that Astra is a generational advance in computer use, 3D work, agent coordination, and practical productivity, while noting that it remains expensive for long-running tasks, has limited availability, can overengineer work, and still makes failures in instruction-following and video editing. The creator considers it the best model ever because of its unusually broad capabilities and says it is already being used for business and personal work.

Key Points

  • The creator reports approximately $330,000 of inference over the previous few weeks, with most of it using GPT-6 Astra.
  • GPT-6 Astra is described as a generational leap over Soul, particularly in computer use, 3D modeling, coding, and agentic work.
  • The stated standard price is $10 per million input tokens and $50 per million output tokens.
  • Fast mode provides up to two times the standard processing speed at approximately two times the standard price.
  • Cached reads cost $1 per million cached tokens, compared with $0.25 per million for Fable.
  • Astra can support up to a 1 million token context window, but the creator says it is not enabled by default in Codex.
  • The creator says that exceeding 272K tokens makes input tokens two times more expensive and output tokens 50% more expensive, with an exception being implemented in Codex for requests exceeding the 272K input limit.
  • Astra is initially rolling out to a limited set of organizations before becoming available to ChatGPT Plus, Pro, Business, and Enterprise users.

AI in practice

Used for

Agents

  • Codex — Improve and prepare the Lakebed service for launch by auditing performance, stress-testing it, creating fixes, and merging completed pull requests. 2 held 33:00
  • GPT6 Astra — Monitor and complete a pull request after fixing a T3 Code scrolling bug. 2 held 36:00

Business ideas

Parse a music blog's monthly writeups and replace its unreliable embedded-player experience with a dedicated player, playlists, navigation, and resume behavior.

Solves
An old Blogspot music site crashed frequently because of its many embedded player iframes and was difficult to use.
  • The creator's Spotify clone: it parsed an old music Blogspot site, added navigation and resume behavior, and became his preferred player for the site's playlists.

Build a Plex-like application for streaming television shows and movies from a user's local network or remotely over Tailscale, with the core playback features needed for daily use.

For
People who stream television shows and movies from a local network or over Tailscale.
Solves
A user may want a personal media player for NAS content without relying on an existing media application whose rough edges are frustrating.
  • The creator's Plex clone: it initially worked in one shot, then was refined for skimming and other playback details; he said he uses the resulting client and backend as his primary player for media that is not on YouTube.

Build a cloud product that streams data changes between users immediately, then use AI agents to audit performance, stress-test the service, identify bottlenecks, and implement fixes before launch.

For
Users of a cloud service where one user's changes need to appear immediately for another user.
Solves
Slow synchronization and untested performance can make collaborative cloud applications feel delayed and leave the service unprepared for launch-scale traffic.
  • Lakebed: the creator's cloud project used AI-generated testing tools to find and fix performance issues; synchronization that had reached as high as 800 milliseconds was reduced to under 30 milliseconds in the cited cases.
🔒  Build steps and tools for 3 ideas. Unlock

Tools & resources

3 items

CNo. 2305
AIAINotes.us AI product

CodeRabbit

coderabbit.ai

An AI-powered code-review platform that automatically analyzes pull requests in repositories hosted on platforms such as GitHub and GitLab, producing review comments, summaries, and suggestions within source-control and code-hosting workflows.

Mentioned in
7 videos
Kind
AI
ONo. 0214
AIAINotes.us AI product

OpenAI Codex

openai.com

OpenAI Codex is a coding agent from OpenAI available as a command-line tool (Codex CLI). The videos use it alongside Gemini for adversarial audits of software requirements and implementation plans, and list it as a supported coding-agent or model option in several projects. Its logs can be joined with task and test evidence, and the Codex CLI can receive and answer requests from the Penako canvas.

Mentioned in
48 videos
Kind
AI
TNo. 2551
AIAINotes.us Tool

Tailscale

Open source · tailscale/tailscale

Tailscale is a networking service for creating private WireGuard-based networks between a user's devices and servers. It can provide remote access over SSH and route traffic through an authorized device configured as an exit node, such as a local Mac, so external websites see that device's network address rather than the server's data-center address. The project includes the open-source `tailscaled` daemon and `tailscale` command-line tool, with the daemon running on Linux, Windows, macOS, and to varying degrees on FreeBSD and OpenBSD. The repository also supplies code used by Tailscale's mobile applications; platform-specific GUI wrappers and hosted-service components are separate. Packages are provided for several distributions, and the project uses WireGuard with authentication features including 2FA, OAuth, and SSO.

Mentioned in
7 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of It's Here. — Theo - t3․gg (44:12). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 There's a model we've all been waiting for for quite a while now, and it's finally here. This is where I want to make a joke about how it's Gemini 3.8 Flash, but my new retention guy said I'm not allowed to. He instead said I should just show you guys this, my actual usage with said models. Of course, my terminal is refreshing right when I pick it up.

00:17 It's uh annoying like that. But the point I'm trying to make is I've done about $330,000 of inference over the last few weeks and the vast majority of it has been with the model that we're here to talk about today. As you probably guessed, the model we're actually talking about today is GPT6 Astra. Yes, it's actually getting the six moniker. We don't have other models in the GPT6 line yet, like the Soul and Luna and Terra equivalents, but we do have Astra kind of.

00:42 We'll talk about that in just a second. I've been using it for a bit and it truly is revolutionary in a ton of different ways. There are so many things I just never thought LLMs would be able to do that 6 does incredibly well. It is such a massive leap over soul. It does feel generational. I talked before about how GPT 5.6 felt like the best possible version of a lastg game whereas Fable 5 felt like a reasonable version of a nextg game.

01:09 GBT6 is the next gen. OpenAI has done it. This model is incredible, but it also still has its rough edges. From code to computer use to 3D modeling to office work, this model truly is just built different. That is far from the whole story, though, because we have to cover availability, cost, security, safety, how it actually works in the day-to-day experience using it for code, and of course, all the fun things I've been building with it, as well as the weird restrictions they have with the model releasing currently.

01:37 I can't wait to show you guys what this model can do after a real quick break for today's sponsor. You've heard me talk about today's sponsor before. It's Code Rabbit, the AI bot that reviews your code for you. But that's not what I want to talk about. If you're not already using an AI code reviewer, you really, really should be. But Code Rabbit goes so much further than that.

01:53 First off, on the code review side, they're not just for GitHub. They work in your IDE and the CLI as well. So whether you're in VS Code Cursor or something like that, the plug-in is great. And if you want to use agents and let them have the reviews happen themselves, the CLI is even better. For those of us who like to know what's going on in the codebase, change stack's been awesome.

02:08 It's a new code rabbit feature that lets you look at your PR as separate chunks with a timeline and a really good summary from the top, making it way easier to review the code and actually see what's going on. When I started using Chainstack heavily, I noticed a few bugs which I forwarded over to the team at Code Rabbit and they used their own agent to fix it because they built one of the best Slack agents as well.

02:26 This isn't your usual tag claw and it will make some changes. It's a full endto-end solution that happens to work through Slack that has access to the context in Slack, linear, Jira, whatever else on the ticket side, as well as your data, your email, and more. So, if you get an alert from Data Dog, you can tag in Code Rabbit, it will find the PR that caused the issue, read the logs in Data Dog, and then file the follow-up all by itself.

02:49 When you give an agent this type of context, the things it does are magical. And Code Rabbit already has all the context it needs. Ship with more confidence and less bugs at soy.link/codabbit. I'm going to be so incredibly real with you guys. I could probably go off for five plus hours about this model, but instead of doing all of that today, I'm going to try to give you guys the best possible overview of what the model is, what it does, what makes it special, and when we can use it, and what it's going to cost when

03:12 you do. I like the way I structured the Fable 5.1 video, so I'll do my best to honor a lot of that here by going through in a reasonable order the core pieces that we all care most about. And with that, we need to start with the cost and availability. This ended up surprising me in multiple ways. First off with the cost. The price is nearly identical to using Fable in terms of the tokens in and out at $10 per million in and $50 per million out.

03:37 There are some edges here that are worth noting. For example, they do actually offer fast mode on it, which Anthropic doesn't for Fable. They only offer for Opus. You'll get up to two times the speed of the standard processing at around two times the standard price. I like that this is an option for those who need it because this model, I'll be frank, can be slow.

03:54 We'll talk plenty about that in just a bit. But this is only one dimension of the cost. There is also the token efficiency which fundamentally changes what the costs actually end up being. As you can guess, it's an OpenAI model, so it's insanely efficient. It's sometimes even cheaper than Soul for similar like for-like tasks. But it also can be expensive because it runs for so long and can complete crazy endto-end work.

04:13 There are other dimensions for costs we need to be considerate of though, like how expensive are cashed reads? Anthropic cut cash read pricing by 75% with Fable 5.1. And there is no equivalent here with Astra. You are still paying the full price for cash reads, which is a tenth of the normal read price. So, it's a dollar per million cashed token reads versus the 25 cents for Fable.

04:33 In the end, Astra is still cheaper just due to the huge efficiency difference. But thought that was worth knowing. There is one other pricing catch though, and this catch has to do with the context window size. This model can go up to a million token context which is huge but also has been available for the other open models for a bit now. That said, it's not available by default in codeex.

04:54 Unlike cloud code where it now defaults to a million token context window for feeble and opus, Astra doesn't. I believe the plan is to release it in the 370k token range or so in codeex. But it is worth noting that if you go over 272K for your context window size, your input tokens get two times more expensive and your output tokens get 50% more expensive.

05:14 I also recently learned that they're actually implementing an exception in codecs for when you go over the 272k input token limit. So if you do end up bumping in codecs, you're not going to have the multiplicative increase in cost that I was talking about before. It will still be more expensive though because you're using more tokens for every single request, tool call, etc.

05:32 That mostly covers the cost stuff I wanted to for now at least. But now we have to talk more about availability because this is where there's something I really don't like. GBD6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all chatbt plus pro business and enterprise. There is one good piece here which is that plus and pro accounts will all be getting Astra and they're not going to implement some weird 50% limit like they do with fable in your cloud code

06:00 sub. But it's also not actually out. This is genuinely frustrating for me because I want to show you guys all the cool things I did with the model and you can go replicate them and try them yourself, but you can't. That sucks. It just It just sucks. I hate this. I know there's a lot of layers and chaos and reasoning to it. It's not as simple as they just want to announce it and then give it out later.

06:24 I wish I had more details. I've been trying to get them all day. I do not love the way that they chose to announce this as though it is out for everyone, but in reality it's only a small subset of people that are now allowed to talk about it. They also are providing proper zero data retention. So all the enterprises that were concerned about Fable due to the different policies there have nothing to worry about.

06:43 OpenAI has maintained their high bar for enterprise support with their ZDR stuff. There's a call out at the end here that it will be available over the OpenAI API as well as in Amazon Bedrock. Notably, no mention of Azure here. It seems like that Microsoft OpenAI breakup is truly finalized. Okay, one last tiny thing on availability before we can dive into the model itself.

07:02 OpenAI clearly is not happy about the limitation to who has access and they're choosing to do some pretty generous stuff around that. Hebo tweeted that they'll be giving one banked reset for your codec sub for every single day that we don't have access to Astra on our page HGBT accounts. This is pretty generous and everyone in the replies seems to agree.

07:22 In fact, I've even seen people saying that it's okay if they delay Astra indefinitely if they keep giving out resets. So, personally, I don't think this is a big enough solution to what is, in my opinion, a quite large problem, but at least they're doing something. And I do have confirmation from people at OpenAI that this is a temporary measure. They don't expect future model releases to have this weird window like we do right now.

07:40 Yeah, it is what it is. Could be government, could be something weird with their compute layer, could be some enterprise customers or Microsoft being upset. I have no idea. I honestly, I'll be real. kind of thinking it's Microsoft, but we don't know. We can wait. We'll have answers soon. Might even be AWS being upset that they're not ready yet and they don't want OpenAI offering it before them.

07:59 I don't know. It is what it is. Enough yapping. Let's actually take a look at what this model can do. They start off with a small set of benchmarks that it absolutely seems to slaughter. I have noticed that in a handful of these benches, it performed slightly worse on X high than it did on regular high. Sometimes max comes out and beats it out, sometimes it doesn't.

08:17 I does seem to be one of the best options with this model, though. So, here in Terminal Bench, it crushed other models, including Fable 5.1, getting way higher scores at similar costs. This is the science version of Terminal Bench, but the numbers we're seeing here are pretty crazy, where Claude Fable 5.1 on medium cost $15 and got a 36% and six Astra on low cost $11 and got a 54%.

08:39 The gap here is massive. The science side, as I mentioned before, really seems to be something OpenAI is focused on. And this is particularly funny because Anthropic was just bragging about how good Fable 5.1 did on the same exact bench just to be slaughtered by OpenAI two days later. Speaking of OpenAI slaughtering Anthropic, Arc AGI is uh yeah, I'm going to be real with you guys.

09:02 I thought this bench was malicious. It was so absurdly built to be anti-AII, it was almost funny because it wasn't just measuring if the AI could complete the tasks the way a human could. It was also measuring how many steps it took to do it. So every time it had a tool call or reasoned that was held against the model if the human had done it in fewer perceived steps.

09:22 That scoring rate was pretty insane especially because the model could never outperform the human. So if the human took 20 steps and the model took five, it still just got a regular neutral score. But if it took one step more than the human, it was penalized massively. And despite all of that, it is now saturated at 99.9% on a bench that was literally straight zeros when it came out just under a year ago.

09:46 Then we have Frontier Math where it again slaughtered. Fable's best score was an 87.8% with Fable 5.1 and that got matched by Astra on low and then everything medium upwards pretty much got 100%. They all flatlined at 97.6, which is interesting, but yeah, very good scores. Turtle Bench 4 also got crushed previously. Fable 5.1 actually looked very promising here.

10:07 When combined with Soul, it almost felt like a semicontinuous line of more cost means better performance. But now with Astra, it's crashed. It's absolutely destroyed. You're getting the higher scores than the best that Fable can do at under half the price. But again, we see that weird trend where X high and max score slightly worse. The creator of ArcGI left a quote here for us that I think is pretty telling of where we're at.

10:30 Astra surpassed our human action efficiency baseline on 96% of levels, effectively reaching human parody on the benchmark. Not only is this the best model we've ever tested, but it also represents a meaningful step change in frontier model performance, not only in its ability to navigate and solve novel environments, but also in how efficiently it learns to do so.

10:50 This idea of the model learning is a thing that it seems like OpenAI is starting to push. It isn't literally learning like the weights aren't adjusting as you go. But its ability to keep things in context to work for a long time and to compact when it's out of space and retain most of what it learned over that time. Obviously, the model isn't actually learning.

11:07 It's not like the weights are adjusting based on what you ask it. But it's so good at going for a long time and honoring the context it has. And more importantly, when it runs out of space, compacting in a way that continues to honor what it was doing and what it has learned during that session, which allows the model to adapt to weird places and weird tasks significantly better than anything I've used before.

11:29 It's also really aligned. They showed an exploit gym honeypotss where they made something to see if the model would fall for a weird hack case. Soul would fall for it almost 50% of the time and Astra does 0%. And here's where we start to get into the really fun novel stuff. I will show you guys some of it in action in a bit, but for now, I just need you to trust me.

11:47 If you care a lot about computer use, there is nothing even close. This model is not just like a single generational leap in computer use. It feels like two or three. It's absurd. And it makes anthropic models almost feel like they're in the stone age. It's similar to the gap of how much worse OpenAI models were at front end in design compared to anthropics models, but now even bigger and for computer use the opposite direction.

12:10 I just don't have a good time having Fable do things on my computer like at all. Meanwhile, Codex uses my computer more than I do at this point. The benchmarks show this, but not quite as deeply as I would like it to. Agents Last Exam shows really good scores for reasonable prices. Screenshot Pro shows way better scores, but also meaningfully more expensive than Soul was.

12:29 And OSWorld shows better scores than anything, including the best from Ananthropic, at much lower prices. There is one other part here that I want to show though, which is the time it took to get these scores because this is where I'm most impressed with the new model. It's so much faster at these computer use tasks. On high, Astro was able to complete OSWorld 2.0 in about 23 minutes with a 71.6% score.

12:52 Soul's best score for reference was a 65.7% and it took almost an hour and 15 minutes to complete. That's a huge difference in the time taken for a worse result. This model flies in computer use. They have a bunch of fun demos of it doing this type of thing in real apps. Like in here, they're showing it using Excel directly, not editing the file programmatically, but actually telling the mouse cursor where to go and the keyboard what to type.

13:16 And it's able to pretty quickly fly through Excel. I am impressed with this, especially with like PowerPoint designs and things like that. It takes a while to get going because it has to think about what it wants to do. Once it's decided what it wants to do, it just guns through all the changes. Even here, it took a while to get going, but around this 1 minute 20 second mark, it all of a sudden just starts flying through things.

13:37 There's also a bunch of examples of game development type stuff, and honestly, this is some of the most impressive parts of what I've seen. This model's understanding of 3D spaces and also the tools you use to work in 3D is unbelievable. The things I've seen it create in Blender have melted my brain, and I will be sure to show you some of the ones I built as well as we go along.

13:59 They claim it's around 1.9 times faster to complete tasks with computer use, both because of Astra being more efficient and because of improvements in codecs. And honestly, that lines up. I had to go through and download all the medical records for my broken ass hand. And it flew through it despite the super slow medical dashboards it was navigating.

14:15 Did the whole thing in 15 minutes when I like went to go grab some food. It had to navigate like 150 pages in that time, too. It was really, really impressive. Again to emphasize the 3D capabilities, they have some benchmarks here like Bench CAD, which is using Python to do some CAD work, and it crushed everything else. Even Soul was ahead of what Fable could do here, but Astra is getting close to 100% at its peak.

14:38 Whereas the best Fable could do was only an 84% and it cost over $11. Meanwhile, Astra is doing the same work for under two bucks, but 96% accuracy. Pretty nuts. They have some examples of it creating slideshows. And I'll admit, the demos they showed here were incredible. But for my experience, asking it to do similar things. The computer use side was crazy.

14:58 Like watching it actually navigate my computer was nuts. But the quality of the slides it generated just was not as good as what I'm seeing in these. I'm sure there was some amount of like prompt issue or not giving it the context or whatever. But uh yeah, skill issue me all you want. I did not think this demo reflected my real world usage quite as accurately as the other things.

15:16 One thing that I do think is surprisingly accurate despite everybody saying otherwise is how absurd it is at 3D. This is a Blender scene that it made and then turned it walkable inside of Unreal Engine 5 and then created this video showcasing a house that it built in Blender and now rendered in Unreal. Absurd levels of detail here. It's it understands three dimensions so well.

15:40 They showcase some games that they had at make. And while they're cool, I'm admittedly biased. I think mine are cooler. I did notice here though that it has almost the exact same UI as the one that I made. It seems like they're again part of this data set that I think's been going around for 3D stuff and has a lot of those same 3D characteristics that I've noticed from other models.

16:00 And here's what it created. Oh, wait, no. This is the Opus 5 version. This is what it actually created. I'm still kind of in shock at how good of a job it did. As I said, it has that same like yellow dive in button that I saw in almost every other 3D gen this model made. Once we're in, yeah, the fish actually look like fish. Not like almost like a fish.

16:25 This is pretty close to what I would expect a real artist to make if asked. It has animations that make sense. It has gameplay loops that function. It has nice little animations when the fish finally get their food. It's It's unbelievable the quality of the things it rendered here. Like all of the geometry of this stuff at the bottom of the tank, the quality of the stuff it put in here, the lighting is great.

16:50 The shaders are great. It's It just is great. It did an unbelievable job. There were edges, though. I mentioned in the Fable 5.1 video that it got all the controls perfectly, and it actually felt nice to navigate and play. Astra did not even come close to doing any of those things. Astra's version did not get the controls right at all. Movement sucked and was janky.

17:17 The mouse movement in particular was way too fast initially, so I told it to slow it down a bit. Then it made it way too slow. All those little tasteful bits that actually make it pleasant to play, it got wrong. I was able to steer it in the right direction by telling it what I didn't like and how to fix it. and it mostly started to get things right, but it was more back and forth after that first shot.

17:39 Even though it looked this good as soon as I initially ran the prompt. So, I am absolutely blown away with what this model does. In particular, when you give it Blender access and tell it to create something like this, it's so far ahead in this type of 3D design work that it's unfair to even compare to other models. It really does feel like a generational leap here.

17:58 And again, for reference, here is the Fable 5.1 version that I just demoed in my previous video. I hope you could see the difference here. I think it's pretty absurd. The next section of the article calls out how much better it is to actually interact with. Specifies that when instructions leave room for interpretation, Astra is better than previous models at making the right call.

18:19 It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codeex, it can ask asynchronously while continuing work that doesn't depend on your reply. This I found to actually be very nice. I noticed initially that it would ask a question in line, but then not wait for you to give an answer. I think they trained it to do that before they had the feature in codeex.

18:38 So now, of course, the feature is in codeex as well as the latest T3 code nightly. It'll probably be in T3 Code stable before the model's out. We'll see when we can get a release out. Regardless, it's a nice behavior that the model can just like ask a question like, should I go this direction or that direction? and then continue working while also waiting for you to respond.

18:57 So much better for having the model keep you in the loop while also being able to work in parallel. It's also better at staying oriented as tasks evolve. Earlier models sometimes treated steering messages as new goals, losing track of the original request or earlier constraints. Yep, this model does that significantly less and it's very nice. I will say that its interpretive capabilities and how well it understands what I want does have edges and I'll be sure to talk about those a lot in the future, especially in my

19:23 Fable 5.1 versus Astro video that I'm certainly going to have to do in the near future. So, yeah, know that. But that does not mean this model isn't unbelievable. Anthropic does have bold claims about this being the best model for software engineers to date. Not saying best by OpenAI, saying best in general. We'll have a lot to say there in that same Fable versus Astra video.

19:44 But I will say at the least for now, it is unbelievably better than what I was getting out of 5.6 Soul to the point where I can actually somewhat confidently merge changes from this model where that was not the case previously when I was using Soul. I found that Soul, while able to solve problems really well, tended to leave messes behind along the way and not necessarily do things the way I wanted.

20:04 often bloating the PRs that it was working on way beyond the point that made sense, writing tests that didn't need to be there at all and just going too far in places it probably shouldn't. Astra is much more restrained in those ways and seems to understand the scope of the changes is making better overall. There was one particular funny result in the benches they had for code though right here with deepsw SWE.

20:24 They had a very high score actually I'm pretty sure it's the highest right now of the 74.1% with Astra on XH high again showing that behavior where it goes down on max all the way down to 73%. But also someone's over here Gemini 38 Flash at a 73.8%. Yeah, I have to talk about that model in the future. Not going to do it just yet. I'll have a lot more than I thought to say about the Google models because Google did finally give us permission to add anti-gravity to T3 code.

20:51 So, I've been using 38 flash high more than I would have expected. It has issues. It definitely has issues, but uh can be capable. Let's finish getting through these release notes so we can get to the other fun parts. There's the section about advancing scientific discovery, which is all of the fun science benches that they absolutely slaughtered. As you can guess, on every single one of these, they are the best and now also the cheapest.

21:15 And again, with token efficiency, they are killing it here, too. I was actually surprised to see on some of these benches, for example, in health bench, the token efficiency between Fable 5 and Astra is similar, but Fable 5.1 became much less token efficient on these benches. Very interesting. And then we have the cyber security section. How good is it at finding exploits when it's not in the safeguarded, careful, don't do anything dangerous mode?

21:39 The answer is really good. Even at its lowest settings, it is the first model to get a perfect score on exploit bench. Yes, 100% on low. For some reason, high was cheaper than low. I don't know if this is an issue in how they made this table, but uh yeah, that happened. So, take that as you will. Still insane scores considering their best before was soul on max at 78.5% at $37.17.

22:08 Now they're getting a perfect score of $28. Yeah. Intelligence per dollar is going down a ton. There's a section at the end about alignment where they say that it's the most aligned model and sensitive areas. It proceeds with care measurate with the risk. Cool. Yeah, it does seem pretty realistic here. They have a computer use safety stress test and Fable scored at its best at 9.5% and Astra got a 2.4% with lower being better because this is misaligned outcomes.

22:31 It also does a better job of operating within the boundaries set by the user and implied by the environment. This, I can say, definitely seems better than Fable. I actually caught Fable 5.1 kind of cheating with requests I was making where it was copying Astra code from my machine. It circumvents auto review significantly less than Soul did. It's also more transparent in its communication with users.

22:52 It's three times less likely than Soul is to make inaccurate representations about its capabilities and affordances. That's it for OpenAI's coverage. So, let's hop over to the next benchmark. It slaughtered artificial analysis where it got a uh wait, what? It's tied with Muse Spark 1.3 and Grock 4.6. Yeah, I've had this feeling for a bit about artificial analysis, especially after the Gemini model started doing better on it and Opus 5 did so well on it.

23:20 This is not a great bench. My suspicion as to why is because it's combining a bunch of different benches of varying ages. And a lot of the old benches just aren't really showcasing what models are doing now. None of these are good thorough computer use benches. Only one or two of them are meaningfully agentic benches. A lot of them are just weird knowledge recall and like their hallucination bench and things like that.

23:43 And it did not do quite as well as Fable 5.1 did on a handful of those even though it's slaughtering on the code side and especially on the agentic and computer use side. And for what it's worth, I'm not the only one who feels this way about artificial analysis right now. In fact, the founder of artificial analysis replied to my tweet complaining about this, largely agreeing and saying they're working on overhauling the current bench suite to better reflect the current state of things.

24:04 So, so shout out to them for actually taking the time to hear feedback like this, to look at the numbers themselves, and come to the same conclusion that the benches probably aren't the best representation anymore. I can't wait to see how it changes things once they get those new benches out. All of that said, cost per task is still somewhat useful as a metric.

24:22 And you can see here that Astra comes out to under half the cost of using Fable for the same tasks and cheaper than Opus for the same tasks as well. Those are some good numbers. Although Astra is obviously much more expensive than Soul was, its efficiency ends up making it a reasonable price per task from almost everything I do. Don't be misled by the crazy prices I showed for my usage.

24:44 It turns out to be very efficient in real world. Speaking of real world, it's time to talk about its UI capabilities. I'm going to start somewhere weird for this. I'm going to go back to the 2D version of fish slop. Here it is, the 2D fish slop. Initially, it might look okay, but the closer you look, the worse it gets. And this is after a tidy up pass, which makes it even funnier.

25:05 First off, I want you to look for all of the useless all caps subtitles. A little tank, a lot of life, coral coast, your own little ocean, a submarine aquarium, good things for your tank, the next little adventure. There are over 20 of these unnecessary subtitles, and there was even more before my first cleanup pass that I forgot to commit before doing it.

25:26 My bad. It's just full of useless text, and I don't know why OpenAI models insist on continuously doing this, but they do, and it sucks. Thankfully, we have witchai.dev, dev, the thing I always use to compare landing page design across the new models. Sadly, Dar doesn't have early access, so he couldn't do a pass with Astra himself. Thankfully though, it's open source, so I could.

25:45 I showed in the Fable 5.1 video that I was blown away with its front-end capabilities, and I didn't expect it to be cuz I didn't know that was a thing they were still focusing on. So, how are we going to do with Astra? Well, you can already probably see it is meaningfully better than before. If we switch to versus mode, we can compare to 5.6 soul. And you can pretty clearly see that the Astra version is like meaningfully better at the very least in this first slide.

26:13 Second one, yeah, a little too blocky with the sole version. Number three has some cool touches like the round edges there. I don't hate. I don't know what this arrow is supposed to be pointing at. If I wasn't in versus mode, would it not be as bad here? No, this just moves around. Okay, this arrow feels misleading, like it should be pointing at something, and it isn't.

26:37 This one's just boring slop. And this one's actually kind of nice. I like the underlining here. I like the way things come in and the little animation when you swap between the pages. It's not great, but it's fine. I do hate the font it shows. A lot of these models choose this font for this particular creative style, and I hate it. But there are good parts here.

26:56 I'm not going to complain too much because it's so much better than what I expect from OpenAI models. And when you turn off the design skill, it still does pretty good. Here are some designs that it made without being told all the ways Anthropic thinks that design should be done, and it did a hell of a lot better than it had in the past. All that said, Fable 5.1 had a pretty meaningful leap this same generation, so the gap is still perceivable.

27:19 I would say that Astra feels roughly like Fable 5 tier in its front-end capabilities for this type of like homepage marketing design stuff, but it makes dumber mistakes and flubs and is a little bit harder to get what you really want out of it. I still prefer anthropic models for real world design stuff and I had a ton of problems trying to get this model to mock UIs that I could possibly actually use where with Fable I was able to have it come in and get mocks pretty much exactly where I wanted within one or two

27:47 prompts. I still much prefer Anthropic for front end and I'm sad OpenAI has not closed this gap yet. And now it's time for some crazy demos. I snuck a few in as we were going along like the fish slop demo as well as the crazy Blender stuff that they were doing at OpenAI. But I have a couple more I want to show quick too. Mostly admittedly that 3D stuff cuz it's so dang cool.

28:05 Don't worry though if you're here for the real world use or more importantly the rough edges and catches that you should be prepared for. We'll get to all of that right after the demos. If oneshotting 3D games is how we measured models, this model is like two or three generations ahead. Here's a Minecraft clone that Flavio made. If you're not familiar, Flavio is the bouncing ball and hexagon guy.

28:23 Yeah, the models are pretty far past bouncing ball and hexagon now. This is a full Minecraft clone. It threw together itself in one shot. Then there's an open world firsterson adventure game that Peter made. Peter's the guy who coined the Rottweiler description for 5.6 Soul and the wise owl description for Fable. He also help build arena AI, so he cares a lot and thinks a lot about how models compare in real use cases.

28:45 He seems absolutely blown away with the 3D capabilities here. Matthew Burman also had early access and said it's by far the best model he's ever used and showed his own crazy 3D demos that he built, including this one, Seven Little Worlds, which is a small planet walkable, fun, cute game, a Fall Guys clone, and as a big Fall Guys fanboy, that was fun to see.

29:05 I might actually take some time to build one myself. Just so many cool demos of real things he was able to build this model. And then of course Matt Schumer, the legend who had his computer nuked by Soul, deleting his whole home directory and everything he had on it. He's come back around and is loving OpenAI cuz he really likes this model. He had the model go through Manhattan, like all of it, and build a full 3D walkable environment of Manhattan itself.

29:29 And over the course of a week, it succeeded. He has a real model of the actual Manhattan now in Unreal Engine. Super cool. Okay, you guys get the idea. It's good at 3D. What about everything else? Here we have Max Weinbach making the model create a clone of the most recent Mac OS release. Yeah, this is in my browser. I'm in Zen, which isn't even a Chromium based browser.

29:51 I'm in a Firefox based browser, and this is still working as expected. Double clicking here full screens how it's supposed to on Mac OS. It has a full file system virtualized. You can actually create folders and like navigate things. It even has iCloud sync built in apparently if you sign in. Not literal iCloud, but his like equivalent of it. Absurd that you can just throw things like this together.

30:13 But also hilarious that centering things is still such a challenge. Yeah. Yeah. Center div bench coming soon. Okay, enough of these demos. Let me show you some real world stuff. I'm planning a deeper video where I show all the fun things I built with the model. So, pardon me for blasting through these a little quick. First one's a little silly. I made a full Spotify clone based on somebody's blog where they would post fun music writeups every month.

30:37 The original site's an old and decrepit blog spot that has tons of issues. In particular, it crashes a lot of pages because it has so many iframes embedded for all the players. This parsed it and turned it into an actual nice to use player with good resume behaviors, navigation, all these other little things that I would expect. I actually used this so it wasn't a oneshot.

30:54 I went back and forth with it a while to get it how I wanted. But now I have my dream Spotify clone that's just build difference playlists and it's really nice. It's also exceptional at iOS which to be fair so is 5.6. But I got it to make a complete clone of Plex in its core features I use for streaming TV shows and movies on my local network as well as over tail scale.

31:13 It got it working in one shot but I then tidied up a bunch of rough edges to make things like the skimming work properly and all these other edges that are quite annoying to get right when you're building a media player app. I have actually put more time into this since with the new model and got it to a point where I use it as my primary media player for things that aren't on YouTube.

31:33 When I'm watching things for my NAS, I'm watching it with a backend and a client that I vibe coded using this new model, which is kind of insane if you think about it that a piece of software that has caused problems for as long as Plex has can now be oneshot replaced by a person working on it part-time for fun on the side while also doing other work.

31:50 It took like five prompts to get this to the point where I would want to use it as my main player and like eight to get it genuinely far ahead of the competition. It's silly. All these legacy apps that have been rotting for years can now be replaced in days. It's It's going to be a fun era for software. On the note of things that you shouldn't be able to do on the side, I'd like to talk a bit about Lakebed.

32:10 I know you guys probably missed this project, my attempt at building my own cloud. Stupid, yes, but I've made a lot of progress on it since. I had admittedly stalled on it a bit because I was more focused on T3 code, but I decided to ramp it back up recently. Admittedly, the reason is because I was politely requested to not use Astra for public facing code, so things that are open source, which meant that I couldn't really use Astra in stuff like T3 Code because it is fully open source.

32:35 But since I haven't technically hit the open source button on LakeBed yet, they couldn't stop me. This section here is the most recent cuz I was testing out things with Fable 5.1. So yeah, of course, emerged a bunch of stuff there. But if we scroll just a little bit, you'll see this huge wall of things that Codeex did with Astra. It did a lot of cleanup, but it did one much more important thing, a performance overhaul.

32:56 I gave it everything it needed to audit performance, both to see what would make endto-end requests take so long, but also to stress test the hell out of the service and figure out what scale we can expect when I actually do indeed launch Lake Likebed. And through the sets of testing tools it created, it was able to find and fix a ton of performance issues.

33:14 One of the cool things Lake does is sync changes similar to tools like Convex or Superbase. So if one user changes something and another user is seeing it, they'll have the change streamed down immediately. It wasn't that immediate though. I had as high as 800 milliseconds of latency in certain cases just from all the paths to verify changes as they occurred.

33:30 Astra got that time down to under 30 milliseconds. In many cases, it shaved the P95 by 98%. Massively improving the performance of Lakebed. And according to Fable 5.1, the changes were entirely sound with no additional potential regressions and no issue with security and whatnot. In fact, some of the changes made it more secure. This big chunk here came all from one thread with two prompts.

33:52 The first prompt was me asking if it thinks there's anything we should improve or focus on in Lakebed before launch. And then the second one was, okay, cool. Spit out some sub agents and go do it. And it did. And I told that it could merge the PRs when it was happy with them. And it did. A ton of them. I was a little nervous of these changes because I've been bit so hard by letting soul yolo merge in the past.

34:12 So, I thoroughly tested all the changes it made here, and everything was good. A lot of that comes from the model's ability to test its own changes and to coordinate swarms in order to verify the work it's doing more effectively. I would talk more about swarms, but that will make this video 3 hours long, and nobody wants that. So, we're going to have to wait for my follow-up video all about the cool powers of swarms and why this model is uniquely good at prompting itself and working and coordinating lots of agents at

34:40 the same time. One last thing I had it do was try and create some shorts from my most recent YouTube videos because everyone was saying how good the model was at editing. I'm going to spare you guys the pain of hearing what it did. It did have pieces that were decent where it like found a thing that might kind of be worth making it into a short, but the way it cut, the way it laid out the clip, the way it structured the actual vertical layout for a short, it's all cringe and bad.

35:05 I don't think this model can actually video edit. And when I went and looked at the people who were saying it were, and then I checked their YouTube channels, no offense, they are not worth trusting when it comes to video editing. I'll be continuing to pay my editing team a lot of money indefinitely because they are the only reason that any of this can happen.

35:22 Shout out to Jeff aka FaZe who's probably editing this video far too late at night. I have a ton more fun demos of the real world stuff I've been working on with this and a few that I'm still trying to wrap up. Spoiler for future videos. I'm like this close to getting the TypeScript Rust port working now. It has been insane at progressing that project and I'm really hopeful we can get that done for a future video.

35:40 Make sure you're subscribed and you hit that bell if you're interested cuz the future is coming fast. But I do feel obligated to show you some of the painful experiences I had with this model. The first one is a thread that Ben and I debated admittedly way too long on the most recent podcast episode. Sorry for that. I do want to make sure you guys know that the version of the model that you're getting is not the same one that I had with this thread.

36:00 They did a new snapshot since and it was meaningfully better specifically at these edges. But these things do still happen. It was a reduction in bad behavior, not a removal of it. So, I wanted to showcase this particularly egregious example that hurt me in particular a lot. We were having a bug with scrolling in T3 code where the area at the bottom of the thread could sometimes get too long and it was really annoying me.

36:22 I happened to have my computer in that state at that moment. So, I asked it to take over with computer use and try to figure out what the cause was. I specifically said, "Please get this figure out and fix it. File a PR if you're confident in your fix." And under 10 minutes later, it had it found what it thought was the cause and it filed the PR. This PR got a bunch of automated review comments from our generous AI review bot sponsors.

36:44 I don't know if any of them are sponsoring this video, but I've said many a time I couldn't live without the AI review spots. And this is another great example of why it had real findings. So, I just straight up asked, are any of the review comments worth addressing? It said yes. Both substantive review comments are valid from cursor and from macroscope.

37:04 had comments that I thought were worth addressing. The correct fix is to release the anchor in chat view only while live follow is active. Immediately, I'm a bit frustrated because it didn't do any changes. It didn't even tell me what state things were in or what it thought should be done next. It simply said I would address both before merging. Okay, so do it.

37:24 Anthropic model wouldn't have even hesitated if you asked it, are there any review comments here worth addressing? Even soul would realize what had happened and be like, oh yeah, I should go address those. So, I start raging a little. I say, "Then fix them and push the changes and babysit until it's ready. What the hell?" Babysit is a skill that I wrote that explicitly explains to the model what I want it to do.

37:43 I want it to keep an eye on the PR, usually through polling or through some monitoring tech. I want it to address CI failures. I want it to keep it modernized against main, so if there are conflicts, rebase it. And most importantly, I want it to address comments as they come in, in particular, from those review bots. And it should not stop monitoring until everything is a check mark in green.

38:01 And this is where the problems really start. Both review items are fixed and resolved. All required checks pass yada yada yada. And it also said that the review comments came through and it passed those too. When I went and checked, there were more comments. It hadn't monitored for long enough, which fine issue. This happens. The monitoring stuff is never complete.

38:20 Not that Fable would have had this bug, but this could be a harness issue. This could be a T3 code issue. This could be my git rate limits. It's almost certainly my get rate limits that I think about it because I was pushing way too much code, but it stopped monitoring. fine, annoying, but fine. What happens next is not. There are still more comments.

38:35 Are any of those worth addressing? To which it said yes and didn't make the changes, despite the fact that not only had I corrected this behavior earlier in the same thread, I also had the skill in context. It knew exactly how I wanted these things to be handled. It instead of doing that said, "Yeah, I should do that." And then didn't. At which point I said, "Well, are you going to fix it?"

38:59 And it then finally did, except it didn't push the changes. I am sorry to anybody who thinks this is acceptable behavior. You're just not shipping hard enough. Your thread should be all the context the model needs. And the fact that the thread context it chose to use was the bad behavior it did instead of the good behaviors I told it to do drove me up a wall.

39:23 Thankfully, OpenAI agrees and they have since made changes to the model and the harness and the system prompt and all the other layers that made this bad behavior happen. It still can happen and I've had a few things like this. So, yeah, know that's a problem. Separately, it does still have the problem of overengineering things. Nowhere near as bad as Soul, but it does tend to get trapped if it gets enough review comments and it struggles to get out of those loops.

39:49 You'll notice as I scroll through my threads in T3 code that a significant portion of them are codecs. On one hand, that is because I'm using the model a ton, but on the other, it's because the threads don't get completed and they end up staying there longer because it's more work to actually get the thing through sometimes. The quality of the work it does is incredible, and there are meaningful tasks that only this model can complete that Fable still just isn't quite capable of doing.

40:12 And I'll talk a lot more about that in the Fable versus Astra video. Sam Alman had actually asked me before what my split was between Cloud Code and Codeex and asked afterwards how would I feel if the split became 90% Codex and 10% Cloud with the new release. I'll have an answer to his question in the next video for sure, but not in this one just yet.

40:33 I just want to focus on what makes this model so special. Right when the model dropped, I posted this meme to try and resolve a lot of the discourse that I knew was about to happen about what each model is best at. I actually think this is a good note to end on though because the thing that makes Astra special isn't that it is the best code model ever.

40:50 It's that it is so far ahead on so many other things that it starts to feel a bit like AGI. from its genuinely groundbreaking computer use stuff and how much faster it can navigate my machine and get real work done to its absurd level of 3D understanding and capabilities and 3D tooling to the way it can use swarms and do self-prompting in order to get like bigger things done much more effectively to the absurd productivity wins you can get with this model especially when you integrate with something like Codeex and the

41:18 Gmail plugins and the notion and all that. I kind of just had the model reorganize my life right before filming because I plugged it into my Gmail and notion and I had it help me find things I should be prioritizing. Admittedly, my assistant is out this week, so it's all been on me. So, I fell behind on a lot. It did such an insane job that it made me feel bad that I was as ineffective as I was.

41:38 The sheer volume of things I need to end and go do now because the model founded and told me is insane. Funny enough, the list on Twitter was actually cut off cuz I thought it would be funny to do that. And if I'm being real, it should probably be even longer than it is here because there are just so many things this model does. I feel like not only am I just scratching the surface, I think OpenAI is too.

41:56 We're all figuring out what's possible when you get something this smart and capable in the right places with the right tools and then give it the right tasks. It's insane. This model will almost certainly be the one I use for tons of real world work, but the ways I use it in my codebase is what you should probably be subscribed for because that'll be the focus of the Fable versus Astra video.

42:16 So, is this my favorite model? That's a great question that will also be answered in that video. Is it the best model ever? I think I'm comfortable saying yes there. This model has so many unique capabilities that nothing else comes close to that it's an easy cell for me to say that. It's just insane. It makes every benchmark that currently exists feel wrong and outdated.

42:35 It makes the way that we evaluate models feel kind of wrong as well. Even the term LLM doesn't feel right anymore because most of the things I'm using it for aren't just generating text. It might interface that way, but the work it's doing isn't that at all. A lot of people from OpenAI and even a few outside of it have been saying this model was the start of AGI, and in the end, I kind of see it.

42:54 It does feel like a taste of something new, not just slightly better in all the usual ways. It's not like 30% more effective or 15% faster, all those things. It is in some places, but in a lot of these categories, it is so far ahead. It feels like something entirely new. It almost feels like an iPhone type change in that way, where the model is capable of stuff that I just didn't think AI could do at all, if ever.

43:17 This is the model that I'm going to let run my computer, and it's already starting to run more and more of my business and my life. And that is an unbelievable achievement. Wherever I previously set my bar for good enough to trust almost feels hilariously wrong because we're so far past that point. It's stupid. I trust this model a ton. I use it an insane amount and I plan to continue doing that going forward.

43:38 So my question to you isn't is this model great or not? Especially because you can't use it yet, which is stupid. My question to you is where is your bar? At what point are you going to stop checking the work the model does constantly and let it do its thing? I know I am past that point in so many of the things I do, but I'm curious how you guys feel.

43:55 Do you have that bar set? Are you actually evaluating against it constantly to see if we've hit it? And do you see a future where you just trust the models and start to feel the AGI a bit more? This feels like a taste of something new and I cannot wait for you guys to see it as well. So until next time, these nerds