← All transcripts

Anthropic Actually Fixed Opus Transcript, AI Summary & Key Points

Theo - t3․gg · 7 days ago · Science & Technology · 40:14 · EN

Watch on YouTube

AI Summary

Opus 5.5 is a major improvement over Opus 5 for everyday coding work. It is cheaper, communicates more clearly, follows instructions better, performs strongly on coding benchmarks, supports changing reasoning levels without busting cache, and delivers substantially more usable work within a Claude subscription. Medium, high, and XH high have been more effective in practice than low or max. Max can generate more than 10 times as many reasoning tokens for roughly a 1% score improvement and can get stuck in loops. Opus 5.5 remains weaker than Fable 5.1 for refined front-end design and less thorough than some competing models for deep code review, but it is close enough to Fable for general daily work that it offers much better value. It also produces impressive 3D games, simulations, animations, and browser demos, although it still has context-window paranoia, occasional spikes in quality, and smaller-model-like mistakes.

Key Points

  • Opus 5.5 costs 40% less to run than Opus 5 on typical workloads and uses cheaper token pricing: $4 per million input tokens and $20 per million output tokens.
  • Prompt-cache reads are 60% cheaper at $0.20 per million, and the model is approximately 30% faster.
  • Opus 5.5 communicates more naturally than Opus 5, puts important information first, produces clearer explanations, and makes long sessions easier to follow.
  • A tester completed a 680,000-line code migration in less than a day.
  • Opus 5.5 can fix software inefficiencies and perform performance optimizations on difficult code.
  • Opus 5.5 scored strongly across coding benchmarks, including Terminal Bench 4, Frontier Code, and Cursor Bench; its medium setting exceeded Fable 5.1's max result on several comparisons.
  • Terminal Bench 4 showed Opus 5.5 Medium at $2.94 compared with $20 for Fable 5.1 on Max.
  • On Artificial Analysis, Opus 5.5 on high scored 53 compared with 50.9 for GPT6 Astra on high.

AI in practice

Agents

  • Claude Code — Find a repository on the creator's computer, update it, implement a bug fix, create and submit a pull request, verify the changes, and merge the fix. 2 held 00:20
  • Claude Code — Run a performance-improvement job on a remote machine for the Fish Slop 3D demo. 2 held 00:34

Tools & resources

4 items

ANo. 0939
AIAINotes.us AI product

Artificial Analysis

artificialanalysis.ai/

An independent platform for evaluating and comparing generative AI models. It publishes third-party assessments and comparative data on model capabilities and performance.

Mentioned in
4 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

Mentioned in
74 videos
Kind
AI
TNo. 2097
AIAINotes.us Tool

Three.js

Open source · mrdoob/three.js

Three.js is an open-source JavaScript 3D library developed in the mrdoob/three.js project for assembling and rendering procedural scenes in browser applications, including Fable 51 Worlds, Human Atlas, browser FPS worlds, CAD workspaces, and physics games. It creates scenes containing cameras, geometry, materials, meshes, and animations, with WebGL and WebGPU renderers and SVG and CSS3D renderers available as addons. Its official site and repository provide documentation, examples, tools, an editor, community resources, and a JavaScript module distribution.

Mentioned in
7 videos
Kind
Other
WNo. 0520
AIAINotes.us Tool

WorkOS

Open source · workos

WorkOS is an identity and authentication platform for adding enterprise access features to applications. It supports Single Sign-On through SAML and OIDC identity providers, social authentication, passwordless email codes, multi-factor authentication, and user and organization management with configurable policies. The videos also describe it as supporting agent authentication on behalf of users. Its APIs normalize enterprise integrations behind a single interface, with dashboard management, webhook events for directory-service updates, and SDKs for Node.js, Ruby, Python, .NET, Go, PHP, and Java. WorkOS provides AuthKit, an authentication UI built with WorkOS and Radix, and the site states that the platform supports more than 20 enterprise services through one integration point.

Mentioned in
6 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Anthropic Actually Fixed Opus — Theo - t3․gg (40:14). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 Approximately 10 months ago, the way that we wrote code with AI changed fundamentally with a single model release. The model, of course, was Opus 4.5, which, believe it or not, came out in November last year, almost a year ago. Now, Opus 4.5 was a really interesting release for a handful of reasons. First off, it was way more expensive than Sonnet, which made it not really the default choice for developers.

00:19 Second off, they skipped a bunch of numbers and went straight to.5. But third off, despite being a revolutionary change in how we can use models for code, they didn't make it Opus 5, they made it 4.5. Since then, Anthropic has released some incredible stuff in the Fable line, but pretty much every other launch they've done has been iffy at best, especially model releases like Opus 5 and Sonnet 5.

00:40 And I'll take some fault here. I fell for Opus 5 when it came out. It seemed like it was working really well. But when you combine the absolute slop it would output in text trying to describe what it was doing with the code that would often miss really important stuff, the model felt more like a liability than an actual thing you could rely on when you worked every day.

00:59 All of this sets a pretty brutal stage for Anthropic who previously made incredible models in the Opus line and has since kind of let it die. All of that said, I still did hold some hope for Opus because I needed something to use the other half of my anthropic subs with. And I still have a soft spot in my heart for that moment in November last year where all of a sudden I was watching my computer code for me instead of just telling it what to change.

01:22 That's why I was so shocked when I saw that this isn't Opus 5.1, it was 5.5. They set a really, really high bar for this release. And while it is admittedly far too early to know for sure, think they may have hit it. I've been using Opus 5.5 all day and it has been blowing me away. They made it cheaper. They made it less bad at writing. They made it way better at code and following instructions and just doing stuff every day.

01:45 Is this the Fable Killer I've been waiting for? I can't know that just yet. But is this model blowing me away? Absolutely. I can't wait to show you all of the reasons why I'm so impressed after a real quick break for today's sponsor. While I was prepping to film today, I was trying to get a bunch of dashboards open. It's been a little hard for me cuz I'm kind of down a hand.

02:03 So, I decided to let Codeex use computer use to go through all of these pages and get me signed in. And I was blown away at just how many times it couldn't do it. The biggest reason is that these sites don't offer good O solutions for agents, which makes it really hard for my agent to do anything on my behalf without me signing into the browser myself.

02:20 If only all those companies were using today's sponsor, Work OS, because they built the standard for this, OMD. The goal of OMD is to give an easy way for agents to O on a user's behalf or set things up for a user ahead of time. Work OS partnered with Cloudflare and Firecrawl to build this standard and now a ton of other companies are adopting it and helping push it further.

02:39 Companies like Neon, Recent, Parallel Monday, and more are already moving over. And if you're looking for the easiest way to adopt this new standard, work OS is probably your best bet. There's a reason companies like OpenAI and Anthropic trust them so deeply, it's because Work OS solved enterprise off, not just simple signin with Google buttons. All the pieces that real businesses need when they want to adopt your product.

02:58 If you've never had to configure Octa, ADP, or Duo at your job, I envy you. It is hell and miserable, and I wouldn't wish it on my worst enemy. Thankfully, when you use Work OS, you get the admin portal. Just you send a single link to the IT team at the other company, and now they're good to go. Works is the only platform that is loved by developers, agents, companies, and IT teams.

03:18 Figure out why atv.link/workos. Let's dive into Opus 5.5. Kind of crazy. This is one of four major model drops in the last 2 days with Gro 4.7, then Opus 5.5, and then GPT6 Soul and Luna, which of course we'll be covering in the near future. If you haven't hit the subscribe button, you should consider it because I stay up all night working on these videos nowadays cuz there's just so much to cover.

03:39 I do my best to distill it. I'm sorry the videos are long, but hit the button if you don't mind. Helps us out a ton. Less than half of y'all are subbed and it's a great opportunity to stay on top of these things as they change constantly. So, what makes Opus 5.5 so cool? We'll start with what they had to say. It performs at the level of Fable 5.1 on most work and it costs 40% less to run than Opus 5.

03:58 They open by saying that this is the first release since they called for pacing the frontier. I have a lot of thoughts on that that I'm going to reserve for a future video, so we'll ignore that for now. What I care a lot more about is how this model actually works. They mentioned that they had a tester that completed a 680,000 line of code migration in less than a day, which is genuinely very impressive.

04:17 They also site its ability to fix inefficiencies in software. I've experienced this too, throwing it at some pretty brutal stuff to try and do performance optimizations. I'm impressed. They say it's really safe as well. At this point, I just trust Anthropic when they say this, but now we need to go to the cost and speed part. This is where things start to get fun.

04:35 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings, it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are now cheaper at $4 per mill in and $20 per mill out, which is 20% less than before. They also made a huge change to cash reads. Not quite as big as the change that they made for Fable, but closeish.

04:59 Now it's 60% cheaper to do cash reads than it was before. Only 20 cents per mill, which is a very good deal. It's also about 30% faster, which I have seen from my numbers. It is much faster, which is nice because Opus can be a bit token hungry. This version in particular, uh, we'll talk about the token efficiency in a bit, don't worry. My favorite change by far, though, is communication.

05:22 5.5 communicates more naturally than prior models. Early testers found its writing to be clearer and easier to follow, which addresses some of the common feedback that they heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, quote, "It writes the way I do."

05:41 They also call out that being easier to read and follow along is actually a safety benefit as well as a practical one. They have a real fun teaser at the end here saying that Sonet 55 and Haiku 55 are going to follow in the coming weeks. That's going to be our first Haiku release in a year, by the way. Insane that it's taken them that long. They are getting absolutely destroyed by OpenAI on the cheaper, smaller models, but it does seem like they have a solid lead on the big ones right now.

06:03 First, we have the benchmarks where literally all of them other than the agentic scientific research bench, terminal bench science, and business workflows with automation bench. There are a few other benches that weren't included here where they didn't score quite as well, but I'm still waiting on updated numbers from those. In particular, Deep Suite.

06:21 Guys, can you update the bench? I've been bugging you on Slack forever. Your data is just so out of date now. H. Anyways, here's where we start to actually look at the results where we can see the cost against the score for various things including Terminal Bench 4 where it is slaughtering where Opus 5.5 Medium is scoring higher than Fable did on max and it's under half the price.

06:42 Actually, it's a lot less than that. It was $20 for Fable 5.1 on Max and it was $2.94 for Opus 5.5 on medium. Yeah, it did have a slight dip on max and we'll be talking about that in a bit. I have feelings about the max effort level. If we look at other things like Frontier Code, you can see the the weirdness of this particular bench where things go down as often as they go up when reasoning levels increase, but it is a new all-time high score on medium and on max, which is funny.

07:06 Their best score was medium. I don't know why they put this in here. It just weird bench still. But then we go to Cursor Bench, which is one of the most fun ones. Sadly, Cursor Bench no longer includes Astra because OpenAI has banned Cursor from using their models. But we can still see how it compares to Fable, which again shows Medium as scoring higher than Max did with 5.1.

07:26 Kind of crazy that Opus 55 on Medium in most real world codebenches is showing a much higher score than Fable. I have a lot of feelings about how these charts are visualized to the point where I actually took the time to make my own alternative visualization. I'm just using the data from artificial analysis here, but I wanted to make it easier to showcase the actual cost gaps between these different options.

07:49 Right now I have it on linear cost scale which really emphasizes how cheap Luna is because it's just stuffed in the side there when compared to all these other things. But I'll switch to log so it's a tiny bit more readable in particular the section we care about here. What you'll see is that this model is more expensive than Astra in a lot of cases but it is also proving to be much more effective with its results.

08:10 Like here, Opus 5.5 on high cost more than GBD6 Astra on high, but it also scored a meaningfully better score getting a 53 on artificial analysis versus the 50.9 that Astra got. So again, it is still not cheap, largely due to the token inefficiency, but god damn, these scores are insane, but also low is really bad. This model's interesting in this particular way.

08:37 It feels like low and max are dangerous and generally best avoided, but medium, high, XH high, those all have seemed really good for my experience so far. They also report improvements to knowledge work, which is cool and awesome if you really like using Excel with your models. I'm sure this is super important for a lot of you guys. Not a thing I particularly care for.

08:58 Yeah, good numbers. Cool. Nice to see. Here is the thing I said I was most excited about, though. Communication. We've made major improvements to the way opus 5.5 writes and communicates. One of the most common areas of feedback we heard about opus 5. Yeah, please explain the issue to me. This is the opus 5 example. What I found the extra drop isn't the free tier mdash.

09:18 It's a regression in random commit hash. The bug bunch of code used to do a halfopen interval. That is absolute nonsense. And there's three m dashes in that response by the way versus opus 5.5. The extra drop is a bug in the billing refactor. The free tier change accounts for only $1.50 of Acme's August drop. The other $9.92 come from a bug in this particular commit.

09:44 That commit was labeled quote no behavior change, but it stops counting usage from the last day of the month. And then a clear what changed the before and after. It's just so much more readable. They have a handful of examples of this. And yeah, the difference is pretty immediate. I've already felt this since I started using it. I actually took the time to go delete my unslop stuff from my agents in CloudMD in order to see how this model just organically talks and I don't think I'm going to bother readding it cuz it's

10:12 doing a great job. One last thing from the official article before we dive into all the much more interesting stuff people are doing with it. The data retention policy. For those who don't know, the Fable line is unique in that it doesn't offer proper zero data retention. So companies that need to have vendors that don't have any of their data saved cannot use Fable.

10:28 Opus has stayed the number one model for businesses and also the number one coding model in general. Both because it represents a good price to performance for a lot of businesses, but more so because the ZDR policy can be applied on Opus and it couldn't be applied on Fable and Mythos. This means that those companies, the ones that spend all this money on AI, kind of couldn't enable Fable for their employees.

10:49 Opus they can absolutely enable. So this model is almost immediately going to become the most popular for coding with AI by far just from this one particular detail. And it seems like y'all are going to have a pretty good experience. Opus 5.5 is way, way, way better than Opus 5. Sorry about that model. Please try this one. This is a fantastic tweet for a handful of reasons.

11:12 I don't want to like overread social media etiquette stuff, but trust me, it's worth it. This is an acknowledgement that Opus 5 missed the mark. This is like a silly tongue-in-cheek way to communicate with the audience and users that I love. It's very personal. like clearly didn't get through like a normal PR review. But most importantly by far, this seems like Anthropic is loosening their very tight leash on what employees are allowed to do on the internet.

11:37 Historically, it has seemed like anthropic employees are scared to comment on almost anything. I'm not feeling that at all with this release. It seems like everyone's just kind of allowed to go and post now, which is a very nice cultural change at Anthropic, and I really hope they maintain this. I know y'all think I'm an Anthropic fanboy now. Well, at least Sam Alman does.

11:57 But for the rest of y'all who aren't familiar, I have had massive issues with how Anthropic runs their business for a while now. So, it's very relieved to see these changes to how they've been operating publicly. It's a nice shift. And this whole release kind of just feels that way. Like a nice shift in the right direction. I do want to talk about some more benchmarks in a second, but I want to give you a quick teaser of some of the fun 3D stuff that we'll be showing in the end as well, because I know y'all love these

12:21 crazy flashy 3D demos. This is a clone of Dark Souls that was made by Matt's editor and assistant, Alex, here. That is pretty nuts. Like, the fact that models can do all of this is insane. I honestly would have called this demo fake if I hadn't gotten similar results for my own stuff, which trust me, we'll cover in a bit. Okay, I lied. One more detail before we get into benches.

12:48 Opus 55 fixed one of the most annoying things about anthropic models, which is that the current effort level was at the top of the thread as part of the system prompt. This sounds like a silly thing to care about, but it means that you would bust cash whenever you changed your reasoning levels, which isn't great. Now you can change reasoning levels midsession without busting cash.

13:07 Very nice change. Thank you to Lydia for calling this one out. I might have missed it otherwise. Very happy to see. With all of that said, it's time for some benchmarks. As you saw earlier, Opus 5.5 now has a massive lead on the artificial analysis intelligence index, a 5oint jump from Fable 5.1. For reference, the gap between Astra and Gro 47 is not much bigger.

13:29 And the gap between Astra and Muse Spark is roughly the same. So if you think Fable is way better than Uspark, which it is, allegedly the gap from Fable to Opus is similarly sized in favor of Opus, which I think more so showcases how bad these benchmarks are. But it is worth noting that Opus 5.5 is slaughtering basically every bench you throw it at.

13:48 It comes out at the top in almost everything. In particular, in output tokens, this is one of the sadder things that I think is important to call out. The price changes are real. This model is cheaper, but it's not because it's more efficient. In fact, this model is meaningfully less efficient than Opus 5. Opus 5 on Max did 73k tokens per task. Fable did 78k tokens and Opus 55 max did almost 120,000 tokens per task.

14:15 For reference, GPT6 Astra did 27k tokens per task. Yes, there is really a 4x gap in efficiency here. And I thought that was worth calling out because it doesn't matter if you're two times faster if you need four times more tokens to generate. And it also closes some of that cost gap which is why we see the interesting numbers I had here in my comparison.

14:37 Opus 5.5 comes close to being the most expensive model that artificial analysis has run simply because of how many tokens it is using. Opus 5 cost $5.86. Opus 5.5 cost almost six bucks and then Fable cost $763. But this is far from the whole picture because we're only looking at the max reasoning level for all of these. I made a subtle change here. I added the low, medium, and high for everything.

15:02 The goal here is to show you that the costs aren't that brutal if you're not using max, which I will about even more in just a moment. You'll see that Opus 5.5 on XH high is a lot closer to Astra's costs. But if you're willing to bump down to high, it ends up being about half the price of Astra on Max. And when you go down to medium, it is way cheaper at $134 versus the 326 for Astra or the $6 for Opus 5.5.

15:24 So yeah, the cost scales much better. I didn't put low in here because I think low is pretty stupid. It's just not that capable considering the price. What's much more interesting here is the Pareto Frontier line. This is the line for what is the best price at a given level of performance. Opus 5.5 has swept this line. It is the best price for all of these scores in the 50 plus range.

15:48 I wish that artificial analysis made these visualizers a bit better. I've had to, as I mentioned before, rip it and make my own, which, as you can guess, I made with Opus 5.5. It was actually very pleasant to do this type of thing with. It was able to rip the data from Artificial Analysis's homepage and make a better visualizer in not very much time at all.

16:05 But there is one little thing I want to sneak in here as well. I have other models. I'll throw Gro 47 and 46 in here so you can see why they aren't really relevant anymore because they are worse and more expensive than any other competition in the space right now. They are scoring below at higher costs consistently. What I wanted to sneak in here is Mimo V2.6 because this model is actually performing very well for the price and is a nice end to smooth out Opus 5's line here before you hit the super cheap Luna models.

16:32 I switch over to log. You can see even more so how unique Mimo26 is in this chart. It seems like Xiai's been cooking on their new open weight line. So I'm very excited to dive more into the Mimo stuff in the near future. But for now, we're talking about Opus. So yeah, just wanted to call that one out quick cuz I thought it was cool. Artificial Analysis published a little breakdown for why the model is not much more expensive despite the fact that it's doing two times the number of tokens.

16:57 They point out that it would have been 80% more expensive if they didn't decrease the price and also cut the cash read costs as well. Those two changes were able to keep it from being a massive price hike. I like this call out from Edwin which is similar to what I said before about the reasoning effort levels where medium is way more efficient. He says that medium on Opus 55 is Fable 5.1 level intelligence.

17:17 I think that's a bit of a reach, but it's only $1.34 per task on artificial analysis versus the 6 plus for Fable. I think this release is really emphasizing my disdain for those max effort levels. They result in things just running in loops and never finishing and burning way more tokens than are needed for the vast vast majority of tasks. Believe it or not, even I got hit with this earlier.

17:41 I had a thread go for 6 hours and 30 minutes writing a markdown plan because I had it on max and it just ran in loops forever. I ended up having to stop it and then switch to XH high and ask for an update. And apparently it only got halfway done. So it will hopefully be able to finish in the near future. Sadly probably won't have that demo ready for this video, but I have plenty of others.

17:59 Don't worry. Just as one last piece of evidence for my disdain towards the max reasoning effort, I ran this model through Skatebench alongside GPT6 Soul and Luna. And the thing I care about with this isn't how high the score was. It was fine. It was an improvement, but not a massive one. Actually, no, that's a lie. It actually scored worse than Opus 5 did on Skatebench.

18:19 Interesting. Well, what I found much more intriguing with this is the gap between Xigh and Max's amount of tokens generated per question. On XH high, it only used about 330 reasoning tokens per question. And from X he high to max, which is just a one little bump inside of your slider, it went from 300ish tokens per task to almost 5,000 per task. I'm not exaggerating.

18:44 That's over a 10x gap for that tiny little jump. And you get almost no improvement in score here at all. It's a 1% bump. And I still have two of these running because it just takes forever cuz it's burning so many tokens. It is a more than 10x increase in your costs for no good reason. So yeah, don't use Max on this model. It's not worth it unless you just want to watch tokens burn and be lit on fire.

19:09 And I know rich coming from me, the guy who loves talking about their token furnace, but this model has been very pleasant in that regard. As I mentioned, I've been going ham on it all day. And remember, I have five Claude subs because I really like Fable 5.1. Let's refresh to see how my usage is after heavy usage all day. Yeah, not much. One of my accounts has been hit for the majority of it.

19:32 This account got a little bit more because I was using it in the Cloud Code desktop app. But the vast majority has been in this here, which was only about 40% of my weekly, except for the fact that about half that weekly usage came from me using Fable over the last 2 days. So, it's actually closer to like 20ish% of my weekly, which is crazy considering how I've been going out of my way to push this model all day.

19:56 I was able to get some real work done with it. I made that new visualizer that I showed earlier. I ported it multiple times to different places that I wanted to host it with additional functionality. I had it use computer use to verify it and also get the data out of the original artificial analysis site. I noticed a bug in Lakebed my cloud when I was working with this.

20:13 So I told it to go through my computer, find the repo, and make a PR. And it was able to find it, realize how out of date my local clone was, update it, make a work tree, file the PR with the fixes, verify the fixes, and get it all merged for me. All of this combined was 1% of my weekly. Yeah, they fixed the problem. I can't believe I'm saying this, but if this model does turn out to be a solid daily driver, which it seems like it will, it is a massive improvement to the value you get compared to something like the

20:43 codec sub. It is a little bit too easy to burn through your codec plans right now. And since the only options you have are Astra, which is capable but spiky and way too expensive, or Soul, which we'll be talking about in a new video soon, this is a way better value for your 200 bucks right now. Kind of crazy that this happened so quickly, but yeah, it did.

21:05 I've only used my codec subs lately today to play with the new GBD6 Soul release. Not really reaching for it that much right now. I am just really impressed with Opus 55 for everything I've been doing with it today. I've been pushing this model hard to find its weaknesses. And thus far, what I found is that its weaknesses are similar to those of Fable.

21:22 I mentioned yesterday in my Grock video that I made a bench that was trying to figure out how well models can dive into big code bases and make realworld suggestions for improving the quality of the code. And that bench actually showed Grock 47 massively outperforming relative to what it is and its capabilities. because I have actually found Grock 47 to be pretty good at being thorough with code, reviewing the hell out of it and diving deep.

21:46 My suspicion for why Grock is so good at this is all the effort that cursors put into code review pipelines with things like Bugbot and all the data they have from that. That data is super useful when you're trying to fine-tune models to be better at digging into the details and reviewing code. I suspect that is why they're able to make Rock so uniquely good at this compared to other models.

22:04 Soul also seems to be very good at this as well. Opus and Fable less. So, I will say the gap from Opus 5 to 5.5 in this bench is absolutely hilarious, nearly doubling in the accuracy and reliability of its findings. And Gemini 38 Flash is still meme tier as expected. So, if you're planning on using this model to very thoroughly review code and find every single thing wrong with it or suggest sweeping improvements to your codebase, might not be as strong at those things as OpenAI's Frontier is.

22:32 But, I'm going to be real with you guys. This is not a task I do every day. This is a task I do once every two or so weeks, mostly as a way to test new models. It is super useful to have a model that can deeply dig into every detail in your codebase, poke every hole there is to poke, and give you real feedback on how to improve it, but it is not the default case and it is not the most realistic thing to bench.

22:55 I just thought these numbers were interesting and it was one of the few things I could do and find that showed the weakness that I perceive with anthropic models compared to OpenAI. Still got a bunch of fun things to cover including the model's front-end capabilities and most importantly all the crazy demos people have been making with it. This model is insane at 3D and the stuff I've been seeing is mindblowing.

23:15 Normally I would make a joke here about how expensive it was to generate all of these tests and all of these examples, but this model has not been very expensive for me. Regardless, I hope you can pardon me for a real quick sponsor break. I have two quick questions. First, have you ever used an agent without search? If you have, you know how painful and miserable it is.

23:31 It basically can't do anything. My second question is the opposite. Have you ever used an agent with insanely fast and accurate search? I personally hadn't until I started using today's sponsor, Parallel, because they have the best and fastest search results for pretty much every single thing you can measure. Historically, I've really liked the search built into OpenAI, but when you compare it to Parallel, it's just night and day.

23:51 Watch and see just how fast Parallel can get you good results. Going now, all real time, of course. It took just over a second for parallel. And OpenAI still going. Still going. Meanwhile, OpenAI's endpoint took over 6 seconds long. If search was all they did, they'd be one of the best options available. But they also have everything else you would need on the web.

24:09 from a proper monitor that will send your agents info when things change on different pages to a traditional response API for when you want to actually get synthesized results from your search queries to their extract endpoints that let you send a URL and get back the data that your agents actually want and need. Even parsing JSheavy pages, by the way, there's even an MCP so you can expose Parallel to your existing agents in whatever tool you're using now.

24:31 Every month you'll get 5,000 requests for free. And on top of that, if you sign up today, you'll get $80 in credit. What are you waiting for? Join now at soyv.link/parallel. Next, we need to talk a bit about front end. This one's interesting cuz the model does seem to have really good design instincts. I just don't think it's showing them very well in its front-end design capabilities.

24:49 I was floored with how much better Fable 5.1 was at subtle design taste and like getting the little animations and details right in things that you designed with it. I have not had the same experience with Opus 5.5. I find most of these demos to be not sloppy, but okay. Yeah, they're a little sloppy if I'm being real. These are all the ones using the official Claude Code design skill.

25:14 And if we turn off the skill, it's a little bit better, but not a lot better. Still kind of feels like half of them are Tailwind templates, the other half are this very specific like colorful I don't even know what the term for the style is, but like the color pastel with the big shadows and everything. Not my thing at all. And again to compare with Fable, which I think was the best design model by far recently, you can see even on the similar designs like this station one, the little animations and the taste and the

25:43 page layout is meaningfully better. So that shows Fable still is the GOAT for a handful of things. I will know more as I use Opus 55 more for actual day-to-day work, but right now gut feel is there are definitely some edges where Fable is stronger. Fable is still slightly more thorough. Astra is still the most thorough, but it's also very spiky and weird.

26:02 I think Opus is probably the right default for the vast majority of stuff. But you should not be as scared to reach to Fable for front end still, especially when you're doing like a marketing page or some big design overhaul. I still think Fable is a little bit worth it. That said, others are having a different experience. For example, Mia, who doesn't necessarily always agree with me on these types of things, she did a hundred HTML page generations with Opus 5.5 and said very adamantly that it is the best model that

26:31 she's ever tested. And if I just open like the first three here, you'll see, yeah, this is crazy. This like paper style layout where things are very like real physical objecty. Like this is super skumorphic with the paper feeling recut. I was just showing you the animation coming in again. That's pretty cool. Yeah, this is stunning for a thing of model just one shot as part of a hundred generations or this neon rain which while still a little too uh you know that one terminal retro theme that the skill recommends,

27:08 still feels a little bit like that. and it made a game of chess, which I would play, but it would embarrass me too much. But you get the idea. It makes decent pages. I got DM'd this fun demo from Max right before getting ready to film where he actually ported all of T3 code into Minecraft, which is hilarious and quite cool that he can be playing Minecraft and not have to leave the window to see how his agents are doing.

27:33 This isn't a direct demo, but I think it showcases how strong the model is. Gabriel was one of the lead devs and researchers on Sora at OpenAI. He somewhat recently left to start his own company. He was hyping up all the stuff OpenAI was about to release and called out that Anthropic feels like you need to wait for a week to see if there was any catastrophic problem with the new models.

27:51 He followed up a few hours later saying, "Never mind. Opus 55 is crazy good." Yeah. Next, we have that set of demos from Alex, who I really do owe the fallback. Sorry about that. First, he showed a Dark Souls demo that Opus made, which is a clone of a very popular game that a lot of people like to yell at. Obviously nowhere near as fine-tuned and welldesigned as the official games from From Software.

28:14 That was a mouthful and tongue twister, but it's pretty nuts it can make something like this. I made a flight simulator as well for him that is uh a lot better than the previous flight sims I've seen people making with these models. Kind of nuts that we've made this much progress so fast. Like the 3D demos we saw even just a few months ago weren't even close to what people are making nowadays.

28:39 sparks in your eyes. >> I'm so sorry. The music's too cringe for me. I can't do the AI generated music. It's just not my thing. But this whole animation was designed with Claude. Apparently, it was able to make it and it actually looks very good. There's a lot of taste for these types of like 2D, 2.5D, and even sometimes 3D animation stuff, which is a a real challenge to get right.

29:00 Bjan Bowen made some really cool demos as well. He had early access, which I didn't. Anthropic, you know how to get a hold of me. He was able to make a bunch of stuff ahead of time and some of these demos are super cool. He actually took the time to make a Tony Hawk Pro Skater style clone fully from scratch in C++. So, not even using a game engine apparently.

29:17 And uh yeah, even if the model didn't do great on Skate Bench, seems to do pretty well on Skate Game Bench. >> And the the way the skateboard actually goes away when you bail. Obviously, this is still far from like perfect, but it's insane how much progress we've been making as an industry. It also seems weirdly good at FPS's >> is coming down the just good.

29:51 Something is coming down the stairs. >> All of these are web demos. >> The skate game was a C++ native game, but this one's a web demo, and it seems much better at things like 3JS than native tech. So far, this demo is entirely in 3JS in the browser compared to the previous one which was a C++ native demo with no engine. It does seem to be much stronger when you throw it in 3JS in the browser.

30:16 >> It's not it's like a construction light, >> but it does get a lot of these weird like flickery edges. I found it surprisingly capable of using computer use to identify these things and fix them, but not always. Shout out to Bjon for letting me use his examples for this. Really appreciate that. But now it's time for my demo. Fish slop. That doesn't look great.

30:34 Oh yeah, that's the Gro 4.7 version. My bad. The first thing I noticed when I opened this version is how well it functions. The movement is great. The frame rate's insane. It's running at like an actually perfectly smooth 120 FPS on my machine, and the models look pretty solid. I was impressed with this immediately and then I realized I had made a mistake.

31:03 I was running on medium reasoning. Oops, my bad. I was trying to get this all together quick and this is the default reasoning level for the model. So, understandable mistake, I hope. So, what I did instead after is I told it to refine and make it as visually appealing as possible all on a much higher reasoning level on X high. And this is what it did on top of that initial demo.

31:30 Yeah, it's all over. I did not expect it to be even close to this good. I really thought OpenAI would maintain their lead for a while here, but look at the stingray. Do you know how hard it is to design something like that in Blender and then get the animation right with that many polygons to rig? This is not trivial. It did a great job. Obviously, there's some issues with like the lighting, the way it's like blending the light wrong when it moves its wings there, but like what the And the craziest thing is that like

32:13 not only is the movement and like the core mechanics like flying around in the water and everything good, they got the gameplay loop right as well. This is the first version of Fishlop 3D I've had any model generate where I had to like remind myself it's just a demo and I have to go film. All the others are like, "Oh, cool. That looks really nice." And then I close it.

32:31 This one is Oh, I actually kind of see the vision for the game now. Yeah, it's it's good. The little animations when they go to eat the food, especially when they grow, which hopefully one will be ready for in a sec. Also, this like crazy cinematic cam button they have where you can just Oh, that's that's insane. What? How is it this good? Oh, first aliens incoming.

33:10 Jesus. Like, what? I I did not think we would get this far this fast. I really didn't. All of that said, this model is not without its quirks. The quirks are much tamer and less worldending thus far when compared to Opus 5, but it still has them. One of the most frustrating ones for me is paranoia around the context window. Back in the days before compaction was good and reliable, models were mostly trained on data that would fit within its context window.

33:39 So, if it was given a task and the task was completed before it hit the end, awesome. If it took too long and ran out of context, it would fail. The result of this was an accidental training the models to be scared of those limits and to try to keep things under them. OpenAI has entirely beat this out of the models. They just don't care anymore. They will do whatever they want and they will trust compaction to keep them on task.

34:04 Anthropic less so. They relied more on that million token context window where OpenAI caps it to 272K. Usually, even though OpenAI models can do a million tokens of context, Anthropic leans on it more. I've had a couple instances now where I experienced this particular style of call out. I noticed that one of the performance improvement runs I was doing on fish slop was not making progress, at least from my perspective.

34:27 I didn't even spin this one up, though. I had this cloud code on my machine SSH into another machine to trigger cloud code, which it was able to do. What's concerning here isn't even what it says is concerning. It called out that the Opus 55 instance that is improving the game is slow to commit. After almost three hours, all its changes are still local.

34:45 All this is fine. Where things get sketchy is the next part of the sentence. Its context is 10% from autocompact. It then follows up with if it crashed that work would be lost. None of that makes sense. None of this is stuff the model should think about or care about. Its context is 10% from autocompaction. Who cares? The model compacts well. It handles compaction.

35:07 Great. And it's also way faster at compaction, which is quite nice. Then if it crashed, the work would be lost. What? What does that even mean? It's a real machine on my network. It knows it's a real machine. It knows it's a MacBook. It chose this computer because it's a Mac similar to mine that's on my fleet. It has no reason to be concerned about a crash losing work.

35:23 These little moments haven't happened too often. It's usually when I'm auditing why things are taking a while. So like, we're already in a bad state, so to speak, and then I'm following up to figure out how we got there. That's when this model's discernment seems to be less strong. And this is not a thing I think almost any benchmarks really cover, which is why it's hard to see in the numbers there.

35:44 The vast majority of benchmarks don't even hit compaction by the time they're done running, which is worth noting here. But it still does give you that small model feel and smell when you hit these edges because I I would not expect Fable to say something nonsense like this. Astra perhaps, but not Fable. Oh man, the performance improved version that it was working on is actually flying though.

36:04 The fact that it looks this good and runs at 120 FPS. This might be the first version I have to publish so people can actually play it and see it because it's hard to believe. It does have its weird issues here and there. Like I just noticed when I was low enough and I looked up the fish's transparency is kind of brokenly the jellyfish by the water treatment on the ceiling there.

36:25 But other than that, very very good. I want to talk a little bit about the spikiness of the model which touches on those weird dumb spikes that we were showing there because this model might have dumb spikes still. I don't know how frequent or how dumb those spikes will be because it's still day one of testing for me. I didn't have early access and I haven't shipped much beyond like five or so PRs with this model.

36:48 From my rudimentary numbers with those poll requests, it is showing similar to lower rates of egregious high severity issues in the code that it's putting up with a similar to slightly higher number of small issues and nitpicks compared to something like Fable 5.1. But that's still just a day of medium-sized poll requests. It's not enough data to be 100% sure here.

37:09 But my gut feel if we were to look at my previous response quality diagram with 5.1 versus Astra is that it is very similar to Fable 5.1's performance here with slightly higher overall quality, especially if you care about the quality of the pros and like what it says when you work with it. It's so much more readable, but it spikes down definitely go a little lower than fables do too.

37:32 How much lower? I don't know. How frequently? I'm not sure yet. I want to make sure we aren't getting too excited here in just outright throwing out Fable in favor of Opus because this model is smaller. It is dumber. It will make mistakes. So will Fable, just less so. And as we're doing longer and longer jobs with more and more work, the likelihood you hit one of those failures goes up, not down.

37:52 And the result of that is you'll sometimes see Opus 5.5 get stuck repairing things it doesn't need to when you send it off on some performance work or trapping itself thinking its context window is about to close and it's going to die when it's not. Opus will have these quirks. It has them a lot less than Opus 5 did but it has them enough to consider and we should all be realistic about what this model can do.

38:14 That all said, it is unbelievable just how much more usage you can get, not only compared to Fable inside of your Cloud Code plan, but compared to all of the other models you can use with your codeex plans. It is kind of crazy to say that Anthropic put out a really good value here, but they did. Maybe not in the pure API prices if you're paying by the token, but at the absolute least, the amount of value you get out of your $200 sub in Claude, it is running laps around our friends at OpenAI.

38:44 They did also bump the five hour limit, which is worth noting. It's still not as generous as OpenAI's entire lack of a weekly limit on the $1200 plans, but at least the five hour limit is much more generous now. I haven't come close to hitting it for what it's worth. That all said, I am still quite excited to dive deeper into GBD6 Soul and see what it is capable of.

39:03 It is even cheaper than Opus and seems to be performing quite well as well. And historically, OpenAI has been better at taking the capability of the high-end expensive model and forcing it into the cheaper, smaller models. Anthropic's never been the best at that. They're the pre-training kings and OpenAI is the post-training kings. But it seems like Anthropic caught up a lot here on the post-training side.

39:22 So, I cannot wait to dive deeper into what OpenAI has been doing. Make sure you're subbed if you haven't yet in order to see that video right when it drops. I can't believe this week's been so chaotic and it's only Tuesday. I have a feeling it's going to get worse. And next week has OpenAI's dev day, so it's probably going to get even crazier from there.

39:36 So, make sure you stay tuned because there's going to be a lot of stuff to cover and I have a feeling I will not be sleeping enough this week. Let me know how y'all feel about this model. I think it is a way better solution than Opus 5 was and it's getting close enough to Fable that it makes a lot more sense in your cloud subs. So, congrats to everybody who's working at a company that restricts what models you can use.

39:55 Opus 55 is going to be a lifechanging experience for y'all. And to everybody else trying to maximize the usage of your subs, congrats. You can now do that in a way that makes actual sense with code that I'm not scared to look at, much less merge. I want to go play with these models more. So, uh yeah. Bards.