← All transcripts

I did not think I'd like this model Transcript, AI Summary & Key Points

Theo - t3․gg · 2 days ago · Science & Technology · 31:47 · EN

Watch on YouTube

Answer

No. Sonnet 5.5 should generally not be selected directly for everyday work; it is most useful as a fast, inexpensive subagent orchestrated by Opus 5.5 or Fable.

AI Summary

Sonnet 5.5 is faster and cheaper than its predecessor, but its pricing, token usage, and reasoning behavior make it a poor direct choice for most day-to-day coding and front-end work. Max reasoning can increase token usage by 1,500% from X high while potentially reducing performance, so max should be avoided; low should also generally be avoided. Sonnet 5.5 becomes valuable as a subagent that Opus 5.5 or Fable can use for codebase research, architectural analysis, task decomposition, and other investigatory work. In a deep codebase-analysis benchmark, Sonnet 5.5 performed slightly better than Opus at roughly half the cost and completed the work in around 5 minutes. Its game-generation results were among the strongest observed, although the run used more than three and a half times Opus's input tokens and ended at a similar cost. OpenAI needs to respond because Sonnet 5.5 contributes to Anthropic's advantage in model capability, efficiency, price, and model orchestration.

Key Points

  • 00:47 — Artificial Analysis Intelligence Index results place Sonnet 5.5 as more expensive and less capable than Opus 5.5 at every level.
  • 03:23 — Anthropic's announcement gives Sonnet 5.5 a 30% speed improvement and up to 30% lower cost for most work compared with Sonnet 5.
  • 05:17 — Sonnet 5.5 scores 70.6% on Terminal Bench 4, compared with 10.3% for Sonnet 5.
  • 06:07 — Sonnet 5.5 is strong on long-horizon work and image understanding and is the first Sonnet model to beat Pokémon Red when working only from screenshots.
  • 07:29 — Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens, and $0.20 per million cached-read tokens.
  • 10:01 — Cached reads make up under 5% of Fable 5.1's cost, closer to 20% of Opus 5.5's cost, and over 50% of Sonnet 5.5's cost in the cited comparison.
  • 11:36 — Reasoning settings act as token budgets rather than fixed reasoning levels; complex tasks can benefit from extra budget, while simple tasks may use only 5% to 8% more tokens between low and X high.
  • 12:28 — A benchmark showed a 5% token-usage increase from low to X high but a 1,500% increase from X high to max.

AI in practice

Used for

Agents

  • Build a playable fish-themed game from a prompt. 2 held 26:20
  • Use a fast, inexpensive model to investigate a large codebase and help plan complex work. 2 held 24:13

Tools & resources

2 items

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of I did not think I'd like this model — Theo - t3․gg (31:47). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 It's hard to ignore just how well Anthropic's been doing recently. From Fable 5.1 being my favorite code model ever released to Opus 5.5 taking over pretty much everyone's workflows, my own included. I haven't chosen Fable for any tasks for over a week now, which is just like mindblowingly crazy to now with yet another release, and this one seems strategically targeted to screw over OpenAI as much as possible.

00:19 Sonnet is back and 5.5 is looking a hell of a lot more promising than Sonnet 5, which if you don't remember was one of my least favorite model releases like ever. Sonnet 5.5 is a model I did not expect to care about. I legitimately thought I would just kind of skip over this one and not talk about it because I've been so impressed with Opus's price and value and just everything I've been doing with it.

00:41 But I was wrong. Sonnet 5.5 is an incredible model. I am actually very very impressed with it, but not in the ways you might expect cuz when you look at the numbers, it doesn't seem that impressive. On the artificial analysis intelligence index, it comes out as more expensive and dumber than Opus 5.5 at every single level. So, why am I so fond of Sonnet 5.5?

01:02 It's going to be hard to justify, but I think I can do it well. The simple way of putting it is there's finally a cheapish model from Anthropic that makes sense in the modern era, especially when you compare it to offerings from OpenAI that have historically been closer to this price range. This model almost feels squarely targeted at OpenAI, specifically at GPT6 Soul.

01:23 There's a reason I'm filming this video and not a GPT6 Soul video, and it's not because Soul is so impressive. Trust me on that. I'll explain what I mean and more after a real quick word from today's sponsor. Staying on top of what's happening in your codebase has never been harder to do. AI has made contributing to most codebases way easier, but it's made keeping track of what's going on in them way harder.

01:41 I can't tell you how many times I had some weird issue in a codebase. I was like, "What the hell happened here? Who would ever have merged this?" And then went to the notes and saw that it was actually changes I made with my agent that I merged because I was blindly trusting the AI reviewers. Today's sponsor is one of those reviewers. It's Code Rabbit.

01:59 But I'm not here to talk about how good Code Rabbit's reviews are. Spoiler, they're very good. I want to talk about a new feature they introduced which has made actually keeping on top of the changes way easier. It's the Code Rabbit change stack. Now when you have Code Rabbit set up on a codebase, it will give you this simple button that you can click even signed out to see what changed in the PR in a way that makes way more sense.

02:18 Instead of a pile of alphabetically listed files, you get a description of what changed and why. Instead of a jumbled mess of commits that nobody's going to click through, you get this stack of the different things the PR changes with descriptions of what is changing and why. And if you look at this, you can see pretty immediately how much easier it is to read through the changes that were made in this poll request.

02:38 It even summarizes these different parts so you can understand what's going on in each of them. In this part, it's the detection for the package install and making changes to how T3 Code is persisted on your machine. Here is where we fixed the update logic that was breaking for a lot of users. There's even a security blast radius diagram that makes it way easier to see what changes are risky in a given PR.

02:57 I also really like the activity view, which is a super quick way to see what's changing in a PR to catch up without having to scroll through the mess that is the thread view in a given poll request on GitHub. Don't let SLOP take away your understanding of your codebase. Fight back at soy.link/code rabbit. As I was saying, the benchmarks don't necessarily look great depending on how you frame them, but I want to start with what Anthropic said with their official announcement post.

03:19 This is the Claude Sonnet 5.5 announcement. Introducing Sonnet 5.5, the second model in the 5.5 family. The clear upgrade over Sonnet 5. It's 30% faster and up to 30% cheaper for most work. I don't love comparing this model to Sonnet 5 because on one hand, it's not going to showcase the benefits that well. Like the speed difference and the cost difference isn't that big a deal.

03:40 Sonnet was relatively cheap. But also because Sonnet 5 was a garbage tier model and is not a good target to compare against. I said the same thing with the Opus 5.5 coverage. It made no sense to compare it to Opus 5, which is why I'm thankful they compared it to Fable so much. But here again with Sonnet because Sonnet's the smallest that they're working on right now.

03:57 Then there's Opus, then there's Fable. Haiku is eventually going to happen, but uh we'll see. They said it's coming in the next few weeks. I didn't think Sonnet would be here so quick, and honestly, I'm pretty impressed with it. Sonet 5.5 is faster, lower cost as a compliment to Opus 5.5, where Opus is built for complex work requiring careful judgment.

04:13 Sonet 5.5 is strongest at well scoped everyday tasks. I will say it actually goes quite a bit further. Make sure you stay tuned for the fish slop demo. I I'm very impressed with it. That's all I'll say for now. It's also good for fixing bugs and creating polished documents, slides, and spreadsheets. I don't know how many of y'all are actually creating polished documents, slides, and spreadsheets with your models, but if you are, let me know in the comments.

04:34 I am actually curious. And when you're on the way there, if you could hit that little red button, it does help us out a bunch. A lot of y'all aren't subscribed, and you'd be amazed at how much it helps the channel. And it's also a great way to keep up to date. and seeing that uh OpenAI dev day is in not very much time. In fact, it may have already happened is where you're going to want to be to keep up with how OpenAI is fighting back.

04:52 And trust me, they're going to be fighting back. Anthropic also claims this model has a sharp eye for design. This is an interesting call out because uh yeah, I'll show you the designs in a bit. It's not quite what I was hoping for there. And then they call out Haiku 5.5 which is going to be built for high volume and costsensitive applications will be joining Claude 5.5 family in the coming weeks.

05:14 So how does Sonnet 5.5 improve over Sonnet 5? First off it gets the best score on Terminal Bench 4 to date which is kind of nuts cuz it's a Sonnet model but it's a huge jump. It got a 70.6% where previously it got 10.3. Like this just looks silly seeing sonnet 5.5 as the highest score ever in terminal bench. On one hand it makes me skeptical of terminal bench.

05:38 On the other I'm just kind of skeptical of benches at this point. It's really hard to measure the capabilities and the ones I have been doing on my own have on one hand felt even sillier but on the other hand have better reflected my vibes when I use these things. So cover my benches in a bit cuz I was also pretty amused by the scores. So it's seven times better a score for terminal bench.

05:59 which means it's like actually viable for code. It scored slightly below Opus 5.5 on GDP Val, which is a bench I don't really care that much about. And it's strong on long horizon work and image understanding. It's the first Sonnet model to beat Pokémon Reding and working only from screenshots. Feel like a lot of models can do this now, but it is cool to see these like cheaper tier models being able to do a long running task like playing a game.

06:20 Seems like Anthropics had a couple like crazy internal unlocks around their RL stuff recently because previously any model that wasn't their biggest and best. They didn't really get the right vibe is all I can say. Like it's didn't feel like they actually understood the work they were being given. It more felt like they were robots trying to operate in a very specific set of things.

06:40 Now these just feel like slightly dumber, way faster versions of Fable. And I did actually get some of that vibe from Sonnet, which is still crazy to me. It really feels like they saw everybody else distilling anthropic models, got jealous, and decided to do it themselves. And turns out they're good at it now that they've been doing it more heavily.

06:56 Another important piece, collaboration. Not like how well do our agents collaborate with each other. It's more about how well can they write and format their outputs to us. They have been taking on the like clawisms and all of the horrible Claudes stuff that everyone hates and doing everything they can to get rid of it. And as a result, Sonic 5.5 writes way more clearly than their previous generation of models.

07:17 It feels like a better partner for collaboration than Sonet 5. The speed helps too, although I think they pushed the speed bit a little too hard here. The numbers are not as impressive as they make it sound. And now we have price. Sonet 5.5 is priced the same as Sonnet 5 at $2 per million in and $10 per million out, as well as 20 cents per million tokens for cash reads.

07:38 But it typically needs far fewer tokens to do the same work. It's up to 30% less per task than its predecessor. Man, are there some rough edges to this statement. I I'm not going to call it outright a lie, but it is intentionally dancing around some important pieces. in particular, MAX and cash token reads. As I've talked about extensively, cash token reads and cash writes tend to be the majority of costs for our day-to-day agent code work nowadays, which is why the 20 cents per million token cash read number here,

08:10 it's a little bit concerning because all of the other models that Anthropics put out recently, the two Fable 51 and Opus 55, got huge discounts in cash read costs. This model didn't. When you look at the pricing chart at the bottom here, you'll understand why I'm concerned. If you look at input token costs, Sonic 55 is $2 and Opus is $4. So, it's half the price.

08:31 Same with output tokens, $10 to 20. And even the cash rate costs is roughly the same factor here with $ 250 for cash rights for Cloud Sonic 55 and $5 for Opus 55. So, why am I so upset? The first row cash reads. Previously, I was saying that cash reads are a very small percentage of my usage. That was the case for Fable 5.1 because Fable 5.1 dropped the cash read cost by 90%.

08:58 Meanwhile, Opus also dropped it, but only around 60%. Usually, cash reads are 90% off. So, if it's $4 per million input tokens, it'll be 40 per million cashed input tokens. It's $2 per million input tokens, it's 20 cents per cache million input tokens. Pardon me for using the Google AI summary here, but every website reporting on this pricing is garbage.

09:17 In particular, Anthropics. I don't know why they don't have a good like model API dashboard like OpenAI does. I hope that they can throw up a prompt quickly to build that for them cuz this is garbage. Fable 5.1 cost $10 per million in and $50 per mill out which is double the cost of Opus 5.5. However, cash reads are 25 per million in. Remember all these other numbers for Fable 5.1 are double but for cash reads only 25% more expensive.

09:44 Everything else 2x cash read 25%. So sonnet to opus doubles price for cash rights for input and output tokens. Opus to fable doubles again. But the cash read cost doesn't change between sonnet and opus and it only goes up 25% for fable. What this means is cash read becomes a more and more prominent part of your cost when you go down the model ti where it rounds out to under 5% with fable.

10:09 It's closer to 20% with Opus and it's closer to like 50 plus with Sonnet 5.5 because that number is disproportionately large when you look at the other numbers. It would have been really nice if they could have knocked cash read down to like 15 cents or god forbid 10 cents. That would have been insane. But they didn't. Which means that the costs for using this model don't end up being as much cheaper as you might think when you look at these numbers because normal input tokens are barely touched with agentic work

10:39 because we tend to read from cash and output tokens are a very small portion of the overall costs. I feel obligated to call this all out because I've looked at the numbers for doing like forlike work with opus versus sonnet for real code tasks and sonnet consistently comes out as if not more expensive than opus does. In fact, with the artificial analysis intelligence index, Sonnet 5.5 cost roughly the same as Fable 5.1 for realworld tasks.

11:05 Flashbang warning since I know a handful of you like those. We're going to artificial analysis. Yeah, this is insane. Sonnet 5.5 is neck andneck with Fable 5.1 for the most expensive run they've ever done on artificial analysis. That should kill any reason to use this model entirely, right? Kind of. It does kill one thing. Give you a hint. It's one particular word you can see here.

11:26 starts with M and ends in max. Don't use it. I'm trying to explain why I hate max mode so much, and I don't want to make this whole video about it, but I'll do my best to explain here. Reasoning levels aren't really levels. You're not saying you should reason this much. They're budgets. They are allowing the model to reason up to a certain amount. So, when you set low, you're saying, "I want you to do this with a small amount of reasoning tokens."

11:46 When you set medium, high, or x high, you're saying, "I'm okay with you using more reasoning tokens up to a certain point." But you're not necessarily increasing the amount of tokens used. I've had benchmarks where the gap between low and X high was like 5 to 8% tokens. Like it's not a big difference if the tasks are simple. But if the tasks are complex, then having the extra budget can absolutely help.

12:09 So what is the issue with max? The issue with max is that it's not really increasing the ceiling for how many reason tokens are allowed to be used. It's more so increasing the floor for how many have to be used. It's effectively telling the model it's not done until it does a certain number of reasoning tokens. And the result is that I've had benchmarks where from low to X high you get a 5% increase in token usage.

12:32 And from X high to max you get, and I'm not exaggerating, a 1,500% increase in token usage. Literally 15 times more tokens for that one little bump at the end because you're no longer letting it be done with easy tasks. And this can often actually hurt performance because if the model is told to keep overthinking the thing, it's going to overthink the hell out of it.

12:55 It's going to start second guessing itself and is going to start getting wrong answers because you're forcing it to think too much. It worked more like this where it slightly increased the floor. I still wouldn't recommend it, but it'd be fine. But it's not. It's forcing the model to do too much reasoning. And I can prove this very easily. So on 5.5, max reasoning effort was $7.60.

13:17 If we compare to something like, I don't know, Astron Max, $3.26. So, Sonnet was two times more expensive than Astra for the same bench. That sounds insane until you realize Sonnet 5.5 can also be run on, I don't know, Xi, high, god forbid, medium. That looks a little less bad, right? XI is still a bit more expensive than I would like, but it's pretty close to Astra's price.

13:42 It's cheaper, but still more than I would want. But once you go down to high and medium, you have crazy low prices, $18 and 59 respectively for running the same benches. Realistically speaking though, if we go look at the intelligence scores when you bump down these reasoning efforts, sure X high now is roughly Astra level, but Sonnet 55 was roughly three points higher than Astra.

14:06 Do we actually think Sonnet's going to be that much better than Astra? Okay, that I feel bad asking that because realistically speaking, Sonnet does have fewer dumb spikes that I get so frustrated with. So, I actually personally would take Sonnet over Astra for my day-to-day code work. Call me insane. I would just use Opus, but yeah, wanted to call this out because people are looking too closely at the max reasoning efforts and I really don't think anyone should use them.

14:30 Like, I have not been shown a good enough use case for why max effort makes sense for most users. Pretend it doesn't exist. It'll make your life much easier. Back to the benches quick before I start covering the actual intricacies of using the model. In particular, its speed, which I don't love the way it's been reported on so far. Funny coming back here for the last rant because as you can see on some benches like Frontier Code, they put two scores in because XI scored better than Max did.

14:53 Yes, really. Max put it below Soul, but XI put it above Soul. Very interesting. And GBD6 Soul is kind of just DOA, isn't it? I've barely had any reason to talk about it and haven't put it in videos for a reason. It's just not that impressive to me. Hopeful, fingers crossed, we'll get something better in the near future. Everything else here looks pretty good.

15:17 Computer use, it's scoring much better than it did before. Still neck andneck with Opus 5.5. I wish we had the numbers for OS World 2.1 for GPT6, not just Soul, but Astra because I still personally use Astra almost exclusively for computer use. Occasionally a review here and there, but even then iffy. I think it's kind of insane of them to put this chart as the first chart showing performance of the model inside of their reporting.

15:42 First off, it kind of shows terminal bench isn't the greatest measure because Opus 5.5 dipped with max, but it also much worse shows that Sonnet 5.5 is more expensive than Opus on Max. It somehow is outperforming Opus. This is the type of chart that you would see and think something went wrong. Not the type of chart that you would publish is the first chart in the blog post, but sure kind of sad that the only place it is outperforming opus for the cost is in the max reasoning effort, which again you shouldn't use.

16:11 Over to Frontier Code, you can see the same pattern where it plummets on max, but does decently on X high and high as well. Medium and low are pretty big drops. And similar to how I feel about not using max, I don't think you should use any anthropic model on low right now. They actually stopped providing the option to do no reasoning on Opus and I'm assuming now on Sonnet as well.

16:35 Previously there was a no reasoning option where it would just start responding immediately like an instant mode. They don't ship that anymore because these models suck without reasoning. And if they're given too little budget to reason, they will suck even harder as they have proven to in benches like this. So generally avoid low, avoid max. And for the most part, medium's fineish.

16:53 Medium is a lot stronger with Opus than it is with Sonnet. So keep that in mind. Cursor bench was fun. I can look at the official numbers for it here where Sonnet did outperform with high kind of and it does complement the Opus 5.5 curve relatively well. But again, the drop from medium opus to low is just too big. And Opus is already such a surprisingly good value that it's hard for me to justify using much else.

17:20 I also can't help but notice that all of these are score to cost and none of them are score to number of tokens. There's a reason for that. If we hop back over to Cursors Bench and click tokens, you'll see that Sonnet does have yet another greatest of all time score, their token usage in Cursor Bench, where they did 271,920 tokens per task, putting it ahead of even Opus at 218K and Gemini 38 Flash at 162K.

17:49 It's almost double the number of tokens to Gemini 38 at Flash, which is insane. It is more than 5x the number of tokens that GPT56 soul and also six Astra used. So keep that in mind. This is not a token efficient model, which should be made up for by the speed, right? Cuz it's so fast. More bad news sadly. According to Open Router, the TPS for using Sonnet 5.5 through anthropic is around 94 tokens per second, which sounds insane when you're used to something like Astra going at 30.

18:16 They also report Opus 5.5 at around 70. For what it's worth, for my numbers using things through the official subscriptions, I have seen around 100 TPS on average for Opus 5.5 and around 150 for Sonnet 5.5. So, it is faster, but it's also not very token efficient. And result here is that when I used it for something like fish slop, it ended up taking quite a bit longer than Opus did.

18:40 Sonnet took 43 minutes to finish and Opus only took 36. And when you measure how much time they spent generating tokens, it's an even bigger gap from 39 minutes to 27 minutes because again, this model uses a lot of tokens. Let's take a quick look at Witch AI, the design showcase made by Dra that has been very useful to see the capabilities of these new models when they drop.

19:02 Reference have also opened up Fables and Opus' runs. Let's take a look at the first set generation. This one is interesting. It turned the page into like a fake app with the little things on the side to feel more like what the product might feel like. It's unique style and that's something I've seen the other models do. Not bad, but a bit opinionated.

19:21 Let's see what else we got here. This one's a bit rough. And it has some issues with the scroll, too, where it just like rotated from today to last month back to today. These little lines in the background aren't my favorite thing. They hurt the readability. Don't love it. Next, we have this card design, and it's hideous. We got this highlighty one.

19:45 Pretty boring. Don't love it. The animations are awful. Then here we have a weird hierarchy of like a geologic. I don't know the term for this type of like cross-section, but yeah, not great compared to Opus. Definitely a little worse, but compared to Fable, significantly worse. I still think Fable 5.1 is the best overall design model in particular for nicel lookinging front-end designs for your marketing site and whatnot, but I have found Opus to be relatively steerable towards good designs and it follows instructions

20:27 around design much better. I've not pushed the limits of Sonnet for design personally very much, but from all of the demos I've seen here, not very impressed with it front end capabilities. still way ahead of anything OpenAI and especially anything that XAI has, but not the model I'd reach to for front end, especially since in like forl like work, Opus often ends up being around the same price.

20:44 If you remember the intro of this video, you might be a bit confused at this point because it doesn't seem like this model is all that great. It's bad at front end. It uses too many tokens. It's faster, but actually slower because of the token differences, and it doesn't seem meaningfully smarter or cheaper than Opus in day-to-day work. So, what the hell do I like this model so much for?

21:04 Well, to be frank, I don't think this is a model that you or I should be selecting. If you are presented options inside of Claude Code between Fable, Opus, and Sonnet, I really don't think many people should be picking Sonnet. I can already tell how the comment section is going to look after I said that. Well, Theo, not everyone can afford Opus. Sonnet's more expensive.

21:21 Shut the hell up. Seriously, I I don't want to have that argument today. We're talking about how these things operate in our actual realworld usage. It's why I'm excited to say for a handful of types of tasks, Sonnet does prove to be meaningfully more cost effective than Opus. Not necessarily the types of tasks I would send a model after, but absolutely the types of tasks that Opus and Fable would.

21:44 Its strengths come as a tool for our other models to use. We shouldn't be calling Sonnet directly. It should be used as one of many things that Opus or Fable will orchestrate when it's trying to break up real complex work. I kind of made a bench for this for my Grock review where I was trying to find the strengths of Grock 4.7. And I did it by having all of these models do a really big deep audit of T3 code, specifically this giant orchestrator v2pr that's been iterated on for far too long.

22:12 It's hundreds of thousands of lines of code. I gave all of these models a pretty detailed prompt asking them to break up all the things being added to orchestrator v2 to propose a strategy where we can get these things landed into main sooner rather than doing it all through the single giant PR. And I found that at the time Astra was by far the best at doing these breakdowns.

22:33 Fable was not great at it. Grock actually was outperforming Fable. And then Opus, Sonnet, and Soul didn't do particularly well. I have overhauled this since and updated the way that it is judged. And those changes have brought Fable up to a meaningfully higher score. Still below Opus and still far below Astro, which by far had the best plan on how to do these things.

22:54 I still find that Astra is just like uniquely willing to dig into the details for things. So, it performed really well here. The thing that came as a massive surprise to me was Sonnet's performance where it ended up being around half the price of Opus. you know what it should be and also performing slightly better than Opus did according again to my automated judging panel that goes through the changes and proposals to make decisions on various axes.

23:19 So remember this isn't traditional code work. This is a deep dive type task where the model has to go through a large code base and figure out what can be changed in it and how to explain it to someone else. This is the type of task that requires going through a ton of different things to come back with good information. It's not testing the coding capabilities in a traditional sense.

23:41 It's much more analytical, I guess, where it's like trying to make good architectural decisions and comprehend what's going on in a codebase. So, the cost per point here is insane. It is like the best value I've seen in this bench by far. More importantly, in my opinion, is how much time it took because it ended up only taking around 5 minutes to do all of this work.

24:01 And like, yeah, sure, Soul was able to do it even faster, but at like half the score. Opus took almost twice as long, and Aster took almost three times as long. So, Sonnet as a tool that your agents can call on to do this type of research to help it plan and scope work. If you give Opus or Fable a big task and they want to analyze the codebase before starting, they can now call on Sonnet to do that and get results that they're more than happy with.

24:28 This is huge and it's one of the biggest strengths and I hope others start to make benchmarks like this because it is such a good way to see these capabilities and to see the strength of what Sonnet is introducing here. Little peering behind the curtain here. As you can guess, I have been testing a ton of models over the last few weeks honestly and it's been chaos.

24:47 One of my bigger tests is that I've been working on rewriting all of TypeScript like the compiler for the language in Rust. There's already a Go rewrite by Microsoft that's really good, but I wanted to see how much further I could push and also potentially get it working in WASM. This was meant to be a Hail Mary project that would never happen, but Opus has fully unblocked and has it going really, really far.

25:05 But one of the things I noticed is that it left over a ton of the slop that Astra and Soul had made, like millions of lines of it, and it wasn't getting rid of it. So, I told it to halt and go clean up all of the legacy slop. And then I had a thread over here in T3 Code where I'm consistently monitoring changes as they come in. And this thread is one where I found all these legacy things that need to be deleted.

25:26 So I asked for an update. How about now? We should have lots of legacy stuff cleared out. And I scrolled. None of this is talking about how much legacy stuff was deleted. And I was really confused. I was like, what the hell? We delete a bunch of code. Aren't you telling me about it? Is this a regression in Sonnet 5.5's behavior? Maybe it's not as good as Opus at like understanding my intent.

25:44 Nope. I was in the wrong thread. This is one where I was always talking about numbers. So this is actually a really good thing is that it was able to make what I would consider a pretty logical decision based on what the contents of the thread are. It was continuing to operate the way the thread had instead of trying to figure out what the hell I meant when I said should have lots of legacy stuff cleared out.

26:05 It just did I would argue the right thing here which is nice. I was about to crash out about it not understanding my intent but it totally does. I asked it how much code has been deleted. It's about 1.6 millions lines of code have been deleted since I started filming cuz I just kicked this job off before filming. Funny enough. So yeah, the real reason I came here was to grab my fish slot thread.

26:22 Both to see how everything came out in terms of cost, but more importantly to show you guys the actual demo. Ended up costing roughly the same as the Opus one did, but sadly the Opus one did actually use a Fable 5.1 sub agent at some point. So it's score isn't necessarily the most accurate. I guess I do have to rerun Opus 5.5 on fish slop at some point to get even better numbers there.

26:44 But the input token difference is insane. Sonnet 55 used way more, like more than three and a half times more input tokens than Opus did, and the resulting cost was actually quite similar. Although, I would guess if I had not accidentally spawned the Fable sub agents, Opus would have even been cheaper. How are the results? It's one of the best I've ever seen.

27:04 There are some ways where it is worse. Like some of the models are just not as good as the models that we were getting with Opus, but it is still without question one of the best I've ever seen. I would argue in some ways it is better than the Opus one and in all ways it is better than everything I've gotten out of Astra and obviously out of every other lab.

27:24 I remember just like two or three months ago being so impressed that Kim K3 could use Blender at all and here I am now with like a bunch of actual like 3D models that you can tell what they are properly. It's not even like bad. Like these fish are kind of cute and have their own like little cute unique design to them. The coral, the seaweed, all of this is like decent.

27:46 I just realized none of the sound is coming through. So, let me figure out why that is quick. One sec. a few audio issues later, but now at the very least I should actually be able to hear and you can hopefully too. One of the things I I can't help but notice it's like all the anthropic models have this, but Sonnet especially does. It just feels good to play.

28:03 Like you're seeing a 30fps YouTube video of this. I might have to start uploading these for y'all to try though because this is running at a buttery smooth 120 fps on my laptop and like all the movement and all the mechanics actually feel pretty solid and balanced. I actually love the animation for the fish, too. It's adorable. Little uh pelican there.

28:28 The craziest thing is the aliens. Where are they? It says that there's an attack. Where is he? Yeah, there he is. The alien is the best looking I've seen by far in any of these demos. Oh, how crazy is it that like I threw a prompt at a model and then 40 minutes later for a few dollars it spit this out? Like if I had paid API prices would have been $16.

28:52 I've paid more than $16 for games worse than this. Shamefully enough. I was a kid at some point. I bought crappy PlayStation games that were made after like the movies that I was watching. We've all been there, right? But like god damn, for the price of a skin in Fortnite, you can make your own game. And that's assuming you're paying the API prices.

29:12 Obviously, this is super super cheap if you're doing it over your subscription, which uh by the way, cuz I I know a lot of people are always curious what the numbers look like there. I did actually run some cost breakdowns for my own usage across my now six cloud accounts. And from my rough analysis of my real accounts that I emptied, which I killed three accounts in the past few days, you get around $2,300 a week of usage on the $200 plan.

29:38 That puts you at almost 10 grand a month for 30 days with your Cloud Code sub on the $200 plan. That's going to stretch you pretty far with these models. While it might not look like Sonnet's going to get you much further per dollar than Opus, that is the case for general work. If you're having it as the model you select to use for things you do, sure.

30:02 But if you have it as a tool, Opus calls to do things like deep dives into your codebase to find specific behaviors or characteristics or confirming a hunch it has about some thing that you found in an API or all of those types of investigatory things. It is really cheap and surprisingly fast, too. I would bet that if you set Opus up properly to use Sonnet at the right times for these sub agents that the result will be Opus feeling way faster and getting real work done for slightly cheaper.

30:28 That sounds like a pretty good deal to me. So, while I don't think you should actually use Sonnet 5.5 yourself, I think it is an incredible addition to this new family of models, and I'm excited to see how Opus chooses to wield it. I will not be using this model much going forward myself, but I do hope that my Opus is able to because there is clearly value here.

30:47 While not in like front-end building work or day-to-day coding tasks, there is a ton of value in how this model can be utilized to confirm things in real world code bases. And I would assume for real world documents, office, and all that type of stuff, too. That's not what you guys are here for. You're here for the code. And I'm hyped to say this model is useful as long as you're not the one sending it prompts.

31:06 But godamn, OpenAI really needs to respond to this. They have lost in all of the places they are strongest. They don't have the smartest model. They don't have the most efficient model. They don't have the cheapest model. They don't have the best model for calling other models. It's rough for OpenAI right now. And I can't wait to see how they catch up because right now it just doesn't feel like a good value.

31:24 The $200 codeex plan gets me nowhere near as much usage and nowhere near as much realworld code as I'm getting out of my $200 Cloud subs. So take that as you will. Fingers crossed we got some fun announcements coming. I have a feeling that things are about to speed up, not slow down. Hopefully this was a useful breakdown. You can better know how to wield this model. And until next time, peace nerds.