← All transcripts

finally a good small model Transcript, AI Summary & Key Points

Theo - t3․gg · 16 hours ago · Science & Technology · 41:11 · EN

Watch on YouTube

Answer

Yes — Haiku 5.5 is finally a small Anthropic model worth using, but only for high-volume cheap tasks and sub-agent work, not for writing code yourself, and its cheapest pricing only applies under a 100k-token context window.

AI Summary

Haiku 5.5 fixes Anthropic's long-standing weakness in small models. It costs $0.10 per million input tokens, $0.50 per million output, $1 per million cache reads and $0.125 per million cache writes — roughly 75% cheaper than Haiku 4.5 to run and 20x cheaper than Sonnet 5.5 on input and output — but requests exceeding a 100k-token context pay 5x those rates. It more than doubles Haiku 4.5 on benchmarks, moves Terminal Bench from 0% to almost 40%, and beats GPT-6 Luna on code and computer use, though it performs poorly with reasoning off. The best use is not selecting it for coding but letting Opus or Sonnet call it as a cheap sub-agent for searching, categorizing, auditing and testing — an Opus commanding ten Haiku sub-agents finished Anthropic's egg-drop demo in under a minute for 14 cents versus 3 minutes and 47 cents for Opus alone. Anthropic also cut Sonnet 5.5's cache-read price from 20 to 10 cents per million, and 20x ($200 plan) and 5x subscribers now get $200 and $100 in monthly API credit. On Artificial Analysis cost-versus-intelligence data, Haiku 5.5 sits neck and neck with Luna, but a pure value stack is Luna, 6.1 Soul and Opus 5.5; within Anthropic alone, Opus plus Haiku covers nearly everything and Sonnet can be skipped. The verdict: Anthropic went from no small models worth using to one that is sometimes worth using.

Key Points

  • Haiku 5.5 is Anthropic's cheapest, fastest and most capable small model, designed for high-volume cost-sensitive work: summaries, compaction, database queries, classification, live customer support and browser use.
  • Pricing: $0.10 per million input tokens, $0.50 per million output, $1 per million cache reads, $0.125 per million cache writes — about 75% cheaper to run than Haiku 4.5 ($5 per million output) and 20x cheaper than Sonnet 5.5 on input and output.
  • The catch: above a 100k-token context the price multiplies 5x instead of doubling — $0.50 to $2.50 per million output, $0.10 to $0.50 per million input — likely subsidizing the under-100k use case and steering heavy Claude Code users away from bankrupting themselves on it.
  • Anthropic employee Lydia suggests manually setting Haiku 5.5's autocompact window to 100k in Claude Code; the setting saves per model, so sub-agents stay in the cheap tier.
  • Sonnet 5.5's cache-read price dropped from 20 to 10 cents per million tokens, matching GPT-6.1 Soul — previously cache reads made up 30–50% of a bill and cost the same as Opus, undermining Sonnet's value.
  • Benchmarks: Haiku 5.5 more than doubles Haiku 4.5 on most tasks, jumped Terminal Bench from 0% to almost 40% (versus GPT-6 Luna's 16.4% on code), and beat Luna on computer use by almost 40%, making it a solid small computer-use model.
  • Speed runs 100–200 tokens per second depending on provider; about 180 tokens per second observed in Claude Code sub-agents. Prime reported worse results swapping Luna for Haiku with reasoning off for automations — Anthropic models are trained to need some reasoning, so for categorization with reasoning off, Jev is the better choice.
  • On the Artificial Analysis Pareto line, the best intelligence-per-dollar stack is GPT-6 Luna, 6.1 Soul and Opus 5.5, with Mimo V2.6 Pro an underrated openweight addition; within Anthropic alone, Opus 5.5 plus Haiku 5.5 covers nearly everything and Sonnet 5.5 can be avoided entirely.

AI in practice

Used for

Agents

  • Audit and categorize ~1,500 open pull requests on the T3 Code repository, generating an HTML summary of findings. 2 held 27:53
  • Design and test egg-drop contraptions in a simulated environment. 2 held 22:21

Tools & resources

7 items

ANo. 0939
AIAINotes.us AI product

Artificial Analysis

artificialanalysis.ai/

An independent platform for evaluating and comparing generative AI models. It publishes third-party assessments and comparative data on model capabilities and performance.

Mentioned in
6 videos
Kind
AI
CNo. 0016
AIAINotes.us AI product

ChatGPT

chatgpt.com

ChatGPT is a general-purpose AI assistant that can generate movie ideas and expand them into story and storyboard prompts. It can also produce software projects from natural-language prompts, such as a Python dating application with user profiles and swiping logic, and help explain electronics and hardware-development concepts. It is also cited as a general-purpose AI capability that consumer products can package into more personal, character-based experiences.

Mentioned in
88 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

TypeScript
Stars
★ 149,337
Forks
25,438
CNo. 2305
AIAINotes.us AI product

CodeRabbit

coderabbit.ai

An AI-powered code-review platform that automatically analyzes pull requests in repositories hosted on platforms such as GitHub and GitLab, producing review comments, summaries, and suggestions within source-control and code-hosting workflows.

Mentioned in
13 videos
Kind
AI
GNo. 5463
AIAINotes.us Tool

GitHub API

In the AINotes directory

An API for accessing GitHub data and operations. In the cited experiment, it was used to fetch open pull requests for bulk auditing; the video notes a rate limit of 5,000 requests per hour.

Mentioned in
1 video
Kind
Other
JNo. 4815
AIAINotes.us Tool

jq

jqlang.org

jq is a lightweight command-line JSON processor for slicing, filtering, mapping, and transforming structured data. It can compose and pipe JSON output, including results from MCP clients such as mcpc, using an expression language modeled in part on tools such as sed, awk, and grep. jq is written in portable C, has no runtime dependencies, and is distributed as a standalone binary for multiple platforms.

Mentioned in
3 videos
Kind
Other
TNo. 1299
AIAINotes.us AI product

T3 Code

Open source · pingdotgg/t3code

T3 Code is an open-source agent-harness control surface for coding agents running on a developer's computer. It provides iOS and Android mobile apps, a web app, and an Electron-based desktop app for controlling locally configured Codex, Claude Code, Cursor, Grok Build, and OpenCode agents while using the user's existing provider subscriptions. The project can be launched with `npx t3@latest`, which starts its backend and local web interface on the machine. It also distributes desktop builds for Windows, macOS, and Arch Linux, and supports remote access from a phone or another machine. The repository describes the project as early-stage and warns that bugs are expected.

TypeScript
Stars
★ 25,759
Forks
6,631

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of finally a good small model — Theo - t3․gg (41:11). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 Anthropic is pretty good at making big models for coding. We've seen that with models like Fable and Opus and all the way back in the day with models like Sonnet 3.5 and the idea of tool calls kind of making us software engineers use AI the way that we do today. But there's always been a bit of a weakness with their lineup. It's the smaller models. And as cool as Haiku 4.5 was when it came out, it came out in October of last year and it was expensive at the time.

00:24 Now it's just laughable. And I have put a lot of effort into adjusting all of my tooling, all of my agent and claude MDs and all those things to make sure every tool in my system makes it very very clear to all my agents that Haiku 4.5 should be avoided at all costs because it's useless. It's just not very good. That's why I made this post a few weeks ago where I said that Anthropic has no small models that are worth using right now and also Duncan OpenAI saying they have no large models worth using because at the time

00:51 they had Astra and 6 Soul. They didn't have 61 soul yet. And of course, Google gets the fun jab at the end saying none of the Google models are worth using. But the thing I want to focus on here is this idea that Anthropic has no small models worth using. Historically, Anthropic puts almost all their effort into the biggest model and then poorly distills it into the cheaper ones.

01:10 That's why when they had Mythos and Fable, the quality of Opus plummeted and we ended up with 46, 47, 48, and then eventually Opus 5, which was terrible. That was all because of their focus on Anthropic's large models. That curse has been broken though with both Opus 5.5 and Sonnet 5.5, which is why I'm so hyped for today's introduction of Haiku 5.5.

01:34 They've been teasing this one for a bit. And I I had high hopes but low expectations because this is one of the biggest missing pieces. the model the other models use to search for files, to triple check changes, to do all the things that a model shouldn't waste tokens on if the model is expensive. I got to the point where I had taught my Claude to call Soul in order to save money because I didn't trust it calling Haiku 4.5.

01:57 So, how did Haiku 5.5 turn out? I don't want to spoil too much, but I think it's fair to say they made something pretty cool. The price is right, the performance seems solid. There's a handful of tricks that make it a little extra special. Especially considering how small this model is. But before we can get started, in honor of the nature of this small model, we should take a really small break for today's sponsor.

02:21 One of the painful side effects of having such good models for coding is that all of our code bases are getting way bigger and more complex. And while AI review tools are good at helping us avoid bugs, are they actually helping us prevent major security issues? And even if it does find one, is it actually going to let you fix it? Because let's be real, most of these models that are this good just refuse outright once you start doing security stuff.

02:41 That's why I've been so reliant on today's sponsor. You've already heard of them. It's Code Rabbit. But I'm showing something different today. But if you remember a few seconds ago, I said that the models can't really help with the security things because they're not allowed to. Well, Code Rabbit has made all of the deals they need to and found all the models they have to to get real security feedback for you as you're working.

02:59 Not only will they give you real security callouts in your poll requests, they'll give you the tools needed to fix them. There's just like a little fix button that will run on their side with the models that are approved to do this. They also have something that's even more important, which is the ability to run deep scans across an entire repo. You click this little button in the corner, you choose the repo you want to run on, and then half an hour to an hour later, you have a ton of real feedback.

03:19 I ended up addressing over half the things that were found in both of these reviews. And if I hop to this findings tab, I'm going to need my editor to start censoring things because the things it finds are actually quite useful to know about. And I'll be real, some of these are bad and I need to go address them. As our software is getting more complex, keeping it secure is only getting more important.

03:36 Secure your stuff at soy.v.v.v.v.v.v.v.v.v.v.link/codrabbit. Before we dive into everything about Haiku, I want to show one thing quick on Slopalytics. We currently show all of Anthropic's current models. Sadly, we don't have Fable 55 yet, but we do have Opus 55, Son of 55, and now Haiku. Let's just show the last Haiku model real quick. Yeah, hopefully not much more needs to be said.

03:58 Let's go through the announcement. Introducing Haiku 5.5. The cheapest, fastest, and most capable small model we've ever released. Haiku 5.5 is designed for high volume cost-sensitive tasks. It reliably handles quick and repetitive workloads, things like summaries, compactions, database queries, and classification requests. It pairs well with Opus and Sonnet as a sub aent on coding work.

04:18 And since it's also our fastest model to date, it works especially well for speed sensitive tasks like live customer support and browser use. Haiku 5.5 is available at a much lower price than 4.5. On average, it's now about 75% less money to run. Along with this launch, we're making improvements to the value of our model range. We are having the price of Sonnet 55's cash reads.

04:40 Woo! I actually missed that detail. Huge. That was the thing I roasted them on in the Sonnet 55 video. That is a huge change actually that makes the whole lineup make way more sense. That might actually be my favorite thing about this release now. Oh [ __ ] I actually didn't know about that. So, while I sit here enraged that nobody mentioned this to me, because I I'm a huge fan of changes to cash read prices, I have stories I wish I could tell that I can't.

05:02 If you know, you know. Let's just say I've been deep in the world of cash pricing as of recent. And the one complaint I had about Sonnet that also made Sonnet 55 much less appealing than 61 Soul from OpenAI is that the cash read price for Sonnet 55 wasn't just too expensive for Sonnet, it was the same price as Opus. So the thing that made up about 30 to 50% of your bill was the same for sonnet and for opus which meant the price difference between these models was not particularly large if you use them for agentic work.

05:35 Cash reads are essential to making these things reasonably cheap. So with that in mind, let's look at price because this is what's exciting to me. Sonnet 55 went from 20 cents per million cash reads to 10 cents which is the same cash read price as a model like GPT61 Soul. finally making those a little bit closer in price. I hope artificial analysis reprices their charts with this new price so that I could show the different numbers cuz that's a really big improvement especially for agentic work.

06:02 Haiku 45 was that same cash read price and then half the price across the board. Haiku 5.5 is a tenth the cash read price. It's 1 cent per mill tokens read, 12.5 cents per million cash token rights, 10 cents for normal 1m read, and 50 cents per mill out. Where previously it was 5 bucks per mill out and sonnet 55 was 10 bucks per mill out. This makes it 20x cheaper in output tokens and input tokens than sonnet.

06:29 Comically cheaper in cash rights as well. About again 20x cheaper and then cash reads are a tenth the price too. There is a catch though that is up to 100k token context. This is a bit of a weird thing because Anthropic used to charge when you broke 270 or 250k token context and doubled the price. They stopped doing that somewhat recently. That makes a lot of sense for tools like claude code because you don't have to like micromanage the context to keep your bill low and cash reads are so cheap it's kind of just fine.

07:00 But a 100k token limit seems very small especially for like longunning agentic work. I see this as a couple things and also I'll just share in advance here. If you go over 100k tokens, the price doesn't double. It 5x's. So you go from 50 cents per mill out to 250 per mill out. You go from 10 cents per million to 50 cents per million, etc. The reason for that info- wise is because when you have a certain threshold of tokens in the context, that is now much more RAM you need to hold it all, which does make it more

07:28 expensive. What I think happened here is they made the price for under 100K way cheaper, possibly like eating some of their own margin there and they're making up a little bit of it with the over 100K. I see this a bit differently though. First off, I see this as a method of making the main use case, which is quick reads and categorization checks like that, as absurdly cheap as possible.

07:50 Something like Jev is the real competitor to a small model used in these ways where it's being handed text and categorizing, organizing, all these types of things. By making under 100K this cheap, they are making it harder and harder to justify spinning up another API when you could just throw a haiku. And then the other side is to make sure they don't bankrupt themselves when people use this inside of cloud code.

08:09 This is a way of steering people towards the intended use case by making it really, really cheap if they use it on short quick response things, but also making it reasonably expensive. still half the price was before, but not for free like it is in the under 100K range if you're using tools like Cloud Code. It is worth noting that anthropic employees like Lydia have come out publicly and said if you want to keep your bill cheap and not worry about this, you can manually set Haiku 5.5's autocompact window to 100K so

08:40 that you're always working within that cheaper token pricing. Apparently, this is save per model. So, this will only apply to Haiku, including sub agents. So if you go do this in your cloud code, if you're working on API billing or you're on like bedrock or something, this allows you to guarantee the sub aents will never cost a shitload of money. So now we need to talk about everything else.

08:57 All the performance, all the use cases, and of course how it performs in my silly suite of benches. As we can see here, it absolutely decimated Haiku 4.5, getting more than double the score in the vast majority of things. And one that my friends at Anthropic mentioned to me is that they were very happy to see that Haiku was no longer scoring a zero on Terminal Bench.

09:18 They made their way up from zero to almost 40%. Also worth noting that as incredible as Luna is, and I still think GBD6 Luna is an unbelievable model, GBD6 Luna only got a 16.4%. So it more than doubled Luna score for code. So if you're one of those people that thought Luna was decent at code, I question your judgment. I question a lot about you, honestly.

09:37 But at the very least, you now have a model that is allegedly more than 2x better at coding for not too much more money. One of the most interesting numbers for me here is computer use. I've mentioned before that OpenAI has been pretty far ahead in computer use. And also, I don't think these benches are great because it doesn't show the gap as prominent as it is in my experience.

09:55 Haiku slaughtered here. It did way better than six Luna did by like almost 40% or so. This makes this a very solid small model for doing computer use work potentially. I haven't tried it yet. I've heard good things though. And when you see the speeds, you'll see why I'm excited because Haiku 5.5 is seeing numbers from 100 to 200 TPS depending on the provider.

10:19 But I got notified by my chat that Prime tweeted he was using Haiku in some work he was doing with reasoning off for automations. He swapped to Haiku. It was a drop in replacement, but he's actually seeing a worse pass rate and a much much higher fail rate as well as it being noticeably slower than what he was experiencing with Luna. This does not surprise me too much because the big thing that he did here is he turned off reasoning.

10:46 He has both six Luna and Haiku 55 on no reasoning. Open AAI has trained their models to be good at that. Anthropic has trained their models to be bad at that. With Opus and Sonic 55, if I recall correctly, I know they does with Opus. I think they did it with sonnet. They turned off no reasoning. So you always have to let it reason at least a little bit.

11:02 Haiku left no reasoning as an option because people use it for work that they just want a response in quick, but it's not very good at that. I would honestly say that for like categorization work, if you're turning off reasoning anyways, you should just use Jev. I have been playing with the OpenAI decisions API and I have plenty of thoughts to share about that in the near future, but for now just use Jev unless you need like vision and then haiku or Luna makes sense.

11:23 probably Hike or probably Luna cuz it's cheaper especially with reasoning off but uh yeah take that as you will for what it is worth when I tried out Haiku 55 on my machines with cloud code in my normal cloud code sub I was seeing around 180 tokens per second for my just traditional prompting working on things like fish slop which of course we'll be showing fish slop soon so while prime's numbers with reasoning off weren't great haiku's numbers with reasoning on do seem very good in this tier which means it's time to

11:51 hop into slopalytics Let's look at tokens per task because I am genuinely very curious and for good reason. It seems like this is not the usual wasteful [ __ ] show that we see from these smaller models from OpenAI. For reference, Sonnet 5.5 on Max did almost 200,000 tokens per task average putting it at I think the highest ever on artificial analysis where Fable did 78K tokens under half as many.

12:15 Yeah, it was 2.5 times the tokens per task with Sonnet 5.5, which resulted in a cost for Sonnet 55 that was slightly higher than Fable 51. The reason for this is an unreasonable setting. It's right here. It's max. So on Sloppytics, you can click the word max, which makes it go away. And now we have a chart that makes a lot more sense because those max reasoning levels are bad.

12:39 And now we see tokens per task numbers that make way more sense and also cost per task numbers that make way more sense. Raichu 55 is now all the way on the left here, cheaper than anything else because I only have X high and max numbers for it. The max number is still pretty cheap, but more than I would want it to be, especially once you add back in models like 61 soul because 61 soul on low is as cheap as haik coup on low, largely due to the token efficiency difference.

13:06 But let's go back to the cost versus intelligence cuz I think this is fun. You also see 45 haiku in here which I am going to turn off because we don't care about that model anymore. can finally be left where it belongs, the grave. And we'll notice some weird things here. All this data is from artificial analysis, and it is far from representative of everything these models could or would or should be used for.

13:24 But you'll see that aggregate across all their benches. Sonnet 55 performs worse than Opus at the same, if not a slightly higher cost, which meant that Opus and Sonnet aren't as complimentary as one might hope. But if we look all the way to left here with Hiku 5.5, things are looking a lot more promising because even though these scores are close, Opus was almost five times as expensive as the run was with Haiku.

13:49 Pretty solid. I have a fun feature here as well, the Pareto line, which is the best performance at a given price point. And this line moves around based on like what is the best you can do at that level of intelligence. And you can see some very interesting things here. Of course, as you expect, Haiku 55 does very well here with X high and max being the best value in that range within Anthropic's line of models.

14:13 Sonnet 55 sneaks in randomly with high because it is smarter than Opus low and cheaper than Opus medium. But after Sonnet high, all the best value according to this bench is just Opus 55 medium high, X high, max. But here's where I must admit I was slightly misleading you because I only am showing anthropic models right now. In fact, I'm only showing anthropic models that are the latest for each of these lines.

14:36 And also, Fable 5.1 doesn't even light up anymore because there's no reason to use that when you look at the Pareto line here. Ready for where things start to get more interesting, though? First off, we can switch to linear pricing where you see the gap isn't quite as big as it seems because the log pricing makes things seem further away on the cheap side and closer on the expensive side than they really are.

14:57 And you can see here that Sonnet X high is a fine value at 2.75 per task, but then Sonnet 55 Max makes literally no sense at $7860 per task. Do not, if you're using max reasoning levels, I have questions. But I also will say very confidently, do not use sonet 55 on max. It makes no sense at all. But again, still just looking at anthropic models. What happens if we put on like Gemini 4 argon?

15:22 Oh, not much. How about 38 flash? Still not much. In fact, it almost looks like we took the ha coup line and moved it to the right and down. So, we don't really need those anymore. We'll just turn those back off. How about uh I don't know, Grock 47. Still not very interesting. Still below all of this. How about six Astra? Oh, they got a point on. Astra on low is around the same intelligence as Sonnet 55 on high.

15:48 So, that adds to the Pto line. But here's where things get more interesting. Let's put on GPT6 Luna. Oh, huh. That just extended the line a whole bunch to the left because Luna is cheaper than Haiku always. It's also dumber always, at least with the numbers they have. But these are complimentary. This is a pretty stable line. If you turn on the PTO line, it has all of Luna, all of Haiku.

16:13 Then it briefly dips to Astra, then Sonnet, then Opus for the top. And if you were to turn off Astra and Sonnet, you have a pretty clear simple line here for the best value per dollar. But here's where things get messy. And turn out the proto line to make this clearer. I'm also going to turn off Fable and Sonic because they don't really make sense here.

16:37 Now, what happens when we turn on 61? So, oh, that's unfortunate. So on the Pareto line and I guess technically Haiku is slightly smarter for meaningfully more money there but then just barely more going from 21.3 cents to 21.4 cents you get a huge jump in intelligence and if we were to turn Haiku off here oh that's a much smoother line. So, if you're actually trying to optimize based on the benchmark scores for the best intelligence per dollar, your stack would look like six Luna, 61 Soul, and Opus 55.

17:14 There is one more hidden gem diamond in the rough that I'm surprised people aren't talking about as much. Mimo V26 Pro is an underrated openweight model. Pretty damn good for the price, but other than that, nothing really comes above this line. So if you're trying to get the absolute most value and you're paying API prices, six Luna 61 Soul Opus gets you pretty much all of that.

17:38 That said, Haiku 55 is neck and neck. Now are Soul and Luna still better per dollar? Slightly usually. However, if you are trying to do everything through one provider, like you want to just use your anthropic inference because you have a crazy spend target with them and you just use their APIs or you're subscribed to cloud code and don't want to also subscribe to codeex or you want models that understand each other better because you're using opus for everything and when opus prompts solit worse than when opus prompts

18:06 haiku. I don't know how true that is but it's real. There's a lot of reasons that you would want a model in that range in this family. And for once, it's no longer irresponsible to use Haiku like it was before. Haiku 4.5 was so bad and so unnecessarily expensive that there was no justifiable reason to use it. When I had my tweet that Anthropic has no small models worth using, open large models worth using, and Google has none worth using, somebody replied, "All three statements are false."

18:35 to which I replied, "Lmao, this guy uses Haiku 4.5." And then he got torn to teeny tiny pieces in my replies because there is literally no reason to use Haiku 45 at this point in September of this year, but also right now especially. Haiku 45 was barely usable at the time and quickly became garbage and just wasn't touched because Anthropic was too focused on the big models.

18:58 Now they're not. And while I do wish this model pushed the PTO line a bit more, I am thankful Anthropic is embracing models in this range because together these all actually look quite reasonable. Also remember Sonnet 55 got way cheaper with the cache read price change. So that might actually adjust things even more. Good news, Artificial Analysis updated their data which means I've updated mine and now Sonnet is shifted quite a bit to the left which makes it much more compelling on the Pareto line.

19:26 Actually not that much more. It does make this bottom sonnet 55 high feel less awkward as a dip. But yeah, that is improved. Sonnet feels less bad now by quite a bit with that change. Okay, enough of the benchmaxing. Let's talk about the model and how it actually works. Okay, one last fun thing from Anthropics release notes here that is actually quite cool.

19:46 They want to encourage people using cloud code to build more things with the cloud APIs. A lot of people, myself included, were scared with this announcement, thinking this might be the end of using your sub in tools like T3 Code because yes, in T3 Code, you can use your sub with Anthropic, OpenAI, or whatever else. If it's Cloud Code or Codeex installed on your computer, it will just work with T3 code.

20:06 And a lot of us are scared because they had said they would kill that and they haven't yet. But they did a cool thing here, an actually really cool thing where if you're subscribed as a 20x user, so you're on the $200 plan, you get $200 in API credit every month. 5X users get $100 in credit. And this is for the API. So you're effectively able to get an API key like you would if you were paying, but they preload it with $100 or $200 every month to throw at whatever.

20:30 So if you want to vibe code an app in Cloud Code that has Haiku or Sonnet built into it and then send it to your friends, you now have $200 of credit for that, which I actually think is really nice of them. I do I say these things? Yeah, [ __ ] it. I chatted with them a lot about this back in the day because they wanted to make sure they didn't screw up things with how like cloud code SDK and agent SDK integrations were going to end up.

20:52 They were going to do something a lot stupider with this. And I'm thankful they talked to people like me and that they listened because they turned what could have been an L into a really genuinely cool thing. This is if you guys have noticed that I'm being nicer to Anthropus cuz they're actually listening now. And this is a great example of one of those things.

21:09 Ooh, actually one more fun small detail here. They're updating the cloud, Python, and TypeScript SDKs to add support for computer use and browser use. That is potentially very useful for us. T3 Code might level up from this. I just saw Elon reply to them with, "Congrats on another great model. I would be surprised if Haiku 55 wasn't nicer to use than Gro 47 if I'm being real."

21:28 Oh, I missed the artificial analysis jump chart. This is hilarious. Beautiful. From the furthest right end here to actually competitive. They are a bit more brutal showing it is not on the Pareto line, but uh yeah, the data is updated. It's more interesting than they give it credit. Regardless, I'm tired of talking about benchmarks. I want to showcase the awesome things this model can do.

21:54 And not just by itself, but when it's used alongside models like Opus 55 and Sonnet 55. The best use cases for Haiku 55 aren't clicking it in the UI or selecting it in Cloud Code and using it. It's letting your agents use it and letting your applications use it. building it in so when data comes in, Haiku will audit it and categorize it or having Sonnet or Opus call Haiku for testing different things out exploring the codebase to find stuff.

22:19 Enthropic made a really cool demo here where they had Opus 5.5 by itself try to build a safe egg drop in this emulation of like egg drop mechanics that they built. And on the other side, they had Opus 55 plus 10 Haiku 55 sub agents do it. And I think the results here are actually quite interesting. The thing that you might notice pretty quickly here is that Opus is only trying one at a time because it's running single threaded.

22:46 And if you were to have it run multiple threads, it would end up wasting way more money. Where on the right, we have Opus plus 10 Haiku sub agents. And it's able to try more things at the same time and save a ton of money and get to a result faster as well. There you go. It took 3 minutes for Opus 55 by itself to succeed at this task. It only designed 25 attempts and it cost 47 cents to do.

23:10 But with Opus 55 commanding 10 Hiku instances, it was able to do it in under a minute with 86 attempts. Way more attempts because it's able to get more renditions out when it's that much cheaper and faster and only cost 14. So for the types of tasks where the majority of answers are wrong, like if you're looking through a codebase to find one file out of a bunch and you can't just do like a rip grip call to find it, he will be way cheaper and faster because it can do a lot of things at once without causing crazy bills.

23:38 So for the subset of tasks where more attempts is more valuable than smart attempts, where you'd rather try 10 times and verify each result, then spend a little more time trying once and hoping it's right. If the gap between a random try from a dumb model and a thoughtful try from a smart model isn't that big, Haiku 55 is going to slaughter. This is a small subset of work and I am hoping Opus is smart enough to know what that subset is.

24:04 I have not pushed it hard enough to know yet. But I am going to go edit my Claude MD to no longer force everything to be opus or sole sub aents because this is cheap enough that it would make real sense. But you guys aren't here just to listen to me rant about sub agent architecture. You're here to see what the model's capable of. So, let's see some of the fun demos people have been making with Haiku 5.5.

24:25 Multiple people from my community have been throwing together video demos cuz everyone knows Opus 55, the model got way better at programmatically creating video. And this example came from Arsh in the community. [music] [bell] This is insanely impressive for the money. God damn. Yeah, I am very impressed. And this cost 60 cents an API cost to make was built in under 20 minutes.

25:05 That's insane. A lot of why Haiku could make a video that good is that the anthropic models have pretty good like taste built in is the best I can put it. Especially around design stuff. I think they just have a lot of like really good design data and have obsessively trimmed that data set down to give the model good reference points. And you can see that even more so with the front-end design capabilities using witch AI by our friend Dra.

25:28 This is so comically better than Grock. Oh, that's cool. Actually, the little animation for how that text came in. I haven't seen anything like that in one of these demos before, actually. It's different, but it's it's cool. It's like a new thing. This is the blueprinty one. I don't love this diagram. It's cool. It like has hover behaviors and changes things underneath, but uh not my favorite.

26:00 This is kind of cute. Not my favorite. I like the fonts I chose. It's not picking cringe fonts as much as models used to. And this one's fine. If we turn off the design skill, how much worse is it? Yeah. Now, now we're back into classic LM design. It seems like the new Claw design skill helps their models a shitload right now. Yeah, the gap to that to this is hilarious.

26:24 I had uninstalled the design skill. I definitely need to bring it back because it is it is much better now than it was before. These are great. And just like for comparison, we can look at Grock 47 [laughter] with the design skill off. It's even worse. My god, like this one still kills me. I I'm still very very amused by this particular one. Oh man, I'm sorry.

26:53 I just I have to laugh a little. Okay, someone in chat said that they legitimately think they might have made better sites in raw HTML in elementary school than what they were seeing from Grock there. I don't disagree. Regardless, like for the the people who are picking Haiku as their model for costsensitive reasons or like you're using it in the $8 plan with open code and such, this is usable.

27:14 It's not great. Like I'm not going to sit here and pretend that this is all you'll ever need doing front-end design and building applications, but like a model this cheap being decent is impressive. And it means like the $20 Claude sub might be usable for code. I'm not going to spin one up to test it. I'm not getting more Cloud accounts even for experiments.

27:32 I think I have too many already if I'm being real with y'all. But uh regardless of all that, good designs. Out of curiosity, I'm throwing this model at one of my favorite tasks, which is auditing my open poll requests. T3 code has over 1500 PRs right now. Actually, it's just under 1500 PR. It's probably 1500 by the time I finish the sentence because we have so many coming in at all times.

27:53 So, I have Haiku 55 auditing all of them. Specifically, in this case, it is addressing all the ones that have performance issues, like PRs that are fixing performance related stuff. And I have another thread here where it's just going to go through and categorize all of them. Sentally told to use six soul for sub agents and then corrected it later. You get the idea, though.

28:09 This is going to go through all my PRs and make me a nice HTML page showing them when it's done. While we wait for that though, I have my favorite demo. We got a new build, a fish slop. O, this one has some problems. First off, no mouse move. Only moving with uh space, shift, and WD. You got to turn. No sound either. I don't know if that's cuz I have it muted.

28:40 Nope. It's just no sound. Yeah, this is a little rough. This is roughly the same tier as the Grock version, which was offensively bad. I will say for this one, other than the lack of mouse move, there is no like just like straight up incorrect offensively bad thing. So, for comparison with um GBD6 Astra, the demo it made was stunning, but it got the like core movement wrong.

29:12 Like the mouse movement was way too slow and it felt awful and the game plays like core loop was a little screwy and there were like pretty apparent bugs in it. This version doesn't seem to have like like other than the mouse move thing where it just doesn't work at all. Nothing else is like immediately like this is just [ __ ] wrong, which is good.

29:30 It's not It's not making bad decisions. It's just not making impressive ones. Oh, yeah. The corners are [ __ ] That's a good point, actually. Yeah, the way this corner works is just bad. Yeah, I'm taking back what I was saying there. This does have some [ __ ] stuff. It's subtle, but it has it. Also, the timing of that second alien attack made no sense.

29:50 It shouldn't have been there that early. It is not honoring the original source the way I was hoping it would. It's not as like immediately offensively pathetic as a lot of other models do things like it's its lows aren't as low as Astros, but its highs are a tenth as high as Astros, if that's I think that's the simplest way I can put it. That build of Fish Slop was about a dollar of usage total.

30:15 It was a shitload of cash read tokens, 9.73 million of them. A bunch of 1 hour cash writes because it always writes 1 hour, which is annoying. new inputs. Apparently, only 102 tokens because it was cashing constantly, and the output tokens were about 30 cents. 47 51 requests exceeded 100k input tokens, so they used Haiku's higher price tier. If I had capped the context, the game probably would have come out much worse, but it would have been way cheaper.

30:40 Look at that. The first PR review is in. They found 40 performance related open PRs. Most are server or mobile fixes. Only a few are clean and high impact. Top of the list below is merge ready cuz Getto Report's clean. So, it thinks that they're ready just cuz GitHub said so. It didn't seem to audit them thoroughly itself. Replaying a command no longer freezes the server retry went from 18 seconds plus to 6 milliseconds.

31:01 I'll do what I usually do here, which is I open up Opus, I paste the PR, and I say something along the lines of this seems like it is a change worth landing. Can you figure out if that's the case? Fix anything that you think needs to be fixed and get this merged if you're confident it's ready to go. Full send. And now either that PR or something like it will be merged momentarily.

31:22 The full send full stop combo has changed how I build. I'm so happy with it. While I wait for this PR summary view to generate, I want to answer one of the questions I see happening a lot in my chat. Would I use this model for code? No. I I want to be clear about this cuz I think you guys are thinking about these things wrong in general. Coding is not what makes LLMs expensive.

31:43 It's not the process of going from a prompt to a JavaScript file you can execute. That is the majority of your costs. Your costs come from everything else largely. Like once the context is in the model and it's cached, generating the right code file and putting it in the right place is relatively cheap. Getting the context needed to do that is expensive.

32:05 Verifying that it worked is expensive. Writing a shitload of code to test all the edges around it to hook into GitHub so that you can monitor the PR when it's up and all these other things. Those are expensive. Regenerating it eight times because you got it wrong the first seven, that's expensive. Those things are expensive. But the act of going from all of the context and details needed to write code to the code file in your codebase, even on the most expensive Fable, that is a few pennies.

32:30 It is not that expensive to generate 500 output tokens once the input tokens are cached. What I'm trying to say here is I don't get why people are looking for a dumb model to do the code part. The code is cheap. The expensive part is all of the prep before you write the code file and all the verification after. And Haiku is a tool your models can use to do those parts more cheaply is valuable.

32:56 And if that middle part the coding has edges around it like you want to try five versions of the thing and you have a verification system where you know if it worked or not trivially where you can run code to know haiku generating five or eight versions to test could be worth it but for most problems the smarter model writing the code is not that expensive and is preferred in general when the model is unsure about a theory and it wants to vet four things to confirm it or wants to go look at eight apps in the world and

33:27 like random open source projects to compare against or read through a bunch of documentation to figure out some facts about some idea you have. That is when haiku is incredible. But the execution of the plan, sure, if you write a really thorough, detailed plan with a really smart model, a dumb model might be able to implement parts of it. Cool. Awesome.

33:46 You already spent the money on the plan though. O, this is an L for me. I hit a rate limit on GitHub. probably my fault because I was doing a lot of these types of things across a lot of data, but the script saved the error body as though it was API data and then failed because it would hit errors when it ran jQ on that. The parallel per fetches that this run did used most of my 5,000 requests per hour.

34:11 So now I'm locked out of T3 code GitHub integrations for a bit. The PR-Star.json files, parsed, find, and aggregate. So their data is valid. Now it's waiting for the quota to reset. Then it will refetch the 181 PRs and fix the script so it rejects responses that aren't JSON arrays. Then it will recount CI and review. Cool. Except for the fact that it didn't trigger anything to spin itself back up.

34:34 It said that it's going to wait until the rate limits up, but it didn't trigger a wait. We would show that in the UI if it did. So this thread is just done until I tell it to keep going. So yeah, this this is why I don't like dumb models. small models will corner themselves like this, make something harder for not just themselves, but for me, like this just ruined my T3 code experience for at least the next 50 minutes until my quota gets reset so I can go back to like doing PRs inside of T3 code.

35:02 But since my rate limits are [ __ ] anyways, let's show off how I would actually do this. I took the same prompt from before, but I changed two things. First, I told it to use lots of sub aents and workflows to review these PRs and to use Haiku 55 for all sub aents throughout the work. This added one more sentence, minimize work that you do yourself outside of actually getting the data and orchestrating the sub aents.

35:27 And now the other change, I'm going to run this on Opus because unlike Haiku, Opus is a smart model and it should not only be able to not hit my rate limits, it should find ways to work around them, too. Here is where it's going to hit those rate limits and it's probably going to find a way to work around it. Here it's here's obus being smart 1495 open PR is a huge number to tackle a sub agents maybe batches of around 15 PRs per haiku agent would mean roughly a 100 agents running beyond the typical small scale workflow

35:54 guidelines but the user explicitly wants quote lots of sub agents look at that it's actually being thoughtful this is going to take longer than I plan on sitting here to show you guys but uh you get the idea I think that haiku is incredibly useful as a tool that you insert into your code and that your agent can use as well, but I would not select this model as a tool I use in my coding agents because I'm on the $200 tier and I would rather just use Opus.

36:23 And as big as the price gap is, you have to remember that when a model can't do something and it ends up doing lots of things in the process, it might end up using more money because if Opus can solve a problem in 100k tokens and Haiku burns a million before it figures out it can't, Haiku ends up more expensive. So for a lot of work, not only is Haiku not able to do it, Haiku will cost more money as it fails to do it.

36:48 And as things have continued to improve across the industry, in particular cost performance with both anthropic and open AI, we've ended up in a fun place where the price gap between Opus 5 low and Hiku 55 Max is not as big as you would think. It's 20 cents to 55. So it's around 2x more money for Opus if you're willing to deal with low. Also, we got the rest of the data for Haiku 55 because they finally finished running it, which makes Soul and Luna still quite compelling with the Pareto line, but uh it does fill out

37:19 for the anthropic lineup quite well here where have all of Haiku for the cheap and most of Opus for the less cheap. This is a compelling line and if the only models you used were Opus and Haiku, you're probably not missing out on a whole lot. No, Sonnet. I showed you guys Sonnet just doesn't offer much with the Pareto line. Actually, it offers nothing now.

37:43 Maybe just barely on high there. Interesting. But yeah, I I think you can avoid Sonnet entirely and just use Opus and Hiku and have a very very good time with your anthropic models. But that leaves us with a question. When should you use Soul? Well, first up, as you see here, Soul is actually quite a good bridge from Haiku over to Opus. It fits in the middle there.

38:02 But that's not how I would think of Soul. I wouldn't think of soul as for things that haiku is too dumb for and opus is too expensive for. I'd think of soul as a different family with similar intelligence. Imagine that you have a team of people that work on the same things with you every day and you have a really hard problem you have to solve and you ask everyone around you and you all for the most part agree on what the path is.

38:24 That makes sense because you guys work together every day. So you've started to think about things in a similar way. Soul thinks about things in a different way. It's like you're asking your friend who works at a different company for their thoughts and it's a unique perspective and I found OpenAI models bring a unique perspective in particular with their thoroughess auditing code.

38:42 I love 61 soul as an auditor. I have set up skills for all my agents where they use Opus as the thing that does the coding and then they consult with Soul for opinions before they put up the PR and I'm landing code so much faster simply because Soul catches all the dumb things opus might miss and by the time the PR's up it's already been scrubbed by Opus and by Soul.

39:05 So, I personally am sticking to my Opus and Soul duo, but when Opus needs to go collect a bunch of data or read through tons of [ __ ] or just categorize things, Haiku is going to be an incredible addition to my portfolio of models. I won't be the one to pick it, but I trust Opus to do a good enough job picking it at the times where it makes sense. And I think the result of all of this is a very compelling set of models.

39:31 This is going to take too long. I am curious what the results are, but uh I think I've said all I have to on this model release. It's a useful tool. It's a nice change to see anthropic making small models that are actually usable and useful for real world stuff. And it's still honestly a bit trippy seeing an anthropic model so far to the left on the artificial analysis cost versus intelligence right there neck andneck with Luna.

39:58 Actually, I mentioned earlier the like Pareto line stuff. We only had X high and max for Haiku at the time. Now that they have the rest, it's so close to Luna that this no longer feels like a big missing piece of the Anthropic portfolio. And with that, I will go back to my opening statement. Anthropic has no small models that are worth using right now.

40:18 Anthropic now has one small model that is sometimes worth using, and that is a very nice change. And with that, all I have left to say is that I can't wait for Fable 5.5. I understand that they're probably not going to put it out for a while because it kind of is outside of the spirit of pacing the frontier, but I'm so happy with this 55 lineup right now, but I'm really hopeful Fable meets the bar they've set with model releases like Opus, Sonnet, and Haiku as part of this line.

40:45 Anthropic is dominating across the field now. And if I was OpenAI, I would be very, very scared. And I would be praying that the next training run for Astra comes out unbelievable because right now I'm honestly struggling a bit to use my codec subs because my anthropic and Claude subs are just so much more useful. Let me know if you guys agree if I'm overexaggerating this gap or if it really does feel that way to y'all. Until next time, he snirts.