Claude became around three times faster in two weeks because Anthropic gave Claude access to usage data, instrumentation, benchmarks, deployment monitoring, and the ability to propose and implement performance changes. The gains came with caching and data-persistence regressions that still required human testing and judgment.
Chrome DevTools is a set of web authoring, debugging, and performance analysis tools built into Google Chrome and other Chromium-based browsers. Maintained by the Chromium/Chrome team at Google, it provides DOM inspection, JavaScript debugging, network monitoring, and performance profiling for web developers.
Claude is an AI assistant developed by Anthropic, positioned as 'The AI for Problem Solvers'. It is a general-purpose AI system used for tasks such as generating code prompts, refining requirements, creating advertising strategy and copy, and processing creative content like storyboarding and video prompts.
Claude in Slack is an AI integration that makes Claude available within Slack.
Datadog is a cloud monitoring and observability platform for collecting and analyzing metrics, traces, logs, and related data across infrastructure, applications, networks, containers, databases, and cloud environments. It provides infrastructure monitoring, application performance monitoring, log management, dashboards, alerting, incident management, security monitoring, synthetic monitoring, and user-experience monitoring. The platform is used to identify broken or correlated systems and investigate operational issues, although the video notes that observability alone does not necessarily establish root cause. Datadog also offers AI features, including investigation, remediation, agent observability, and an MCP server.
TanStack Router is an open-source web application routing tool in the TanStack stack, used to handle navigation in web applications. TanStack describes its tools as headless, type-safe, composable, framework-agnostic, and built with TypeScript.
WorkOS is an identity and authentication platform for adding enterprise access features to applications. It supports Single Sign-On through SAML and OIDC identity providers, social authentication, passwordless email codes, multi-factor authentication, and user and organization management with configurable policies. The videos also describe it as supporting agent authentication on behalf of users. Its APIs normalize enterprise integrations behind a single interface, with dashboard management, webhook events for directory-service updates, and SDKs for Node.js, Ruby, Python, .NET, Go, PHP, and Java. WorkOS provides AuthKit, an authentication UI built with WorkOS and Radix, and the site states that the platform supports more than 20 enterprise services through one integration point.
Searchable transcript of How Anthropic made Claude 3x faster — Theo - t3․gg (01:10:29). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 I know most of y'all think of me as that AI dev vibe coding YouTuber, but believe it or not, a few of you were around back when I was more focused on full stack application development, in particular the website, and really focused on performance and endto-end capabilities with systems. I love that type of stuff. I love nerding out about the details of good quality engineering, data loading patterns, performance issues, browser stuff, all those types of things.
00:20 That's why I'm so excited about this new article that Anthropic just put out about how they made Cloud.AI AI three times faster in 2 weeks, largely by using clot. This is a very fun one for me for a bunch of reasons. Not only am I a performance nerd, I've also been dunking on anthropics performance and general engineering quality for a long time now.
00:39 It seems like Anthropics models are now finally good enough that they're actually able to help with performance instead of hurting it. Which is particularly fun because I posted a video not that long ago about performance issues in T3 code that were caused by Fable that I had to go solve myself. So, I'm very excited to read this because one of the things that's made T3 Code so awesome has been that focus on performance.
00:58 And if Anthropic is finally going to stop using 18 gigs of RAM when I open Cloud Desktop, maybe it's more worth a look. I also just want to see how they're setting Cloud up for success in these types of environments where you're trying to debug deep performance issues. It's a not it's a non-trivial thing to do and a lot of the tools that exist, I'll be real, kind of suck.
01:18 So, I'm very curious how Anthropic was able to get the profiling set up such that they could make the desktop app as well as the website perform meaningfully better all by using Claude. I do have one more thing worth using though, and while it might not be three times faster, it's still worth a listen. It's today's sponsor. Weird question for you. You want people like me to use your product?
01:36 Because if you do, you probably shouldn't do a normal sign-in flow because I'm not going to be clicking those buttons myself nowadays. I do most of my work through my agents now, which means any platform that my agents choose needs to be usable by the agent. In particular, it needs to be able to be signed up for by your agents. Today's sponsor is Work OS.
01:52 And that might confuse you because they are a platform for user signing in, right? Well, they are. They are the best way to have humans and businesses signing into your app. So, if you're trying to get enterprise customers, then you should probably be using work OS. That's why everyone from OpenAI and Enthropic to FA, Cursor, Bolt, Vanta, Carta, and more are all using them.
02:09 But I want to talk about OMD because this is a super super cool standard. The goal of OMD is to make it easier for agents to sign up for your apps. It's a new standard that's been designed by works in partnership with companies like Firecrawl and Cloudflare. And the results speak for themselves. There's already a ton of companies adopting these standards like Neon, Monday, Parallel, and more.
02:25 What makes it so cool is that agents can now register on behalf of a user. So if you're trying to get business customers and you want their agents signing up, this is a great place to start. Get off for all your users at soy.link/workos. Time to dive into how Anthropic made Claude three times faster in two weeks. My honest guess is that the models are much smarter now.
02:46 So, the quality of Anthropics engineering is actually improving for the first time in a while. But super curious what tooling they set up and what they gave Claude access to in order to make this possible. Once Claude can measure something, it can make it faster. So, we kept finding more things to measure. I like this opening. Although I am concerned this might be the new Claude Pros because they did a lot of work to make Claude less cringe.
03:09 But if this becomes the new Claude format, there will be a new cringe sentence structure. They got some nice little animations on the site, too. It's cute. But let's dive into the actual details. This August, we made the core user experience of cloud.ai and the Cloud Desktop app around three times faster in a 2 week sprint. Users have been telling us it was slow, and they were right.
03:28 Hi, it's me. I'm users. Yeah, it it was bad. I did absolutely notice the difference here when I was going to claw.ai mostly to like check the status of my account and my usage before I had my dashboard set up. And I noticed the day it got way faster. It was like a meaningful difference. I didn't think I had a post about it. While that's hunting, we could keep diving in.
03:46 We focused on four journeys that made up 95% of user activity. At the 75th percentile, time to attage on a fresh load for Claude went from 3.1 seconds to 0.55. And starting a new cloud code session went from 0.8 seconds to 0.3. And loading cloud co-work cloud sessions went from almost three seconds to under one second. In aggregate, we estimate that this saves tens of thousands of user hours of waiting every single day.
04:12 This is why it's important. Oh, that's funny. I actually noticed improvements in June. So, they definitely made a bunch of improvements at the time. I even called out that it felt faster on a mobile hotspot than it used to feel on gigabit network connections. I also found when I dug deeper into it that it was using Tanstack now, which was really cool to see.
04:31 Specifically, Tanstack router makes it much easier to handle these types of navigation cases. But that makes me even more curious what they did because they already did the framework upgrade. So, how do they get this performance squeezed out after? Let's dive into the core user journeys. There's launching the app. They had both the web and desktop versions improved here.
04:50 It was, as mentioned before, 3 seconds to launch the website before. Now, it's 550 milliseconds. Desktop app used to be over 6 seconds and now it's just over three. Starting conversations, another thing they focused a lot on and they were able to make huge improvements there. Loading conversations they made huge improvements in. And sending a message, specifically cloud co-work went from a second to 48 milliseconds and cla code on desktop went from 250 milliseconds to 52.
05:13 So this is all in the apps by the way, not the terminal version, but the cloud code desktop app and the cloud code website. I am actually curious how much we feel the difference. That's still slow as balls. Still waiting. Still waiting. At this point, I think it was a CDN failure. Content at CloudAI login may not load or link to file. Great. Refresh.
05:36 The refresh worked. I guess those performance improvements didn't happen on the signed out page, but uh now we're good. So, let's take a look. I'm going to do a refresh. That was pretty fast. Some test prompt. This is one of the ones that pissed me off the most is the flow of submitting a prompt would create the new thread but it wouldn't navigate which meant that if you refresh there was a chance it would disappear forever.
06:00 That came through very very quickly actually. Let's do another and it immediately refresh. Ah, did they fix that? No, that thread disappeared. Yeah, that's here's that bug. I hit it again. This thread that I have open right now did not make it to my recent threads. So, I have this one from before. This one doesn't exist. Make the title proof of my theory.
06:27 I'm going to send this and then refresh really quick. It's hard with one hand. Cool. Did that said what to make the title getting set up for the session. is really thinking on this. But I did manage to kill it getting added to the sidebar at the very least here. [laughter] It had a failed tool call. Yeah, this thread effectively doesn't exist because I refreshed before it would get persisted to the sidebar, which is uh yeah, rough.
06:55 It does let you submit the prompt pretty quick though. Like if I refresh, okay, I tried pasting that whole time. So watch this. I'm going to refresh and then start spamming paste. It took about 700 milliseconds or so. A bit under a second. Maybe about a second actually. Oh, that one was brutal. That one took like two to three seconds. While it's still not meeting my bar, I am impressed overall with the difference.
07:22 It does feel a lot nicer to navigate. I hope they can figure out Oh, look there. My threads finally appeared. Took them a bit, but they did. But yeah, you get my point. It has a lot of like data persistence layer issues still, but it does feel better. I see people in chat saying they need convex. Unironically, yes, it helps with so many of these things.
07:41 I don't know how well it would handle the insane load that I'm sure that Anthropic pushes, but something similar to it with a proper engine that keeps stuff up to date would help a lot. A few months ago, I did a deep dive on how both ChatGpt and Cloud.AI I were handling caching for the data that was being fetched on the page that you actually saw and learned that both are now caching the recent threads in the sidebar through local storage or I guess now in this case now I'm digging into it index DB is a way of using a
08:05 key value store to store the threads you have in the sidebar. All these threads here are all cached when you load which is why when I refresh they come in almost immediately and there's no pop in or flash in. To show you a really silly example of like why this matters, I'm going to go to my other browser and archive a handful of these. Uh, I guess there's no real archive, so I just have to delete.
08:30 That's fine. Delete. Delete. Cool. So, I deleted those two. So, there's still one test prompt. Watch what happens when I refresh. They're not invalidating the cache. It still thinks these two threads exist when they don't. How do I get this to refresh? If I click that doesn't even do it. What? I was hoping this was going to be a video about how much Anthropics improve their engineering quality.
08:52 I'm almost relieved to see nothing has changed. What the I I have these two threads that when I click, you get an error, but it never revalidates the sidebar. You know what? I bet I bet if I send a prompt here that will Nope, that updates the status for that one. But these two are still broken. Chat just said coding is solved, by the way. Yep. Yep. Coding solved.
09:13 AGI is here. [groaning] They dis No, they didn't disappear. They're still here. Two broken threads. Are you kidding? I I hope you'll understand why it's a convex shell now. It solves all of these types of problems. Even T3 code doesn't do this. We don't even have a server to sync against. It's all on your devices. Oh, they disappear. When I was complaining, they just poof vanished all of a sudden.
09:35 I was even looking. God. Reminder that according to Sam Alman, I'm an anthropic shell. So yeah, take it as you will. >> [sighs] >> Well, that answers a bunch of my questions. They made performance better by making it so that things never get updated. Clever. They're caching far too aggressively, and they're not invalidating the cache anywhere near often enough.
09:58 For reference, in T3 code, we do this a bit differently when we're in the loading state because we haven't fetched the new data from the backend yet. I show it in a faded blurry state because I don't want to show you confidence in data that isn't accurate yet. The fade and doesn't trigger until we have loaded the real data from the server. Did I say T3 code?
10:18 I'm sorry. T3 chat. My bad. You guys know what I'm talking about. T3 code does everything through servers that you're running yourself. So, it's very different here. The caching like caching just works fundamentally different when there's one database per user and it's on your own machine. For T3 chat, you need to do more. And as incredible as Convex is, and it absolutely is.
10:36 I'm so happy we moved our data over there. It still doesn't give you data as soon as the page loads. I want data as soon as the page loads, which is why we set this up with the local storage cache so that we can show that immediately. I did also notice though that we changed how we are loading in things on the bottom bar. We need to start caching that as well.
10:53 On that note, actually, let me um new thread T3 chat. I notice that it seems like we're not caching the necessary data in order to open the model picker and other things at the bottom of the composer as soon as the page loads. It's kind of jarring to have the sidebar come in immediately, but the model picker and other selectors in the chat take some time to load.
11:17 What would it look like to cache that similar to how we cache on the sidebar? If it's relatively easy, I would like you to implement it and file a PR and babyset it. As much as I like the dunk on anthropics lack of engineering quality, I think that we should go through this article because at the very least they did make a lot of improvements to performance.
11:33 And while they seem to have regressed a bunch of data patterns in the process, we can still learn from the ways that they gave Claude access to the things it needs in order to find these performance issues and fix them. Let's start reading though cuz I am actually excited. Apparently, they did the whole thing through Claude tag. So, they didn't even use cloud code or I don't know something better like T3 code.
11:53 They did this all through Claude Tag, which is the Slackbot. I have seen examples from Boris, the guy who started Cloud Code, using Claude Tag to do all of his work every day. He actually showed an example of a real thread he did to add the mods feature to uh codeex or to cloud code if I recall. I could be wrong on that, but he did something like that recently.
12:10 So, they used Claude Tag using an internal research model roughly comparable to Opus 55. Huh. I hope this is just a snapshot of what became 55 and not some other model that they're claiming is similar to 55. Like they claimed Opus 5 was similar to Fable when it wasn't. But uh yeah, no guardrails though, which is a good call out from Maria because when they're testing the models, they have much less guardrails generally.
12:31 And I happen to know that when you start diving into Chrome internals, it becomes much more likely that you start to hit the safety guards and the stuff that prevents you from sending prompts because it thinks you might be hacking. So yeah, I will put that all out there because I think it's important because I've had that experience. I am hopeful because again I haven't really pushed Opus too hard here.
12:53 Actually, I'm kind of lying there. I did have one instance. Let me find it in my terminal quick. I vaguely remember you mentioning that the other agent hit some type of security flag or guard rail when it was working on the performance improvements. Do you have more info on that? Like what it was trying to do at the time and what guardrail it hit? It was a built-in safety check on remove commands that I hit.
13:15 So, not as bad as I thought. I assumed I was hitting the guard rails because of the weird things that Opus was doing in Chrome to debug performance stuff for fish slop. It does appear to be the case that that isn't what happened here. So hopefully we won't hit too many of those guardrails. I've seen it before with the other models. I haven't done as much deep diving in performance stuff with clawed models lately.
13:36 Honestly, I have still found OpenAI models to be better at this type of uh I still call it the like Rottweiler way of working where it grabs the problem by the neck and strangles it. I have found anthropic models to not be quite as good as OpenAI models in the GPT line for these like deep dive tear apart all the performance problems type stuff. But a lot of this does just come down to the tooling because codeex has much better computer use stuff built in and it seems like it's able to figure out how to control Chrome
14:01 and get these traces out of it a little bit better. Bullets still aren't great at it from my experience though and I find myself still digging in to the details myself when I'm deep in performance stuff. So this is all why I'm so curious how Anthropic did this. They said that through cloud tag, they were able to find bottlenecks, build benchmarks, ship improvements, and watch every deploy.
14:19 They steered by setting goals, making trade-offs, and improving every change. With that approach, they merged more than 3,000 changes without a single customerf facing incident or roll back. To be fair, I did just kind of showcase things that might have been worthwhile uh a roll back or an incident, but uh to each the own. This is a post that covers how they did it apparently safely.
14:38 Very interesting. Before the sprint, they created a new channel in Slack with the following standing instructions. At Claude, your job is to facilitate all things related to the performance of the Claude AI website and desktop app. Your responsibilities include monitoring deploys for performance regressions, assessing the accuracy and comprehensiveness of existing telemetry, maintaining well-curated observability dashboards, proactively implementing solutions for observed issues and low-hanging fruit, proposing
15:03 performance project opportunities, and communicating with your human teammates. The ultimate goal for this channel is for you to become as autonomous as possible, but today we know that isn't yet possible. This is a terrible prompt. They didn't even mention to make no mistakes. No wonder it had those regressions. If they had just told it make no mistakes, why would it had no problems?
15:21 We asked Claude to analyze usage data through the data dog MCP server. They got that dog in them. It identified the four highest impact user journeys. The app launches, starting conversations, loading existing conversations, and sending messages between web and desktop, as well as across our products. Those journeys came to 13 distinct measurements.
15:38 To establish our baselines, we added instrumentation until they were directly comparable. Each started with a user interaction, ended once the result was rendered, and disamiguated client and server work. People underrate how important this part is to fixing performance issues. It is much harder to fix what you can't measure, even if you're not an AI agent.
15:57 I cannot tell you how many arguments I would get into in my previous jobs where I was like, "This thing is slow as hell." They're like, "Well, we looked in our dashboard. and only took 200 milliseconds. And then I went and profiled it myself and learned that they were counting from not like when the page was loaded to when the thing happened. They were counting from like when their component rendered in the UI to when it went from loading state to done state.
16:20 And if it took 8 seconds to get there, then your trace doesn't matter. And so many places get their profiling wrong with all of these types of things. I cannot even comprehend how many times I have been shown data to prove me wrong about a performance issue just to learn that the data was not actually reflecting the real world performance at all. Now let's see how they actually got the work done once they did the instrumentation.
16:40 We kicked off the sprint with a list of about 20 handpicked projects each targeting specific journeys. Claude estimated the impact of each project in milliseconds and we aggregated those estimates to set our targets for the sprint. Some of the projects were fairly large but we thought we could probably achieve most of them within 2 weeks. They hit 12 of their 13 targets by day three.
17:00 Did they fall for the classic mistake? I was planning on having Opus do the classic rewrite of ping.gg to make it modern. I'm curious about a thing quick though. How long do you think these changes would take if we did them all in one pass? Here we are. The the classic anthropic overestimation. As great as models are now at writing code, understanding code, their understanding of agents is still kind of garbage.
17:23 And as such, when you ask them how long will it take to do something, they're wrong. This is an overhaul that I'm pretty sure it could do in a few hours, and it is very confident it would add up to about 25 working days to do it. It is just wrong, but it's cute that it thinks it isn't wrong. I'm going to throw this one a a rough instruction here. I want you to kick off a agent run on my computer left book.
17:53 You should be able to see how to connect through the fleet repo on this machine. I want you to make sure this repo is cloned. All the environment variables that are needed are there and that you can run opus 5.5 through cloud code in T3 code such that I can still see the thread how I normally do in T3 code. My goal here is to do the whole implementation on that machine in one pass with however many sub aents and whatever else are needed with full computer use capability on the other machine because I don't want you to
18:21 be spamming my machine with computer use requests and browser control while I'm using it. If anything does not work as expected when you move this workload over to Leftbook, don't hesitate to ask questions so we can get it working right. I'm actually curious how that work goes. We'll check into that in a bit. But yeah, point I was trying to make by pulling this up is to highlight how bad models are at estimating time frames cuz they don't understand how capable models are, which I honestly find hilarious.
18:45 I'm just trying to do an extended version of the joke that Claude told them it would take 2 weeks and then it actually took 3 days cuz these guys make Claude makes sense. They're bad at judging this too, as we were. The planned projects landed early. For faster launches, we baked a static composer into the HTML so users can type during the React initialization.
19:02 That's a fun one that we did a while back. They also pre-ompiled the V8 code cache so the desktop shell main process doesn't have to recompile it from scratch. Oh boy, they're going deep in the weeds on that one. H, do I want to add that? You know what? Anthropic has been focused on overhauling the performance for cloud.I and the cloud code desktop app.
19:22 I noticed this paragraph in particular has some fun ideas on how they improve performance. I'm curious if any of these would make sense for us. both the embedding the HTML so the composer is available faster but more importantly the V8 cacheed compile I'm curious if we'd be able to implement something like that in T3 code and what type of performance improvements we would see if we did very curious how that goes another fun Reacty change is that in order to make navigation faster they kept the composer mounted between
19:49 conversations and they also prefetched sessions when the user hovered over them and in doing all of that also cut cyber rrenders by 90%. These are the types of things I was doing by hand just barely a year ago in order to squeeze every bit of performance I could out of T3 chat. Just some fun real examples here. I'm just going to click a random thread so you can see how fast it loads.
20:10 That was not as good as I was hoping. If I go to a different one, skill manager name. Okay, that was immediate. That's immediate. That's immediate. I want independent status. Immediate. Yeah, data fetching on Hover is one of the most fun little tricks you can do to make things feel absurdly fast. It's actually crazy how fast things will feel when you make a small change like that.
20:34 I am surprised they hadn't done it before. They also called out that they left room for Claude to identify opportunities and propose new work streams. Those work streams quickly ramped into full projects of their own, which far exceeded their initial targets. So, they set new targets and looked for more things to measure after. We've ended up funding nearly every project in the original project list and more funding in completing.
20:52 Curious the wording there. Let's do a refresh. What have we not explored and what can we hill climb on? Where's the most opportunity at this point? I am open to wacky ideas. Don't underestimate the value of that particular piece of the prompt. I know it sounds crazy. I'll be real. It's because it kind of is crazy. When I was testing GBD6 Astra early, I tried something similar here.
21:15 I was doing a huge performance overhaul for Lakebed, my cloud, trying to really see how far I could push performance on it. So, these small number of user tiny JavaScript apps in the cloud would use as little resource as possible. It did a bunch of runs and the performance was better after the first performance set of improvements that I merged, but it wasn't quite where I wanted.
21:34 So, I wrote a prompt very similar to the one we just read. Sounds like there's more improvements to be made. I want you to really dig deep and find everything we can, even extreme measures to improve performance. I want you to come to me with three different proposals for things we could do. The more extreme the better. Don't be afraid to boil the ocean.
21:50 And this is why I now have a custom runtime for LakeBed. I ended up forking Rusty V8 and building my own runtime in order to make the performance as good as possible. And it was like a 2 to 4x improvement. They proposed the persistent V8 execution service which ended up being its own custom runtime. Build a query compiler and maintain live answers and give each app a durable owner with local data.
22:08 didn't end up doing those other two because I just got so focused in on the persistent V8 execution service. Rusty V8 is a phenomenal library by the way. Thank you to the maintainers of that. But yeah, I effectively told it to just go hill climb, come up with something crazy and the classic boil the ocean and it did it. It made something awesome and the performance is crazy.
22:27 And for those asking where's Lake Bed, it's become my pet project to play with these things with. So, uh, making it like a real product that I can confidently pitch to y'all is less a focus for me now. But it is out there. It works. If you find the URL, which I don't hide very well in the npm package, which is pretty easy to find, too. You can probably use it for real stuff.
22:46 I use it all the time. Hell, I used it yesterday in my Opus 55 video to make a better visualizer for the data on the uh, artificial analysis dashes. So much easier to read this than it is to read the artificial analysis homepage. It's a fun little thing. It's very nice to just like set up apps that autosync and live update and get stuff out. But the performance improvements I got out of this really stupid prompt and also a lot of time spent working with it to build both ways to verify the performance improvements as
23:14 well as ways to verify the behavior not regressing. It was a a war. This is one of the longest threads I've ever done in T3 code. And a lot of pretty crazy run times, too. Like the planning took 20 minutes. The first build was about 2 hours and then something broke. I think I was doing something on the machine so it stopped. I told to continue. Five and a half hours later it had the PR and it actually did the thing and it ran awesome.
23:40 So yeah, don't underestimate something as silly as I am open to wacky ideas in a prompt like this. And that's why the next quote here sounds insane, but I I'm on their side here. Anything can be hill climbed. They're not wrong. From the start, we knew we wanted to iterate faster than our deploy cadence. Claude could work asynchronously for many hours, even overnight.
24:00 We wanted to let it validate its prototypes without waiting for field reads. To achieve that, we looked for other ways to measure performance in the lab. Sam found the first lead. What can we do instead of wall clock timing? Can we measure JS instruction counts for instance? Yeah, for pure JS hot paths, literal instruction counts, run the benchmark under valindrind with node predictable and compare to a checkedin baseline.
24:22 One run, no statistics needed. Okay, so we know this isn't Opus 55. There's one little hint that this wasn't Opus 55. Thank you for including such a useful hint in topic for browser paths. There's no instruction counting under Chromium, but there's a ladder of other deterministic counts. React commits per interaction, function call counts from V8's precise coverage, layout and style recount counts, DOM mutations.
24:46 What do you want first? He said, "Let's explore val grind plus IR and predictable flags in one thread in each of the browser React benches in new threads. Ping me in all of them. You know what we want. Let's go." I I I hate to be so fixated on this. I have not read too many prompts from employees of the labs. I don't know how I feel that my prompting is so similar to the prompting of anthropic employees.
25:10 I am a combination of proud and ashamed in this moment. Audio made fun of me for the way I prompt. I feel better about it now. Anyways, 11 minutes later, five threads were running, each focused on different measurements. Why it take 11 minutes to spin up five threads? Regardless, instruction counts, V8 call counts, React commits, style recalc, and DOM manipulations.
25:30 I think I know what happened now to those weird issues I was calling out before. If you're just blindly measuring the number of React commits or V8 calls, you're going to start turning things off that don't actually affect performance. The harsh reality is that when you're on the page and the page is working, firing more network requests is not too big a deal, especially on like desktop web and desktop app experiences.
25:55 I think both people and agents are a little too quick to look at every single network request and try to disable them rather than find ways to change the hierarchy of requests to get you what you need to interact with the page first and then keep things up to date and reliable after. And the reason I'm calling that out now is giving the agent this type of insight is encouraging that type of behavior because you're not going to see when the page becomes interactive via V8 call counts, but you are going to see every
26:25 request happening whether or not it's blocking. And if you see opportunities to reduce those requests, like I don't know, maybe we shouldn't fetch all of the sidebar data all the time. We already have it in local storage. That would kill six of the 800 calls on this page in this trace. we should turn that off. And then you end up in the state that I showed earlier where the app wasn't syncing the sidebar content properly and you could end up with bad sidebar runs.
26:47 Not the best solution here. It is useful tooling, but you need to also make sure the agent is combining that with real introspection to how the user flow works. I don't necessarily know if agents can do the whole end to end here yet. That's why I still do performance stuff very much in the loop myself. I'll usually start by asking it to come up with theories to improve performance for things that could be causing a specific thing to be slow.
27:13 Like I'll start with it takes too long to start typing in the input box. What do you think we can do to improve it? Then I'll tell it to go make demos with all of those potential improvements and then I will test it myself to see if it feels faster or not. It is far too easy to cause weird niche UX regressions in performance overhauls. And I still do this a lot myself.
27:31 They then asked it to prove that hill climbing against each of these can actually result in measurable wall clock performance wins. We'll unship the benches for any candidates that cannot prove that. Interesting. They also called that they treated every new benchmark with skepticism. Each one had two jobs. First, make a metric that Claude could move in the lab.
27:49 And second, a guardrail in CI with a number that could only ratchet down. This is one of my favorites. Benchmark was flaky or didn't actually correlate with user latency. We threw it out rather than letting Claude climb the wrong hill. Also very good they called that out there. I hate that I'm just going to like keep deep diving into my own stuff, but that's what happens when we talk about performance.
28:08 I get to nerd out about my favorite I somewhat recently did a lot of work in T3 Code's data layer to make it so you don't have to load as much data when you open big threads or you resume the app. That's a big part of why T3 Code feels so fast even on absolutely horrifying hell threads. For example, this build Swift UI app from scratch thread. This thread I've been working on for 4 months and it probably has like 500 prompts in it at this point.
28:32 Watch how fast it loads when I click. I am clicking now. Pretty much immediate. It did have like a flicker down due to a weird scroll state thing. Fun problem for me to fix later, but things load pretty much immediately in T3 code because I had to go the hell and back finding every single place I could squeeze out data that wasn't necessary in the loads.
28:54 I made a lot of these improvements and then noticed things starting to get worse again. And I didn't want that. And I also wanted to be easier to measure when I made improvements. Which is why the most important change I made actually had nothing to do with the performance improvements itself. The thing that mattered is this guy. I added an action that would do a test for a bunch of different request formats that actually exist in the app with some stored data for a real in quotes fake thread to make sure that we are
29:23 not meaningfully going above or below the existing data like baseline that we have for loading stuff and this has prevented so many regressions. I am very thankful I added this because not only will the GitHub action catch these types of regressions by putting it in as a comment, both the agent working on the code as well as the agents doing PR review will notice it and call it out and potentially even fix it autonomously.
29:50 So this gives a signal to the agents when they make mistakes that cause performance regressions so they can fix it before it becomes my problem or worse my user's problem. And once again, I want to emphasize, even if you're not that into using vibe coding into AI coded stuff to ship realworld projects to users, if you're not using it for some of your work, you are missing out so much because using these tools to slop together tests like this where like, yeah, if this was missionritical to have our performance never
30:21 regress, it might be worth really digging into these tests to make sure every detail of them is good. But that was not the mindset I went in for this. It was a much more vague, wouldn't it be cool if we had a better idea when regressions kind of happened. And it threw together this test suite. I think it's a couple thousand lines. I don't know. I didn't read the code.
30:38 And now it has actively prevented regressions in our codebase. So once again, slop is useful even if you never ship it to your users. Back to the article. Wall clock time is what users feel, but it's noisy and milliseconds are too flaky to use as CI gates. Instruction counts were appealing because they were deterministic, but we still needed Claude to prove they tracked wall clock time.
30:56 We asked Claude to drive the countdown on two hot paths. The routine that assembles a conversation's message tree and a scanner for status lines in Claude code output. Claude profiled both with Valgrren and found that a quarter of its first paths instructions were megamorphic dictionary lookups resolving the same message ID three separate times. Going to be so real with y'all on this one.
31:19 Data loading patterns tend to be the vast majority of the problems with these types of things. Like if it is going down the same path three times, that means data is being fetched in a way that triggers rerenders three times usually. An hour later, it cut instructions on both paths by 48% and 31%. And while clock time had dropped from 78% and 44%. We checked in two new ratchets.
31:41 From then on, an APR that raised the instruction counts of these paths would fail CI. the daily job lowered each ceiling whenever the count went down. That's another good idea. If you do make improvements, it's important to update these things because if you don't, then you end up in a weird place where every performance improvement becomes area for somebody to make things worse.
32:01 Again, this is even more painful with unit testing and code coverage. I hate code coverage, period. Like the idea of like having a number that shows you how well tested your code is is silly and dumb. And I had times at Twitch where I had a feature that I rewrote. We took like half a million lines of code and rewrote it in 50k and we had shipped the new version with a feature flag to trigger between the two.
32:23 We had moved everyone over to the new version and it was great. And then I went to delete the other 500,000 lines of code and couldn't because when I did that, the code bases code coverage would go down from 95% to 94. It was not like it was like 94.9 or whatever, but it was enough that it made it so I couldn't merge the PR. Well, you should add more tests to your changes then.
32:41 I did. My 50k lines had 100% coverage. The 500k lines had 96%. But that was enough where deleting it would shift the code base enough to knock us just below the threshold. And I never got to delete this giant old unused feature that we had rewritten because of code coverage. Does the count track the clock? CPU instructions versus wall clock time. Yeah, CPU instruction was 48% less.
33:07 wall clock was 78% less for message tree assembly. That's really the best way to measure that. Also doesn't give us real times for wall clock here. So we don't know if it matters cuz like if you shave 200 milliseconds to 60 milliseconds that matters a lot less than 4 seconds down to half a second. This led them to the central lesson of the sprint with Claude.
33:27 Measuring something makes it tractable. Yeah. Makes it so you can actually give Claude the things it needs to improve once it can measure. Measurement used to be step zero. You'd add a metric, wait for data to roll in, and only then would you start to understand the problem. With Claude, it's just step one of the climb. As soon as Claude has a number that it can beat, it can start optimizing.
33:46 This meant that the highest leverage thing they could do is find more things to measure. This is why I was excited to read this article. I want to see the new novel ways they found to measure things. The loop thread by thread. This is a little cringe. Let me All of this ran in the same Slack channel with multiple engineers and Claude jamming in every thread.
34:03 From there, the sprint settled into a loop. One, someone would open a thread about a slow stretch of the journey, often with a screenshot or recording. Two, Claude would trace the flow and then find or build benchmarks that demonstrate the problem. Three, once it had promising results in the lab, Claude would come back with a PR or sometimes several sized more mdash is great.
34:22 Once it had a promising result in the lab, Cloud would come back with a PR, often several sized for risk and review with anything user visible behind a flag. After it shipped, Claude watched the deploy and read the field data. If performance improved, Claude locked in the win by ratcheting the benchmark down. If not, it turned the flag off and iterated accordingly.
34:42 Then it went looking for the next slow spot in the same journey. So I I'm split on this flow. On one hand, nothing matters if it doesn't actually improve performance for real world users. On the other hand, if you're not testing these things yourself manually as the engineer prompting for the changes and merging the PRs, you're not going to notice the types of regressions that users won't report.
35:07 Like, let's be real here. How many users would have actually reported the sidebar issues I just found in Showcase? Like, who is using Claude in two browsers enough to notice these regressions? And who would be annoyed enough to bother like reporting that to Anthropic in a way that they would even notice? If they're not testing it themselves, you're going to have these types of regressions.
35:24 That's why I put so much work in T3 code to make it easier for us as the T3 code developers to spin up builds to test ourselves. For example, I added a label on GitHub that I can apply to a PR that will trigger a manual build of the desktop app for Mac OS. So, I can just download the DMG and test it directly without having to like fetch the branch, build it locally, and spin it up.
35:46 This means I can even take one of my other machines like my MacBook Air I use for testing and go to GitHub, download the DMG and test it directly there without it affecting my machine. I think these details are just as important if not more so than measuring realworld user performance. And I'm amazed how few people have flows for this. I've talked to everyone from like small startups to independent devs to big businesses that are just outright failing in this way.
36:12 Like if they want to test a change in a PR, that's a 15 to 25 minute process and it's so frustrating. If you're going to put slop somewhere, it should be there. Slop up ways to test your changes yourself with as little friction as possible because that friction makes you hesitant to do all of this in the first place. I am very hyped this made it in here.
36:29 Someone shared a screen recording that showed sidebar rows popping in after the page loaded. Chat and co-work rows resolved at different times, making the page feel janky. None of our existing monitors detected it. The closest that we had was a cumulative layout shift, but each shift only scored about 0.008, well within the good threshold of 0.1. Isaac had the idea to reference the underlying layout instability API directly.
36:52 I don't think you need to add a bunch of measurements for this. You just need to cache that one query. Uh like why are you hitting the like deep layout API for this? You don't need that. I'm very thankful they recently hired Addios Mani from the Chrome team because he is going to clean up a lot of the slop that happened in this arc. Ah, anyways, apparently when they were doing that, they also added integration tests that opened the page with a populated sidebar, held the sidewars data until after first paint, and
37:24 failed on any shift in any named region. Oh, this might be the problem actually. If the data in the browser has some things in it and the data from the network has different things, that will be a layout shift. If you have four threads cached locally and you have a fifth thread that it loads from the network, it will shift the four threads down when that fifth thread comes in.
37:45 Their tests here might have actually caused the regression because doing this right would fail this test. I I love trying to reverse engineer where these failures come from. After the event deployed, Claude read the field data and found that 31% of web page loads moved something after the page was usable without any user interaction. From there, Claude worked through the causes by name.
38:06 A header row that arrived late, a carrot that slid sideways once the user's name loaded, a list that moved when the scroll bar popped in. Claude fixed the top offenders as a batch. And when they were gone, it found the next batch. I do love these types of comparisons. Do you see how much stuff is moving around on this left example here? Yeah, that's rough.
38:30 This is one of the big benefits of using cached data on page load. The catch is if the cache data is not accurate against the real world data, you're going to have some shift. But I think that like the chats and tasks shifting down one or two items because the real data loaded is not that big a deal, especially if you have a fade treatment on that for when that change happens.
38:52 Someone in chat just asked a very good question. How do you avoid this type of issue without looking at the code for testing? I'm gonna phrase it back in a pretty silly way. How was I able to identify this issue with testing without ever seeing any of the code at all? Because I don't even have access to the repo where they wrote this. The blog post that the anthropic team wrote here is probably, I would guess, slightly less detailed than the responses Claude gave them in their threads in Slack.
39:16 I don't know how you build these skills now because the way I built them was going deep in the details to fix all of these things myself by hand for almost a decade way before AI was even a thing you could use for dev. And I am noticing these issues because of my history doing these things by hand before. I don't need to see the code or read the tests to know that.
39:37 I can tell from the PR description or from the response Claude gives me in my Claude code or from Astra in codeex. A lot of this just comes down to yeah this you said it branch intuition and you have to build that over time. I didn't have to read the code to have these potential realizations and I still could be wrong. Of course I might not be correct in that these measurements are enabling or encouraging these specific types of regressions.
40:02 But it's a theory that if I had codebase access I could ask Claude to go prove out for me and then use that to tell the team like hey I like where you're going here but these tests kind of suck. here are the regressions that we're going to cause if we don't shift how we think about this. So yeah, reading the code does not solve this problem at all. You got to read the responses from claude and understand what it is doing when it does it.
40:23 If you're actually reading all the code when you do these types of tests, you're not going to get very far. Because I'll be frank, when I was doing all the data loading stuff in the T3 code like overhaul, I probably wrote a thousand lines of tests in slop for validating my findings for every 10 lines of actual code I wrote. It's probably a 100 to1 ratio.
40:41 And the code I wrote that actually shipped would have been worse if I hadn't taken the opportunity to slop a shitload of tests to verify my changes. And as far as I know, there were no regressions from those overhauls, which is actually surprising to me. I expected it to be pretty brutal and wasn't at all. So yeah, take it as you will. I don't know how you build this intuition now.
41:00 Maybe you just ask cloud questions. I don't know. It's worth hiring somebody who has it if you don't have someone on your team with it and you care about these things and it's probably learnable still. I just don't know how I'd recommend to do it. The next chunk of this blog was actually one of the more fun parts. Scaling horizontally. Once the loop worked on one thread, running it on more was just a matter of opening them.
41:18 It's one of the really fun things about setting up a cloud environment or, I don't know, using something like T3 Code and the remote machine management stuff. Once you have an end-to-end flow that is working for figuring out a class of issue or bug, you can spin up a bunch of threads on a bunch of instances on a bunch of machines and immediately parallelize it.
41:38 We're all hopefully engineers and we've experienced this before. Writing code to do a task once is rarely worth it. But once you wrote the code, which might take hours, even if the task took five minutes, once you've written all that code, you can now run it again and again and again. You can parallelize it. It can do a bunch of and if it turns that 5minute thing into 3 seconds and you can run hundreds of them at once, it fundamentally changes the way you operate with that thing.
42:02 So yeah, sure I put a lot of effort into the testing I built for my performance improvement changes, but once I had built all of that testing, it was trivial to fan out lots of agents with lots of theories to try and solve lots of things. And the result is that my agent automerged like 40 PRs improving performance in T3 code. And the only real regressions were one change to animations on the actual desktop that people didn't like, which I understand.
42:28 I didn't like the change either. And it broke the sliding ticker for tweets on the homepage of the marketing site. But other than that, even Astra and its weird spikiness. 38 out of the 40 PRs, it merged itself. Real performance wins. Pretty cool. Once you have the testing all set up, this is an interesting detail, though. Instead of closing threads once the original request had been fulfilled, Claude would keep going.
42:47 An individual thread would put up 50, sometimes a hundred optimization PRs. Interestingly, it was Claude, not one of us, opening new threads to chase opportunities it had found on its own as parts of separate investigations or nightly jobs. Shelley, one of the engineers in the channel, observed, quote, "The model is a numbers demon." Yeah, it's surprisingly easy to set something like this up.
43:08 Every measurement found something to improve. Claude ran a React hook census and found 6,900 hooks in 900 store subscriptions in the composer's typing path, rerendering on every keystroke. [groaning] Nearly 7,000 hooks for the composer. This is what h maybe we should read the code, guys. I I take it all back. Anyways, Claude counted style recalculations and found a single root has selector that added 24 milliseconds to every single DOM change.
43:40 You know how many of these types of things there are? Like one simple CSS thing causes every change in the browser, even in Electron, to be way slower. It's way too easy to cause these things. And Chrome's tools do not make it easy enough to find these things, sadly, especially GPU stuff. Chrome's okay at helping you with CPU stuff, at event loop stuff, anything JavaScript side.
44:00 Once you're in the GPU accelerated CSS layer, good luck. Have fun. I'm going to choose a different mindset for how I'm going to approach anthropic engineering going forward. Now, this might get me in trouble. I don't care. Anthropic has chosen to be the first company to really really experiment on what happens if you let AI write too much of the code.
44:25 Since they did this in the sonnet 4 days, they have a code base that is full of slop and garbage. They learned a lot of lessons. They made the model much better. And now for the betterment of everyone else, they are using the models to clean up the slop that they made themselves. The cla site and the cloud desktop app and I'd argue even cla code itself are them dog fooding both how much slop can something take and also when are the models good enough to clean up the mess that we created.
44:54 They also found a code path that had a leftover location. Which is like a browser refresh that would cause half a million hidden reloads a day that none of their load metrics could see. Claude read profiler samples from idle tabs and found identical cache snapshots being cloned into index DB twice every minute all on the main thread. Oof. This actually might also be the change I was talking about with the sidebar issues.
45:15 If they made it so that loading data from the network doesn't write to the index DB storage as much, that would prevent that storage from being updated. Great point from Jamie from Convex here. Companies that are good at making models are not necessarily good at software engineering. This is a lesson playing out slow as I'll go even further than this, Jamie.
45:34 There was a deep, as I'm sure a lot of y'all know, a deep rivalry between OpenAI and Anthropic because Anthropic was largely started from the original pre-training team at OpenAI. Daario and friends were all helping build GPT3 and they left because of a lack of alignment between them and the leaders at OpenAI. They still say a lot of this is around safety concerns.
45:54 I don't actually think they were that concerned about safety in the GBT3 days. I and many others have started believing the real friction here was the difference in engineering versus research where OpenAI despite being a research company was led by two engineers Sam Alman and GDB which meant that engineers were the ones commanding researchers to go do research stuff and that meant they were holding them to engineering standards.
46:19 because they were expecting engineering deadlines and timelines and all of that type of stuff. And that caused disdain between the two groups. Since those guys left, it definitely caused some frustration and open AAI towards researchers and they started looking for ways to engineer solutions to research problems. Meanwhile, at Enthropic, it meant that they kind of didn't like software engineers that much.
46:40 And you can still see that in their coms even now. I think a lot of why Anthropic pushed so hard towards this idea of what if we replace all the engineers with AI is because they didn't like engineers, which means they probably weren't hiring the best ones. And I'm going to be very real here. It's some of the advice I give the most. Hire for what you love, not what you hate, because you'll make much better hiring decisions.
46:59 Since they were hiring engineers and they don't like engineers, they weren't making great engineering hiring decisions. It's difficult to automate things you don't understand, and it's difficult to understand something that you don't respect. Absolutely agree, Jamie. Back to the performance improvements because I I love this stuff. We rarely knew where a thread would lead.
47:16 In a sweep for CPU hitches, Claude noticed that highlighting a finished code block could freeze the page for about a second. Oof. Syntax highlighting in the browser. Synthe general is one of those like really annoying problems to get right. It dug in the lab and found the culprit. M dashes. Oh man. Oh man. I I saw someone say in chat that there was an M dashes section.
47:39 I I was not prepared for replies markdown contained any non-Latin one character like an M dash or a curly quote. V8 would store the entire string as UTF-16 which put every syntax highlighting reax on a slower two byte path with a 20line change to copy each code block into a one byte string before highlighting it. Blocking main thread for 0.35 seconds is in my opinion always unacceptable.
48:05 We moved our syntax highlighting off to WASOM because since it now runs in a worker, it won't block main thread at all, which is quite nice. Write me five code blocks in various languages doing some simple demo that's at least 10 lines of code and I'll put them in the chat here so I can see them. You might notice the fade. I added the fade to hide the fact that I was doing the actual highlighting in a different thread to not block the main thread.
48:37 So you can still navigate around T3 code and it still flies even when those things are happening in the background. It's able to syntax highlight various different languages all in a row while you have a bunch of other stuff going without it affecting the main thread's performance at all. So yeah, they should probably be doing that too. The catch here is that you have to load a bigger bundle for that syntax highlighting, but if you lazy load it anyways, it comes in later once and you're fine.
49:04 So yeah, I'm surprised that they didn't get Claude to propose that to them because it would have helped with a lot of these problems. I would argue that since we're not reading the code anyways, main thread being blocked at all for syntax highlighting is bad. Apparently, by the second week, they could barely even summarize their output into daily updates.
49:21 On the busiest days, more than 200 changes would land. Claude kept proposing new benchmarks. About a third of PRs included additional telemetry or guardrails, and each new instrument generated more threads with more opportunities. Working in one channel meant everything happened in the open. We jumped in and out of each other's threads to debate decisions and celebrate wins.
49:39 Word spread. Other teams started to bring their changes into the channel to have them reviewed for performance. New projects were written in subtly more performant ways because of the guard rails that had been added and the new cloud skills that had been introduced. I do really like this idea of gang prompting at companies, like having a channel that can be used by various different people across various different teams with shared context and capability.
50:02 This isn't reading like an ad for claude tag in Claude and Slack, but it is a very compelling one. It makes me want to go set up something similar for us with all the T3 codeb stuff that Maria's been doing. And now we are in the guard rails. We' prepared for the pace because nearly everything we touched was a hot path. the first paint, composer, transcript, etc.
50:21 We'd established our safety mechanisms up front. Every PR went through automated reviews with at least one human approval. Unit tests always came before optimizations, and anything that could cause a user visible problem shipped behind a short-lived feature flag. I don't believe you there. Regardless, once the flag started to pile up, they opened a thread to coordinate their rollout and cleanup.
50:42 Claude classified every flag as a kill switch or ramp and retired each one as soon as it was safe. Across the two weeks, they introduced nearly 200 flags, more than half of which were already cleaned up by the end. That is nice. Cleaning up the slop as you add it is an underrated thing. New performance wins decay in a fast-moving codebase and code chips fast and anthropic.
50:59 Once a project proved a win, they invested in ways to protect it. Static Composer, for example, is brittle by design. We show users an HML copy of the page almost immediately, and we let React paint directly on top after. If the React render is off by even a pixel, all of the magic is lost. So Claude built dozens of guardrails. Static markup is generated by rendering the real React component in JSDOM and tests guarantee that they never drift.
51:24 Integration test suite compares the static page against the React render across 14 different viewport sizes and then asserts alignment within one pixel. A keystroke test types straight through the handoff and fails on any lost reordered keys. And in the field, every handoff report shifts to a tenth of a pixel. Cloud opens a thread for any event with nonzero movement.
51:41 interesting that they have like threads auto open when users experience things by the sound of that. Not everything could be caught in the lab though, so they had to make use of the oldest guardrail in the book, incremental rollouts. High-risk changes were rolled out to employees first, then 1% of users, then everyone. For hours after we released the static composer internally, a teammate shared a screen recording of a layout shift that none of our metrics could see.
52:01 When he opened cloudi in a new tab, the composer would drop mdash, but it wasn't our code. This is a Chrome extension I'm calling it now. Chrome resizing the page, not the handoff from the static composer. On the new tab page, Chrome draws its own 56 pixel footer below the page managed by enthropic.com customized Chrome. When the tab navigates to Claude AI, the footer goes away and the page gets 56 pixels taller, but only 100 milliseconds after the first paint.
52:30 Okay, so it wasn't a Chrome extension, which is the cause of most of my weird hellish issues other than Brave. Brave causes so many problems for us. That all aside, this is weird managed profile stuff that causes niche bugs within Chrome's rendering itself. They are putting the Oh, this is actually probably the bigger problem. They are rendering the composer based on a percentage of page height and a distance from the top rather than a fixed location on the bottom.
52:55 By doing that, they need to fire a JS calculation, which means they are now reliant on the page resize behaviors. And it seems like that didn't or get triggered on time or worse it triggered after the 100 millisecond like repaint occurred or the 100 millcond shift and then it repainted and moved down possibly even before the JS loaded in. Fascinating.
53:15 It was because of Chrome speculative loading. While a user was typing the URL in the address bar, Chrome would pre-render the page in the background and the height of the current tab. Huh. Okay, that actually is is weird, but it makes sense. Yeah, when you're on this page, there's a certain amount of space available. So, if there was something on the bottom here that caused this window to be smaller and you type cloud.ai/ new, this is now preloading in the background and when I hit enter on it, it doesn't have to load
53:43 as much as a result of that. But Chrome's pre-fetching triggering a behind-the-scenes render that is at a different page size is fascinating. And I would guess this is a nonzero part of why they hired our friend Adiasmani from the Chrome team. It's actually one of the more fun ones that they have in here. I think that was that's a really cool finding.
54:02 Yeah, it's from the Chrome pre-render. If I was using Claude in Slack and I had a thread like this and I saw Claude come in and point out a Chrome behavior I didn't know about, I would be the most insufferable human for at least a few hours. I'd be like, "What the actual This is AGI." They also have a section at the end here about steering, which I think should be quite fun.
54:24 The loop was productive, but it wasn't autonomous. Keeping it fast, safe, and on track was our job. And it had three parts. First, ambition. By default, Claude is careful about scope. Hey chat, do y'all feel that Claude is by default careful with its scope? I want a yes, no, or LOL from all of y'all. Thank you guys. I think the point's been made. Regardless, yeah, helping it know it can push further helps a lot.
54:54 You'd be amazed how much like one small change to your claude MD will affect these behaviors too. Like if you add to the claw MD something like assume you have my permission to keep going. If I state what I want the outcome to be, he'll climb till you get there. It's much more likely to actually get there. They called out here that they were confident in their guard rails, which meant they didn't want it to be so reserved and careful.
55:17 So a lot of what they did, especially early on, was encouraging Claude to be bolder. Here's in one of the threads. Claude said, "One small PR to give code the same timing marks that Chad and Co-work already have. I'll put it up this week. Realistically, code's number waits a few days for merge deploy and a baseline window." To which Raymond had to respond, "If you put it up right now, I will get it merged and deployed.
55:36 We have the power to do anything. Please be braver." Pier will be up within the hour. God, Claude is still so bad at time estimates. Speaking of which, look at that. It actually spun up my thread on my other machine in T3 Code to do the overhaul of uh ping. Nice. I don't even have anything set up in T3 Code in this instance to allow a thread to spawn another thread.
55:56 I just got it to SSH into another computer, clone the repo, copy my environment variables, copy the plan MD file that we had made, and then prompt from one machine to another, spin this up, and then it reappears automatically with the remote connections in T3 code. Pretty cool. People aren't prompting boldly enough and they're not telling their agents to be bold enough either.
56:15 And I can already hear the comment section. I know how this goes. Well, Theo, some of us work in real code bases where we can't just merge, slop, and break things. First off, so does anthropic and they break more stuff than I do. Second off, more importantly, the focus of this isn't on merging a bunch of breaking changes. The point of this is proving that you can use the slop to measure things to make regressions and problems less likely and also make it easier to make improvements.
56:44 I have a feeling the comment section on this video is going to literally kill me. So, I'm just going to move on. I do like this call out that they put at the end of the section there. When we started hitting the targets we had set, we noticed the threads would slow down. Sam went thread to thread with the same message. Let's keep driving this down. The targets are not the stopping point.
57:01 What's next? Be ambitious. I know y'all think this is cringe, crazy, not useful prompting. What you need to understand is that these models are effectively word maps and the things it chooses to do because the paths in the model drive it there when it's just being told fix this thing and improve this metric are going to be the safe happy paths to do that.
57:24 But as soon as you start including words like ambitious or boil the ocean or push beyond what I requested or make weird and wacky suggestions, you're pulling the model out of the happy path and towards the sketchier, scarier ones that is correctly trained to go down less because if you don't have the right tooling, then it could break So, if you I know you're saying this ironically, SJO, but I actually think this would work.
57:54 I would unironically guess that if somebody had a model that they felt like wasn't going far enough, if they had a a bot or a hook that would automatically append, "Fuck me up, fam," to the end of every prompt they send, they would actually notice a better experience. It sounds so stupid. It really does. And I know you guys are going to cook me in the comments for this.
58:16 I don't care. It works. Speaking of all of these fun things that people don't really want to believe in, myself included often, the next section is about taste. Every thread had a named human owner, and Claude highlighted any user perceptible change with before and after screenshots or recordings for them to rule on. Should a table fill in cell by cell or wait until each row is complete?
58:37 Should a loading skeleton show up immediately or only after half a second? Is a wordbyword fade on streamed text worth the fifth of a frame budget it costs? Claude looked for ways to shave milliseconds and we weighed the trade-offs. I mean, I can answer all of these questions because I know how browsers work. Should a table fill in cell by cell or wait until the rows complete?
58:55 It should wait till rows complete. I'd argue it should wait until the table's complete. Should a loading skeleton show up immediately or only after half a second? You should do a lot of metrics and analysis to know how long loads take and how wide a range these extremes are. Because the worst case would be that you think the data is going to take a second.
59:10 So you show this half a second in and then 501 milliseconds in it has the data. So you show the skeleton for like two frames at most before you show the real data. You got to tune this one based on real world usage. And if you get really fancy with it, you can time other requests and cache some data locally around how long this request usually takes for this user and then make programmatic decisions around the skeleton based on that.
59:34 And then we have the wordby-word fades for stream text worth a fifth of frame budget. No, wordbyword fade in is almost I I would say just straight up isn't ever worth it. I begrudgingly went from only showing things when the whole response was done in T3 code to only showing things when the paragraph or new line is completed or the code block is done.
59:55 And that allows us to have insane performance still without any of these weird potential issues of like frame time being killed by per token stuff. Yeah. Like I'm happy that they were able to work with Claude to find what seemed to be reasonable answers for all of this, but I now I want to see how reasonable these answers are. I'm testing how tables render in Claude.
01:00:14 I want you to render some test tables with at least four columns and 10 to 15 rows. Can use whatever you want for the data. I just want to see how they render. Let's see how they come in. They're doing rowby row it looks like. Uh, nope. That that was pretty brutally streamed to the side there. Maybe they're doing line or they're doing a cell by cell.
01:00:45 Yeah, that's that's very sell by cell. I might need to switch to Fable just to make it slower. Yeah, that's definitely cell by cell. This might end up being a bit of a self-own here if I do this, but uh I want to see how this compares to the performance in T3 code for a similar request. I prefer that. I don't want to see the table while it's still being created.
01:01:17 I just want to see it when it's done. To each their own, but I prefer that. The one thing I don't like, and I'm still deciding what I want to do about it, is the title comes in first. and then we have to wait for the rest of the paragraph. I almost want to hold title renders until the content underneath is done. That's like the only change I would make with our behavior here.
01:01:36 But yeah, you get the idea. These are complex things that you should have conversations with your team about and really make good decisions on. Rowby row is what chat's asking about. It's fine, but I prefer for the table to be done personally. I will also say that rowby row renders for like code blocks horrible. Terrible. It breaks all the highlighting.
01:01:58 It makes way more render work than you should need anyways. And I don't want to read a partial code block. I just want to read it when it's done. I also find that I don't really sit watching threads anymore. I usually kick off the thread and then go somewhere else and then come back when I have a reason to. So yeah, I now see why they labeled the section taste because this really is taste.
01:02:16 Now we hop to direction. We kept each thread deliberately narrow, focused on one benchmark or journey, and asked Claude to find improvements only within that scope. We thought of threads as 150 hammers seeking nails. Most of our calls were about sequencing and user impact, which surface to prioritize, how to combine threads that were stepping on each other, and when to close a thread that had reached diminishing returns.
01:02:37 On one 900 line PR, we got a oneline reply. Quote, going to gavvel that 2 millisecond per send is not worth the complexity of maintaining this build plugin. They were adding build plugins to shave 2 milliseconds off message sends. This is why you have to be careful about how far you let these things go. Like I I wrote my own V8 based runtime for my cloud, but I had to like push for that.
01:03:01 We're nearing the point where the agents will autonomously do like that. So make sure you look out for unnecessarily huge changes for unimportantly small wins. This sounds like an interesting section. An 8 millisecond budget. One of our side quests shows everything working together to demonstrate an optimization to a reax used in the live syntax highlighting.
01:03:20 Claude attached a screen recording of a longer answer streaming in the lab. In the corner, it added a frame rate readout computed in the page from animation frame timestamps. Well, yeah, for those who don't know, 8 milliseconds is how long you get per frame at 120 Hz. A lot of people have 120 Hz displays, whether it's their iPhone, their MacBook, or any gaming monitor, and even some modern TVs.
01:03:40 It's a good target to aim for because somebody with a very expensive pro MacBook is going to want the page to run at the 120 FPS. This is actually kind of a sick bench. Are we capped at 60 fps? Can we try to drive scroll and stream smoothness up to 120? If I understand correctly, your rig may not support this. Right. Today's rig ticks at 60 Hz because headless chromium does by default.
01:04:03 I believe it can be driven at 120. Confirming that first, then I'll rerun the eval at an 8.3 millisecond frame budget. Update on the 120 Hz rig. It works. Deterministic 120 Hz frame stepping in headless Chrome via dev tools. Begin frame control. Exactly 240 frames with 240 begin frames at 833 mms. This is this is not Opus 55. This is Claude slop. This is hardcore Claude slop.
01:04:28 Then Raymond amazing cook. Yeah, I love this. I also am just now realizing one of the things that I really like about quad tag. Notice what isn't in this thread. Notice that we don't have code. Notice that we don't have tool calls and the responses for those tool calls. Notice that it's not listing all the commands it ran. Notice it's not showing all of the thinking traces for the sections of thinking that it did.
01:04:55 I like this because they don't really matter anymore. I'm considering adding a feature to T3 code that just outright hides them or maybe gives you a small ticker like I've run this many tools so far because I don't care. I just want to know when it's done. I also can't help but note the irony that a lot of the performance things they're trying to fix here are not in Slack.
01:05:15 They're in the Claude web app and desktop app, which they are notably not using for any of this. First off, that's funny because Slack's a better experience with Claude than Claude code. But more importantly, it's funny because a lot of these performance issues are things that they show in the desktop app and on cloud.ai that they do not show in Slack.
01:05:35 Just saying. Once the mechanisms and ambition were established, Claude got to work. Each painted frame had a budget of 8.3 milliseconds. So Claude stepped through a long reply frame by frame, timing each one to find the slow parts. It eliminated some work that was O to the message length, so however many things are in the message, which is a lot of work it was doing per chunk.
01:05:56 By memorizing finished blocks, moving tokenization logic from growing code fences to a worker, and revealed tables cell by cell. Oh, that's cool. Uh we were doing perch chunk memoization for T3 chat in February of last year. It's been about 22 months or so. The moved tokenization logic from growing code fences to a worker. This is the early pieces of what I was suggesting earlier which is using workers and wom to run your syntax highlighting and code block stuff so that it's not blocking the main thread.
01:06:30 Also, you can cache the result of that in memoized chunks on the outside of the like wom side and the worker side in the react world. So, you don't have to re-trigger the WASM on every change. You can also reveal the table cell by cell. Yeah. So, I guess I had the answer to my question earlier about how they're doing the table stuff here. Just need to scroll more.
01:06:48 In that one thread, we landed nearly 60 PRs. This is just the thread about the 8.33 millisecond render times. Long replies were blocking the main thread for 200 milliseconds in total where they used to block for over 750. Ran on about a third as much CPU and it held 120 FPS from start to finish on 120 Hz MacBook. The 120 Hz rig itself became a nightly job with Claude watching for regressions.
01:07:12 Awesome. We love that. Oh, I hadn't thought of that. Agore, good point. We could just use Jev to do the syntax highlighting. That would solve all these problems. Oh, that's actually really funny. They cite the tweet that I I actually people saw what my post about this as dunking. It was not meant to be a dunk at all. I was actually trying to start a conversation around whether or not streaming matters.
01:07:32 This is mine my my quote tweet here. I don't want responses streamed anymore. Just give me the whole paragraph or block when it's done. To be very clear, I'm referring to token by token streaming. You should stream down paragraphs, code blocks, tool calls, etc. Just stop sending me incomplete text. I also called out that I feel obligated to mention this performance work is very hard and I'm impressed about the work they did here.
01:07:53 Particularly funny that I have this when in the end it was all just claude running in a loop. Anyways, when they started this performance sprint, they hadn't planned to hill climb on the milliseconds between frames while streaming, but it turned out that they could count them. So, it being countable meant Claude could climb it. So, what's next? Cloud AI and the desktop app are now three times faster than they were in early August.
01:08:14 and the ratchets should keep them there. But they're not done. The 95th percentile other journeys and very long conversations still have room to improve. In a separate post, we'll also write about some of the side quests that took them upstream during the sprint with contributions landing in Electron, Chromium, Node, and more. They'll share the results internally.
01:08:31 Isaac put it best. You cannot have convinced me that this was possible even 6 months ago. Yep, this would have been very hard to convince me 6 months ago. We expect to keep working this way one thread at a time at any scale. The channel's still going. I'll shout out Anthony. I know I've been harsh towards Claude Desktop, but I know he's been working really hard on improving it.
01:08:47 I did actually use it a bunch when I was testing Opus 55 and it has made a lot of improvements. So, shout out to him and the rest of the team making Claude Code Desktop way better. Anony's awesome. And also, I do love that Boris is encouraging them to be more ambitious. That is a very good thing for him to do as the leader. This was a genuinely really fun read.
01:09:04 I enjoyed this a lot. I didn't learn quite as much as I hoped to, but it did really emphasize a lot of the things I've already been trying to do. giving the model the ability to measure things so that it can fix them rather than hoping it will figure that all out when you ask it to fix things. If you tell the model improve performance, it will do its best based on its theories to improve the performance.
01:09:25 But if you start by telling it to find ways to measure performance, validate those findings and then tell it to go improve the performance, it'll improve a lot. This is really cool and I think we can all learn a lot around how they are prompting, how they're thinking about these changes and how they are validating the work that they're doing because that is going to be more and more important as things continue to change.
01:09:44 If you still think that reading the code is the solution to all these problems, then you're going to have a whole separate set of problems and a whole lot less solutions. Figuring out how to work with and around the slop really kind of is becoming the meta. I know a lot of engineers are going to hate that. I personally find it kind of fun. It's a new type of engineering that requires new types of cleverness in design to set up systems and structures that allow your agents to operate effectively and efficiently.
01:10:09 I'm thankful I'm finding more opportunities to deep dive in these fun engineering topics again. I really am enjoying it. Thank you all for watching this video as well as all the kind words on the recent React Native one. It's such a relief to be able to dive into the engineering topics again where it makes sense. Let me know how you feel about all of this and what you learned in the comments. And until next time, peace nerds.