Searchable transcript of Getting the most out of Opus 5.5 — Theo - t3․gg (28:42). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Theo - t3․gg. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 Opus 55 has been out for a bit now, and the more I use it, the more I think this model might actually be really, really good. I've been very impressed with the code it puts out, with how nice it is to interact with, with how it stays on task for long work, as long as you don't use Macs. I We'll talk about that in a bit, don't worry. Today, I want to talk about how you can get the most out of this model.
00:17 I've had a couple videos like this and they vary in performance, but I think it's some of the more important work we can do, especially when awesome people like Addiosmani, who used to be part of the Chrome team, recently joined Anthropic and is using his depth of knowledge and education capabilities to write something awesome like this. If all goes well here, not only are you going to learn about how you can use Opus 55 more effectively, hopefully I will as well.
00:41 I trust Addie with my damn life if I'm being real. This guy is super legit and when I saw him joining Anthropic, I got really hyped and I am so excited to see what he has to teach us about maximizing the value we get out of the new Opus release. On the topic of maximizing value, we should take a quick break for today's sponsor. If a chef doesn't have access to ingredients, they probably can't cook a very good meal.
01:02 So, why do we think our agents are going to make good results if they don't have access to 80% of the data they need that is currently locked away on the internet? It turns out agents need a way to get that data. They need a way to browse the web. And that's why Browserbase built the browser for your agents. Not just for going to sites and clicking through them, but for getting all the context they need from the internet.
01:21 Whether it's through their search API to get results on whatever arbitrary queries they need to resolve or their fetch API where they need to get some data out of a page that is hard for them to access. How many times have you seen an agent write a curl request, fail, give up, and then go find something else that was incorrect? If you've ever seen Codex or Claude do this, you know how miserable it can be.
01:40 You've probably already burned hundreds of dollars in tokens hitting these errors silently and not even noticed. And if your agents do need real browsing capabilities to do actions on your behalf, whether it's to sign into a page to book something or go find some data that's behind a payw wall or a signed-in state that you need, Browserbase can handle all of that for your agents, too.
01:57 Browserbase's goal is simple. Since agents are good at calling APIs, turn the whole web into one. This means that they handle all the details from scaling up their servers to make sure you're always able to access what you need to catching broken flows in your apps with real agents clicking things and noticing issues to getting around captas, handling off, and so so much more.
02:12 Ready for a crazy lore drop? Browserbase is so good that it's what Google officially recommends for their computer use work. This is a real Google repo demoing computer use, and they recommend browserbase. Figure out why everyone from Theo to Google loves them at soy.link/browserbase. Let's talk about getting the most out of Opus 5.5. Opus 55 works well in the way that you already use Claude.
02:33 A few things do behave differently, though. It works for longer on its own. It tells you plainly what it did, and it thinks before every reply. This guide covers how to work with Opus 5.5 in Cloud Apps and Cloud Code, including how to prompt the model, steer long runs, and check your results. He starts with a really fun exercise. He recommends trying three specific things in your first sessions with Opus 55.
02:55 The first example he has here is a thing I talk a lot about. The idea of handing over the whole task. I call this prompting wider having the model come in earlier and go further. And an important part of doing this is defining good done states. What is complete? Rather than start working on this feature, say I want you to build it so it has these three things.
03:16 Verify it by showing me a screenshot and then file a PR after. That gives the model a clear end state. It's not done until it has made the feature. There is a screenshot of the feature and there is a PR up including that feature and the screenshot. Generally speaking, these higher tier models from anthropic and open AAI work really, really well if you tell them what done is because they'll keep going until they get there.
03:41 The next two things are more vague, so I'm curious how he frames these as we go on. The second one was delete think carefully lines. You don't need to tell the model to think. it already is going to. I don't know if that really fits here exactly, but interesting inclusion. The last piece I don't like the wording for, but it is important. When long runs end, read what it needs from you first.
04:05 This is an important thing, and I actually think the new model does this quite well, where after it does a run, it will tell you exactly what it needs. It does sometimes miss this though, so I'll often find myself prompting accordingly. Here's a real thread I kicked off earlier today in T3 Code that I think showcases what I'm talking about relatively well.
04:19 This is a thread I had started on my machine testing out the new Opus model with Max Reasoning. Six and a half hours later, I realized I wasn't going anywhere. I I really don't think you should use Max on Opus 55. It just This thread ran for way too long. So, I had it on Max Reasoning, and I was running it on my own computer, which was even more annoying.
04:37 As y'all who have been around for a bit know, I have moved pretty much all my agent work off my laptop and onto other machines, mostly my apartment, but a few remote servers as well. So, normally when I use T3 code, I trigger it somewhere else. But in this particular case, I only had the codebase and everything set up on this particular machine and I wanted to see it working too.
04:54 So I ran it on here, forgot about it, went and did other things. Came back and it had been running for six and a half hours. I switched to high reasoning, asked what was going on. I only asked it what's the state of the work. I it didn't continue the work after there. I told it to continue and it did. 10 minutes later, it had finished and was in a pretty good spot.
05:12 I asked how long it thinks these changes will take. That was for another demo and video. But here is the prompt I actually care about. Remember this thread was running on this laptop I'm on right now and I wanted it to run somewhere else. I could have put the manual work in here. I could have SSH to the other computer, cloned to the repo, added it to T3 code, opened it there, SSH in again to go dump all the environment variables, make sure it has browser use and everything it needs set up, and then tell it to go take
05:37 this markdown file that I would copy paste over. Instead, I just told the model what I wanted. This is the prompt that I think does a good job of giving the agent an out when it needs it. I want you to kick off an agent run on my computer named Leftbook. You should be able to see how to connect through the fleet repo on this machine. This was me letting it know how to connect so it doesn't try to do anything sketchy to connect and also so it doesn't do something dumb that might cause it to hit a like security guard or
06:03 flag. I want you to make sure one that the repo is cloned, two all the needed environment variables are there, and three that you can run opus 55 through cloud code and T3 code such that I can still see the thread how I normally do in T3 code. There's a couple pieces here that are important. I want you to make sure is me telling the model, not only do I want you to do these things, I want you to validate that you can and inform me if it fails.
06:26 I also am explicitly giving it permission for very specific important things like its ability to copy my environment variables. If I didn't tell it it could do that, it might be concerned when it goes to do that and then stop and ask me. I don't want it to stop. I know it's going to have to do that for this work. So, I just told it it can so I don't have to give it permission.
06:46 Then we get to the other parts I had here specifically what my goal was. My goal here is to do the whole implementation on that machine in one pass with however many sub aents and whatever else are needed with full computer use capability on the other machine. This is an explicit statement of what I want. Being explicit about what you want is great, but being explicit about what you don't want is even better.
07:11 I don't want you to be spamming my machine with computer use requests and browser control while I'm using it. This is the most annoying thing in the world that actually made me not use browser use and computer use as much. I got so annoyed with codecs just popping windows up on my screen constantly. When I started using this remotely on other boxes, I started liking computer use a lot more when I didn't have to watch it pulling windows up when I'm trying to do other [ __ ] on my computer because I fire and forget.
07:35 I throw up a thread, I let it do its thing, and then I go respond to some messages in Slack or I go play a game through my fork of moonlight. I do other things while my threads run. I don't want them interrupting the other things while they're running. So this specific request to not spam my machine makes it very explicitly clear that even if it thinks remote controlling my machine to access the other machine is a good idea that it knows not to.
08:00 Then we get to what might be the most important piece. If anything does not work as expected when you move this workload over to Leftbook, don't hesitate to ask questions so we can get it working right. It sounds silly, but these models have been RL to [ __ ] hell and back to not give up on tasks. So if the model asks a question, the initial thing has been RL to do is to try and answer the question itself.
08:23 So if it did have questions or issues with my request here, it by default might not do the right thing. It might ask itself, should we try another way or is this still right and then go blast through it anyways. This particular last sentence gives it an out. It makes the model more willing to stop and ask the question if it does run into a problem or it does get confused.
08:48 Thankfully, it didn't run into any problems here. Everything went relatively well. It gave me an overview of all the things that it did in this process. Called out that I have some connection issues with Leftbook remotely right now, which I do actually need to figure out. I have some PRs up that should fix that. But the point here is that I gave the model exactly what I wanted.
09:07 I upfront told it about the things that I was scared it would ask about because I didn't want it to stop and get permission. I just wanted to do the thing. And I also gave it an out where if it didn't have confidence in a path, it could ask me so I could hand it the solution. And the result was that after about 17 minutes, it was able to get all of this set up.
09:23 And most importantly, I now had the thread from my other machine. Here it is. This thread from my other machine appeared and is behaving. And apparently it only did phase zero. Great. Let me see the prompt it ran here. Goal: Build the whole ping rewrite described in overhaul MD in one pass on this branch. Read the whole plan before you start. Follow its phase 0 to 7 in order.
09:49 And it did phase zero and then stopped. I want you to hold off on all questions until the very end when you get through to phase 7 in completion. I don't want you to stop after every phase to ask for permission to continue. I want you to do the whole thing. Continue working on this port until the entirety of it is completed and don't inform me until you have a tail scale link I can click on.
10:15 That is the complete rewrite so I can test it and make sure everything works. Don't worry about how it's hosted. Don't worry about how it's built. Just worry about completing the work as assigned. I am admittedly a little disappointed that that first prompt wasn't enough to get it to behave properly, but now it will be. Now I do not have to look at this thread again for at least an hour or two.
10:35 Yay. So with that all running, let's dive back into this blog post. The first section is how should you ask? It starts with this important piece that funny enough opus doesn't understand. That's why it just had the issues there. You have to say what done looks like then let it run. Name the finish. Give the whole task in one message. Something like the tests pass or every endpoint is migrated.
10:57 It needs specific instructions. The reason that we have to do this is because despite oat was 55 being much better at going on long multiart work it might still stop to ask for permission with a clear finish line it will know when it's done but again as you saw with that example there without that it might just randomly stop gives an example here of what this would look like migrate the payments endpoint from the old client to the new one done means every endpoint uses the new client the old one is deleted and the test
11:27 suite passes stop and ask me only if a test fails for a reason that you can't explain. This is the perfect simple one two three. What you want, what means done, and the out if it needs an out if it does get confused or does something wrong. Next point, stop telling it to think hard. I am amazed people still do this. I've seen so many people say like, "Think deeply about this."
11:50 The models are pretty good at knowing how much to think. In fact, a a video I kind of want to do, but I don't know how to like phrase the ideas and like package it well enough is a adittedly a crash out on max reasoning. Because the cool thing about these models now is that the reasoning levels aren't necessarily think this much. The reasoning levels are think up to this much.
12:13 Low and medium are effectively saying think as much as you need until you hit this point and then stop. high and X high are much higher caps for how much thinking the model can do before it stops and gives up because the thing was too hard. Max isn't just making it so the model can think more. It is removing its ability to think less. And that's what scares me about max reasoning and why I specifically think you shouldn't use it.
12:41 Funny enough, I was talking about this with Addie earlier and he's trying to figure out how they can improve Max reasoning levels in Opus 55 in Cloud Code. But the way I would recommend thinking about Max right now isn't it can think more. It's forcing it to not think less. You can see this very easily here with my runs with 55 using X high and max in Skatebench, which is admittedly kind of silly benchmark.
13:01 On X high, the average tokens per response was 338. The average duration for a response was 6 seconds and the slowest was 31 seconds. On max, the tokens went from 338 average to 5,000 average, more than a 10x. Much scarier, it bumped the average duration to 50 seconds. So, the average max run took longer than the slowest X high run. And the slowest run was 20 times slower at 600 seconds.
13:33 You know what the best part is? All it got out of that was one additional correct answer. It went from 78% to 79%. It cost 13 times more. It used 15 times the tokens. It took 20 times longer in the worst cases. And I got jack [ __ ] [ __ ] out of it. Are there things where Max could help? Like maybe it is falling for a simple trick answer, but if it has to reason more, it doesn't fall for the trap.
14:01 Perhaps. I still don't think you should ever use max reasoning. All the other reasoning levels just change the ceiling. And X high can still be fast. It can still not do much reasoning. It's just a matter of how much can it do, not how much will it do. Max is a will. You are setting it and forcing it to think more. I just leave it on higher X high and don't think about it right now personally.
14:22 But yeah, low kind of sucks. I'll be real there. This model does need to think a bit. That's why they don't have a no reasoning version. So yeah, I'm sure a lot of y'all like the the comment I got the most in my previous video was, "What reasoning level should I use? High or XI? They're both fine. I don't have any issue with either. I can barely notice the difference with either.
14:38 Xi feels like it takes a little longer. High sometimes misses like smaller details that Xi doesn't. Not that big of a gap. I just leave it on higher XI. And don't tell the model to [ __ ] think. It knows that it should think. It is smarter than you probably think. Hell, it's smarter than you probably are. Smarter than I am. Next, we have add to a running task.
14:54 If you remember something midrun, you can type a follow-up while it works. Why does this matter? It matters because the runs are longer, so going back to the start is more expensive. I have done this before. I've had times where I noticed the model was just going the wrong way when I was reading its outputs and just decided [ __ ] it, stop, new thread, copy paste prompt, make a few adjustments, add two more things to the bottom.
15:16 Like, by the way, do these things as well. I don't do that anymore. The modern models, specifically Opus, Fable, Soul kind of, and Astro mostly are way better at what we call steering. That's when you hit send when it's already working and it steers the model in the direction of what you sent while it's running. steering used to have a pretty rough problem where the RL was on a per message basis.
15:40 So if you said I want you to do tasks one, two, and four, and it started working, and you're like, wait, I forgot to mention three. It would then forget about one, two, and four, immediately do three, and then be like, okay, I finished three. And then not do the other tasks from the previous message. Because through its training, it effectively was taught to complete the message and ignore the history beyond how it helped the context.
16:03 They have now been trained better about this to take these interruption messages not as a reason to stop previous work or treat it as done but as what they are steering additional things to help the model go the right way. Here's another one that's really important around design. This one actually bit me a bit when I was using the model. When I took a look at Witchai which I still think is one of like the cooler projects I can use to showcase model capabilities.
16:27 This was built by DRA to show the front-end capabilities of different models. I noticed that Fable 51 had much better designs overall than Opus 55 did. This is like this is the Opus 55 version of this design and this is the Fable 51 version. I hope we can all agree that the Fable version is obviously better and nicer. But I saw other people saying Opus was a way better designer and I was confused because when I took this quick look here didn't seem to be the case.
16:57 Turns out the thing Opus 55 is good at isn't just making a nice design when you say, "Hey, make it pretty." It's much better at going the right direction with design instructions. If you tell it what you want, and more importantly, also tell it what you don't want, it can make really good designs. I like how Addy framed it here. This is another example of like anthropic loosening their death grip on comms, letting the employees say things that the models are bad at.
17:21 The reason this matters Opus 55 is that with no design direction, Opus 55 will fall back on a few default styles. A general instruction like avoid generic looks will mostly just swap one default for another. A list of specific patterns will work much better. For example, build a personal website with placeholder content. Don't use a cream or off-white background, italic accent words and headings, numbered one, two, three section labels, monospace labels, or pill-shaped buttons.
17:48 I don't think you should have to specify that many negatives. But you get the idea. When you tell it to not do a thing, it won't do the thing. And once that's done, you can look at it and decide if you like it or not. If you don't like something it added, put that on the list, and then ask again. Or something I do a lot, I use a tool like Shotter where I take a screenshot, I draw an arrow, I point it, and then I just paste this into the chat prompt.
18:10 Like, I don't know, here, paste, and I'll say something like, "This sucks. Make it better." That works way better than you would think. Next, we have more on steering. Steering long runs in Claude Code. The first point, tell it which stops you want. You can put a short rule in your Claude MD file about when to stop and ask and when it should keep going.
18:28 The reason this is important is because Opus 55 will keep you posted as it works. On long tasks, it'll sometimes stop to report instead of going on, a summary that names the next step without taking it, an offer to continue, or a list of choices that don't actually block the work. It follows instructions that name the stops. So, name the stops that you want as well.
18:47 When a step doesn't need my input, keep going. Put status notes in the same message as your next action. Stop and ask only when you can't continue without me or before anything destructive, deleting data, force pushing, or changing anything outside of this repo. You know what I'm going to do? I'm putting my money where my mouth is. I'm pasting this directly into my cloud MD.
19:08 I already have a section around approvals here, and I'm putting at the end of this. When a step doesn't need my input, keep going. put status of same message yada yada it's the exact thing we just read. Now this will be in my cloudmd but I don't have a script that auto pushes this. I am much sillier than that nowadays. So I hop here. I switch over to opus 55 and I say I just updated the cloud MD in this repo.
19:31 Commit it and push it out to the rest of the fleet. One of the rare instances of typeless screwing up my pros but that is fine. Cool. Oh, it actually does have a dictionary. That's nice. Hopefully this will be a meaningful improvement. I like that. We'll see how it helps. If a run stops with, "Want me to continue?" Reply continue. If it happens often, the rule above will help.
19:50 Yep, that's why I added it. A rule to keep going means fewer stops. So, keep your own check before anything risky or hard to undo. The last line of the rule above does that. He also calls out the example of prepared programming where you might want the opposite. Maybe a oneline plan before it starts and a short recap at the end because you want to be more involved with it or maybe you have another person working on this with you too can be very useful.
20:11 I don't want that. I want to kick off the thread and come back to it when it's done if even. So I very much prefer the telling it to just keep going here. The next point ask it to split big work across sub aents. The point here is that for audits, migrations, reviews across large codebases, stuff like that, splitting the work up into sub aents can make it much easier for the work to get done within the context window and for those different sub aents to go explore different things.
20:35 The reason this matters is because early testers had Opus 55 coordinate parallel sub aents on long audits with little oversight and it was successful. The other thing he isn't saying here is that I've noticed Opus 55 seems less willing to spin up sub aents unless you tell it that it can. So I find myself often telling at the end of prompts, use sub agents and workflows however you choose.
20:55 I trust your judgment there. Even then it doesn't always do it. So if you like know the task is better with sub aents like a giant pile of PR reviews or an audit of your codebase just tell it to you sub agents and it does. The next section is around checking the results. What should you do when long runs end? Look first for anything Claude is waiting on you for like a decision it left open or a change it wants you to approve.
21:17 Then read the rest of the Claude summary. The reason this matters is because Opus 55 reports on its work more clearly than Opus 5 did. So it's actually worth reading the outputs. The updates and their final summaries say what they did, what they found, and what it needs from you in plain language. This is interesting. He actually suggests putting this in the cloud MD.
21:33 End every run with three headings. Blocked on me, changed, and found. I'm not going to go that far. When I do want this, I usually just ask. I often will find myself going back to a really old thread and not remembering what it's for or what the status is. I'll just ask it. What's the status of this work? What do you need from me? Another one I ask a lot and I I'm amazed at how helpful this has been.
21:57 What are the risks of merging this code today? That's probably my most common prompt. I'll often have a PR up from an agent that ran for hours, made a bunch of changes, simplified it, and it gives me like a 300 line of code thing that I don't fully understand or really care to read. And I don't know if I should bother reading it or not until I know how risky it is.
22:16 So, I'll just ask, how much risk do you perceive with this change? What's the worst thing that would happen if we merge this right now? And I have found myself much much better understanding the state of my work when I ask questions like that. Another fun thing you can have the model do is review code. It can review its own code. It can review code from other agents and other models.
22:35 I have personally found that swapping between model families and labs is actually pretty useful. I think that like generally speaking, I find OpenAI models to be better at review still. they just really dig into the details and won't let go of the thing until it's confident it's found everything. It'll report more stuff that doesn't matter, but it will occasionally find things that do matter that the claude models just miss entirely.
22:59 So, personally, when I have Opus code being reviewed, I don't want just Opus reviewing it. I would like for Aster or even like Soul to give it a review as well. That said, 55 does find a lot more than previous Opus models did. I mentioned this in both my Gro and Opus videos. I put together a a bench in quotes that measures how well different models are able to find areas of improvement in the T3 code codebase.
23:26 And I was very surprised to see that Grock 47 came in second place and Fable was quite a bit lower. The reason why is pretty clear when you look at the number of supported unresolved contradicted findings. Astra and Grock 47 both found eight things that could be improved and substantiated and supported its claims. GB6 Soul actually found nine things, but apparently they weren't quite as valid because the judge panel I had set up found it less quality overall.
23:52 Babel found only five things in its two runs. That's a big difference. Babel absolutely vetted the things and they all matter and are worth improving, but it just didn't find as much as the OpenAI models and hell even Grock did. But also, you can see the gap with Opus 5 to 5.5 here. Opus 55 was almost twice as successful according to my judging system.
24:14 and also didn't find any unresolved or contradicted things. Where with Opus 5, it had four supported findings and then two that weren't. So for every two findings that are valid, it has one that's [ __ ] The point I'm trying to make is hopping between model families can be useful here. And as much as I prefer coding with Fable and Opus, the review quality I get out of Astra is also really solid and maybe even Grock, by the way.
24:36 So consider using Grock for your reviews, too. One more important piece with the reviews, though, and I do this a lot. You should ask the model to mark the things that it can't confirm so it can review and verify as much as possible. Ideally, you give it the tooling it needs to open up a browser and check your changes or run the system and use computer use to verify it or a test suite that it can use or build to verify the things that it's concerned about.
24:59 So, the model knows the code works. But if there are things it cannot confirm for any of many reasons, whether it's like a tool that you need to use different hardware for or using an environment variable or it's a more subjective thing, if you ask the model to let you know what things it can't verify, it will. And this is kind of the point we're at now.
25:17 The models know your codebase really well. They know how to operate really well autonomously. They know how to get stuff done. They can even verify their changes. They don't know how to work with a human yet. That's partially because the data doesn't exist. It's partially because this is still a new phenomena. But if you tell the model how to work with you and you tell it what you need and what you're expecting, it'll usually give it to you.
25:35 Especially models like Opus 55, Fable 5.1, and GPT6 Astra mostly. Okay, this is a little bit of a silly call out. This is actually about how to use the cloud apps more effectively. And the first piece is to check the model picker says Opus 5.5. This is making me feel like I shouldn't be reading this article. So, let's skip down to what to do when messages are flagged.
25:55 Obus 55 is the first OBUS model to launch with Fable level bio and cyber safeguards in cloud apps as well as cloud code. Most flagged messages move to older models and your work goes on there. Binding security vulnerabilities in source code is allowed in everyday health and educational questions should still work. These safeguards can sometimes flag legitimate work and we're tuning them to cut down on incorrect flags.
26:17 If you're switched, here's what you'll see and what to do in the cloud apps. You'll get a notice that says you were switched to a different model. Quad answers on that model and the chat stays on it too. If you go back and choose 55 again in the model picker, it might work. It might flag again. Starting a new chat should avoid it. You they also have settings for quad where you can turn off the auto switch and instead you'll get a pause or an error.
26:41 I'll be real, I have not hit as many flags recently. It doesn't seem like that big a deal, but to each their own. One thing that is called out here is that you shouldn't ask the model to show its reasoning. This is a little annoying because sometimes I want the model to explain why it did a thing and my like natural English way of doing that is what was your reasoning for these changes that might flag as you asking it to share the reasoning traces which anthropic doesn't want to do because that can be used to distill
27:08 their models. They don't want that so they hide them. It is what it is. As such, anything that matches reasoning is risking getting a flag. I guess there's a cute little checklist at the end here. Make sure you tell the model what done looks like. Don't tell it to think hard, it will anyways. Design requests should have the styles to leave out as well as what you want it to look like.
27:26 Charts and screenshots should be attached. Yeah, I I think people underrate how useful screenshotting is. Half the time I would need to give a model like context on an error. I don't copy paste the error. I just screenshot the browser and paste that. They're good at reading these things. And Claude Code even has tools where if the screenshot is too high res and it can't read the text, it will crop to the area it needs to get the context it needs.
27:48 The other sections we've spent this whole video going through, but if you do want to read this and check it yourself, the link is in the description as always. This is a pretty fun article and I'm thankful to see once again that Anthropic and I are pretty aligned on the best way to use these models. We're now at the point where we need to give our agents more leash.
28:04 They need to have the ability to verify their changes. They need to be told when done is done and given enough trust to go do the thing. I talk about a lot of these layers in a video coming out soon all about how Anthropic made the performance for the cloud site and desktop app three times faster because the systems you build to verify these changes are just as important as the raw source code itself.
28:22 Hopefully this video will help you maximize your usage of Opus 5.5. It really is a great model and the more you trust it, the more you give it the things it needs to trust its own work, the further it can go and the more you can build. I hope this was helpful and until next time, peace nerds. Also, don't do not touch max mode.