AI-Native Workshop is a repository-based workshop for practicing AI-native software development workflows. Its curriculum covers voice coding, agent sessions, loops and measurable goals, Git worktrees, verification gates and hooks, adversarial review, and scheduled tasks. The repository includes setup skills, exercises, slides, a glossary, a facilitator guide, and an optional local `aie-coach` MCP server with seven tools for check-ins and setup scoring. Its setup workflow runs inside Claude Code and can install and verify Bun, the Codex CLI, and Git; the workshop is organized in four blocks covering voice, loops and goals, gates, and schedules.
Claude is an AI assistant developed by Anthropic, positioned as 'The AI for Problem Solvers'. It is a general-purpose AI system used for tasks such as generating code prompts, refining requirements, creating advertising strategy and copy, and processing creative content like storyboarding and video prompts.
Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.
A development tool that creates a Git worktree and starts a Claude instance inside it, allowing development tasks to be worked on in parallel.
Cloudflare is an internet infrastructure and security platform that provides hosting and delivery infrastructure for applications, agents, websites, and data. Its network combines compute, connectivity, and security at edge locations, with Workers running code close to users and backend data; the company states that its network operates in more than 335 cities and reaches 95% of the world's Internet-connected population within 50 milliseconds. Cloudflare also provides AI crawler controls that let site owners inspect, allow, or block crawler activity, along with pay-per-crawl and a monetization gateway for pages, data sets, APIs, MCP tool calls, files, and search indexes. The described x402 flow uses an HTTP 402 response to request payment, after which an agent pays, retries with proof, and Cloudflare verifies the payment at the edge.
An AI-powered code-review platform that automatically analyzes pull requests in repositories hosted on platforms such as GitHub and GitLab, producing review comments, summaries, and suggestions within source-control and code-hosting workflows.
Fleet is a terminal dashboard for managing multiple AI-agent sessions in tmux. It watches Claude Code sessions through a Claude Code plugin and event hooks, combines hook data, JSONL events, and tmux pane scraping into a seven-state model, and displays sessions grouped by project or tmux session with attention indicators, summaries, notifications, filtering, live previews, and prompt or passthrough controls. It also supports Codex and pi integrations and can discover hookless agents such as aider, Cursor, opencode, Gemini, Amp, and Droid from the process table. Fleet provides a Bun-based CLI for querying and streaming agent state, waiting for state changes, capturing pane output, sending prompts, switching or acknowledging sessions, and integrating status information into tmux. It is distributed as a standalone binary or through Homebrew, with optional tmux sidebar, popup, status-line, and desktop-notification integrations. The repository states that it is licensed under the MIT License.
Handy is a free, open-source, cross-platform desktop speech-to-text application that transcribes speech locally and inserts the resulting text into the active text field. A configurable keyboard shortcut starts recording in hold-to-talk or toggle mode; Handy filters silence with Silero voice activity detection, transcribes with local Whisper or Parakeet models, and pastes the result into the application in use without sending audio to the cloud. It runs on Windows, macOS, and Linux, supports Whisper models with GPU acceleration when available and the CPU-optimized Parakeet V3 model with automatic language detection, and provides CLI controls for recording and startup behavior. Linux text input may require tools such as xdotool, wtype, or dotool, and Wayland support has stated limitations. The project is distributed under the MIT License, although its name, logo, icon, and brand assets are not open-source.
Linear is a project management and issue-tracking platform developed by Linear, Inc., designed for planning and building software products. It provides issue tracking, roadmaps, workflows and integrations with developer tools, and includes AI-assisted features to support planning and task management.
OpenAI Codex is an AI coding agent from OpenAI available as a command-line tool (Codex CLI) that helps developers produce software. It can be used alongside Gemini for adversarial audits of software requirements and implementation plans, listed as a supported coding-agent or model option in several projects, and its logs can be joined with task and test evidence. The Codex CLI can also receive and answer requests from the Penako canvas.
tmux is an open-source terminal multiplexer that creates, accesses, and controls multiple terminals from a single screen. A session can be detached while its terminals continue running in the background, then reattached later, allowing terminal-based processes such as coding agents to survive disconnection. It runs on OpenBSD, FreeBSD, NetBSD, Linux, macOS, and Solaris, and is built from source with dependencies including libevent and ncurses.
Vercel is an application and agentic infrastructure platform for deploying and hosting web applications, marketing sites, platform products, and AI agents. Its infrastructure includes global delivery, deployment environments, serverless functions, fluid compute, a web application firewall, durable orchestration, sandboxed environments, and an AI model gateway. For hosted platforms, Vercel provides tenant isolation, domain management, custom SSL certificates, and preview URLs.
Wispr Flow is an AI voice-input and dictation tool that uses speech recognition and contextual editing to convert spoken commands into formatted text for writing, messaging, and interaction with desktop applications.
Searchable transcript of Lifestyles of the AI-Native — Nick Nisi & Zack Proser, WorkOS — AI Engineer (01:01:28). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:12 Hello everyone. Thanks so much for being here. We are super excited to be with you here today. We've got a great program for you. My name is Zach Proer and this is my colleague Nick Mi. We are both AI engineers at work OS and uh today we're going to share with you some of our our best workflows and things that we've learned over the course of working with Claude and other agents every single day.
00:31 So it's going to include voice coding, agent skills, hooks, and scheduled tasks. And hopefully at the end we'll have some time for QA as well. >> Yep. Uh so the main thing that we want you to take away from this is that uh most engineers just babysit a single session. That is so 2025. uh you should be operating a fleet of agents and we'll talk about the different ways that we do that in real life.
00:53 This isn't uh gimmicky or aspirational things. These are things that we do every day to uh survive at this point. >> Yes. >> Uh so as I said uh so I'm Zach. I'm on the applied AI team at work OS. I would say that my role is uh 60% implementation of both internal tooling and customerf facing applications using AI all day and then about 40% of AI education leveling up the rest of the org.
01:18 Uh, and I'm Nick Nissi. Uh, I'm an engineering manager on the developer experience team. And we are a small but mighty team of two that manages over 25 repos across eight languages. And we're only able to do that because of AI. Uh, it's made it the only way that I can not go insane. [laughter] >> Cool. So, uh, what we have for you today is a a very much an interactive workshop.
01:40 So, if you want to hang out and chill and just watch us present and talk, that's fine. If you want to get hands-on, uh we have a repo for you to use and um you can step through everything. It's it's got skills going to help you get set up with voice coding if you've never tried that before. Uh the way we'd like you to think of this is this is an excellent opportunity and kind of a safe space if you've wanted to try out some new workflows before, voice coding tasks, um scheduled tasks, hooks, etc.
02:04 Uh this is a great place to actually get hands-on with it. And then we'll be available to either help you or answer questions as well. >> Yeah, for sure. Uh so we're going to go over these four things. um in to varying degrees of um depth and we will also have time for um questions and and taking those as well. Um but there's three things that we're going to see uh three main things that we're going to see.
02:27 The slides themselves. There's also this QR code right here that you can scan. This goes to a live glossery and it has a whole bunch of terms that we're going to be talking about. Things like MCP and skills and what else is in there? Loops. Yeah. All of those are defined there. But also there's a little rag chat in there. So if you have questions about that, you can ask there as well to help kind of keep things moving.
02:48 >> Uh but feel free to ask us as well. >> And then we are we are also going to ask you uh for some feedback and some ways to collaborate and take the skills that we're learning today and show what you've learned throughout the course of this next hour. And that's on this live board that we'll see. And uh if you get into the repo, you'll see that and uh be able to contribute to that.
03:10 Uh so Nick made this awesome MCP server that's uh part of the repo which we'll need to share the URL to uh shortly but this will interview you. It's going to it's optional. You don't have to do this. >> It will scan uh your machine if you're comfortable with that. It will take a look at some of the conversations you've already been having uh with agents and get a sense of what you've been working on, how you've been working and give you some just simple rubric scores about uh how AI native you are.
03:33 Uh mostly for fun because we'll do it again at the end and see if there's been any motion. >> Yep. Uh, and it is built set up like a coach. So, it's there to a uh for you to ask questions to and it can try and help uh guide you. Uh, and it tries to answer in ways that we would try and answer those questions. Uh, but also you can ask us as well because AI is not always perfect.
03:55 >> And so this is the the data visualization. So again, opt-in only if you're interested and if you're comfortable. Uh when you are interviewed by that um MCP server or that coach tool, you will uh basically send a simple post to a back end that we've set up in Cloudflare and it will just take a stock of where the room is now and then later on we'll take a look at the end of the session >> and the only data that it sends is things questions that you answer.
04:16 It doesn't uh scan your system or anything like that. >> Yeah. Uh so this is the to get started with the interactive portion though and uh just to again mention that if you want to follow along a great way to do that is watch us you know up here and talk uh listen to us talk up here and and step through it all. will be driving as well on on the screen.
04:35 But then you can also clone this repo and it is designed for you to run claude within it. And that's where you'll get the MCP server and you'll get the skills and the tools and where you can work through it. And then this is an excellent um kind of artifact for you to take with you and then continue practicing on uh after our session. >> Yep. And I think there's a way to leave this screen up here if you want to clone that repo.
04:56 Uh but we could take a question or two while we uh leave that up here if anyone has a question. No questions. Okay, we don't bite. >> Cool. >> Okay, so if you are going the repo route, uh the first thing you're want going to want to do is run claude. However you normally do that, you know, typically you just type claude. And it's going to ask you if you want to trust this repo.
05:25 So trust that repo. And then uh one of the first skills that's already loaded will be this setup uh skill that'll help you install all the stuff for the workshop including uh your first open source and free uh kind of voice uh coding model. So you can say set me up for the workshop or type that in and then Claude will run through the steps to do that.
05:43 And then once you're done with that you can say uh run my workshop check-in as well. >> Y and so this will just give it kind of an idea of where you're at and what uh what's going on. It also the repo is also set up with a number of skills already uh set up and and like plugins for claude. So uh there's an MCP in there. There's a couple of skills uh made by us and not made by us uh like ones from codeex for example.
06:06 Um and it's just an easy way to get started uh and get the repo and and the cloud environment set up in the way that uh is opportun or is best for this workshop. >> Cool. And while we're waiting for a couple more people to get the repo set up, just curious, like show of hands, how many people are already using like a voice coding tool every single day for stuff.
06:28 Okay, awesome. Like a pretty good >> What's the voice coding tool? >> How many for whisper flow? >> Yeah. >> Oh, sorry. Which one? >> Whisper flow, but it got slower. >> Yeah, I had the same issue. >> What got What got faster or what's faster? >> Cool. >> Okay, >> this the the one that we're going to install today is uh an interesting new tool. It's called Handy.
06:48 and uh it'll install down onto your Mac. And the fun thing is that you actually get to select which open source or openweight model you're going to use for transcription and you can balance between uh language support speed. Um but the cool thing about it is there's no subscription because it's running on your machine. Nothing's being sent to the cloud.
07:03 It's fully private, right? So there's a bunch of advantages and we'll talk a little bit later about the the differences in choosing a voice cloning tool >> and it's free. Uh that was a big driver for that one specifically. But if you have a tool already installed that you're using and you prefer, >> please use that. >> Yeah, go for it. Okay, so voice coding, uh, we are huge converts of voice coding.
07:21 I'd say that, uh, in addition to just kind of learning about the LLM ecosystem and tooling and harnesses, this is probably one of the highest leverage moves that both of us have made in past years. Uh, I consider myself a decent typist. I learned to type pretty quickly playing EverQuest as a as a teenager. And uh, I I hit like 90 words per minute if I'm fully caffeinated and I'm hitting like 190 with um, with uh, tools like Whisper Flow, right?
07:45 So, it's significantly faster and we'll talk about how it's not just significantly faster for a single task, but when you multiply that across multiple tabs uh and agents running and checking back in and on work and just guiding them, you you get immense leverage very quickly. >> And not me. I Mavis beacon failed me. I cannot type. Uh but I make up for it in Vim skills.
08:04 >> Yeah, he's being modest. Uh so it could you may if you haven't tried this before uh definitely kind of you know we encourage you to get out of your comfort zone. Give it a shot. Uh I will also share that it's not just about uh coding. You know I use it to respond to you know colleagues and customers and also still ensure that we've got that quality and that everything is formatted properly.
08:23 But uh you know sometimes you might even just be talking through something architecturally and that'll be the largest lift that you get spending an hour in ideation. Nick has a really popular skill that's actually in that repo called ideation and getting to clarity faster and focusing on the outcomes are some of the things that uh voice coding can help you do.
08:40 >> Yep. I want to say a special thanks to Swix and the whole AI engineer crew for this wonderful event in San Francisco. Just an example of of using it. Oh, >> get it running. >> Step one is get it running, folks. >> There. >> There you go. >> I was holding the wrong key, I guess. Uh made this. >> Yeah. So, uh, can be incredibly fast and, uh, we've we've both found that there's there's a value really in just being able to rapidly speak your thoughts.
09:10 They don't need to be perfect and it'll they can be messy and they'll get cleaned up later. Um, but that just gets you to a working artifact even faster. >> Yep. So, uh, once you have that all set up, let's, uh, try and fix a bug by voice. Um, so in your coding agent, you can just, uh, ask it to, uh, for example, extract the off client into its own module.
09:33 Uh, and that's just a way to easily speak it. You're kind of eliminating that friction of having to type everything out. Uh, and moving more at the speed of thought, which if you're a Vimmer, uh, you would be proud of >> indeed. Um and then you know just to reiterate like if uh both of us will tend to run um a terminal throughout the workday and might have eight or more sessions all focused on a different project um you know that has all the context and the tools built up specifically for that and you can kind of drive
10:00 by each one of them and as one's working in the background you're directing the one the next one to go and proceed down the path to get you there faster. So it really compounds quickly. Excellent. Yeah. So you're you're saying the outcome uh not the keystrokes. Uh so you want to think about where where you want to get to and don't have like in my brain it kind of like changes the way that I interact with claude and with with my computer in general.
10:28 Uh because I'm really having more of like a speed of thought thing and it's it's a little bit different than the words that I would actually type out if I were physically typing them out. uh which gets me to a result faster, but also I can catch myself occasionally like stopping and pausing in weird ways because I'm like thinking about the next buffered set of words that I'm going to say.
10:48 >> Absolutely. Um so again, just uh significantly faster if you're mostly typing. Um most of these models that you'll experience whether it's Whisper Flow, even the open source stuff are fine-tuned in such a way that they're capable of picking up exactly the file name that you're using. they're aware of, you know, highly technical terms and operating system tools and utilities and so, uh, it really is significantly more more rapid to get to your outcome.
11:15 >> Okay. And then a quick word on why the different tools or where where does some tools excel. So, Whisper Flow is uh, ideal and beloved and people everywhere are kind of waking up to it and using it, even folks that are not, you know, traditionally developers. And I would say that one of the places that benefits people the most is just the the rapidity of setup and how easy it is to do because if you can drive Mac and if you can download an app on Mac, you can use Whisper Flow.
11:37 Um the trade-off is of course there's you know a subscription involved and uh folks might not want be doing that especially if you're doing sensitive uh work whether it's coding or other types of you know document review or you work in a regulated industry. uh there's downsides to having everything that you say uh and all the file names etc and contents and PII sent off your machine into the cloud.
11:59 Um and so that's where you know using a tool such as Handy can be uh significantly um more comfortable for folks. Mhm. >> And so that that's the tool that we shipped with today for this workshop. We're going to use uh handy and that's just about local ondevice dictation because models are getting smaller and smaller through quantization and it's now possible to run a very performant uh speech to text model just on your consumer hardware.
12:25 I want to bake a cake and there's three steps to that. The first step is to create the batter. I don't actually know how to make a cake. But the second step is to get some eggs. And the third step is to put it in the microwave. >> There's sugar in there. >> Okay. >> Yeah. >> Uh well, you can see my point. It did get it. The reason that you would use a tool like this over just like the built-in like Mac dictation, for example, is that it's actually listening oops to what you are saying and it knew that I was talking in
12:51 list form. So, it made a list and and that is a big piece of it. It's also very very good at understanding like if you're talking about file names or function names, it's going to put those in there correctly and it can autocorrect uh or you can correct it if it puts like the wrong word in there. I tried to say Swix earlier, but I messed up and I thought for sure it wouldn't say that.
13:11 And I could go correct that and then the next time I say that it will autocorrect to Swix correctly. And you can set it up so that um like all of these have this way of defining the context that you're in. So, it knows if you're in an email, maybe you want to type a little more formally. So, it'll use a little more formal text. But if you're in Slack, it'll be a little more casual with punctuation and things like that.
13:33 And it can customize based on the the place that the text is going, which is really nice. >> So, now you try it. Um, did we say the the command to run to get it set up? >> Yeah, it should be already in there. >> Should be next, I think. >> Okay. >> Yeah. So, we're going to have you try it and talk to the repo. There it is. Yeah. >> Uh, so you can just ask Claude if you have that repo set up.
13:58 Just say, "Set set up Handy for me." And it should automatically set that up for you. >> Conference Wi-Fi. Uh, willing. >> Yep. Do you want to open up uh Handy and show it look like [clears throat] this? Yeah. >> Uh, so it looks like this. It's got a cool little uh Mickey Mouse glove or hamburger helper glove. Um, and it's really just these settings.
14:19 I don't know if I can make that any bigger, but um yeah. Uh so the main thing that you'll use is a uh a way to like start the transcription shortcut. In my case, I have it set to my super key, which is control shift alt uh and command al together. Um but typically it's something like um just function by itself often is is one. Uh and you can set that up however you like.
14:49 um option space I believe. Yeah, that's the default. Um and so that's just what you're gonna hold down when you want to talk to it. >> And then you can set up uh different things like pushto talk uh what microphone it should use uh and all all of that. Is there any specific >> if you want to go deeper you can look at models and then if you you know want to get really really nerd out on it you can pick which model is ideal for you go with the recommended settings for the first time if you haven't used it before.
15:15 But just be aware that there's the option here to change it and then it will just get downloaded from the cloud to your machine so that you can run it there. >> So uh way more customizability and and control when you're using um you know kind of like an open source uh product like that. >> Yeah. >> But yeah, we'll give you guys uh a minute or two just to run this and uh if you've not tried Handy before, try to get the transcription um key working and say something quickly into your computer.
15:39 And don't worry about sounding ridiculous because everyone next to you will sound similar and no one's gonna laugh at you here. So >> yeah. And and uh just for reference, Zach and I both are remote employees. We both work from the comfort of our respective houses. And uh so we don't have co-workers to uh distract or annoy with our constant talking. Uh but there are workarounds for that.
16:01 Um, I know of companies that will get like this like little like pencil mic for everyone and they'll just kind of like hold it really close and kind of talk into it or mumble into it pretty quickly. Uh, Zach and I have also been playing with uh the DJI mini mics, like the little wireless lav mics, and I've been in loud coffee shops and I just put that on my shirt and I just kind of whisper into it like this and it can totally hear me just fine.
16:27 And no one else I just look like the weird guy mumbling to myself in the corner. >> Yeah, it's good. uh which I'm totally okay with. >> Yeah. And and then still the the actual transcription accuracy is like 100%. So >> highly recommend those if you've not tried them before. >> Um but yeah, if you run this command uh then the skills in the repo will be found and Claude will do all the correct steps to actually pull down handy for you and get it started.
16:48 >> Yep. >> So >> anyone need more time for that piece or we >> have a question >> or questions so far? >> Okay. All right, one thing that we want to clarify in this is that we only have an hour. So, we're kind of moving along trying to get as much into this as possible. Uh, but you have this repo and you can uh try this out as well if we are kind of going too fast.
17:14 Uh, but also do stop us if you have any questions. >> Y >> So, one thing that you can try for example is just like find a failing test in the repo I brought uh and fix the root cause and it will just kick off. I mean, this is exactly what you would type uh but just a nice easy way to to get going. Zach, also I haven't tried this yet, but Zach is someone who will uh use that remotely or away from your computer while you're kind of pacing, right?
17:39 >> Oh, indeed. Yeah. So, you can also get Whisper Flow on your phone. Uh, and that is really ideal because then if you pair the remote control feature with Claude where you can say any session that I started on my laptop or my desktop should follow me on my phone. I should be able to access it there. you can, you know, be working at your desk for two, three hours, do your morning check-ins and your calls with colleagues and do some focus work and then get up and go take a long walk.
18:01 Uh, I like to go into the woods for a bit and maybe take like an hour and a half hike and I have my phone and so like, oh, this I just figured out because I'm walking like shower principle how to solve this bug. Now I take my phone out and I go and find that exact session and I fire off a 30-cond voice memo and then Claude's still working in the background making me look like I'm an excellent employee who's not literally phoning it in at that moment.
18:22 And then I continue my hike and then I go home. So, uh, >> recorded, right? >> I don't think so. Um, but yeah, highly recommend that. >> Yes. >> Yep. >> Hey, what were we working on? >> Yeah. No, you guys. >> That's a great question and this uh digital uh blacksmith here has got a tool for you called sessions. >> So the the question the question is uh just to repeat it for the the recording is uh how do you manage when you have like 10 different cloud sessions open and and understanding what's going on in what and that
19:10 is a hard question. Uh there's two solutions I have for that. Uh I drive everything through T-M uh and I I built something called fleet actually. I right now have 12 agents running simultaneously. Uh, and I list them out by the the project that they're in and the session name that they're in. And this is like the T-Mux window. That's what's on the left.
19:28 And then in the middle there, that is the Claude summary that it gave. Like it automatically generates this for every session. So it tells me kind of an example of what we're doing. And I can kind of quickly see that. And I can just like scroll through these and I can be like, "Oh, I'm going to go I was messing with my dot files earlier. Switch over to that one quickly."
19:45 Uh, and then bring it back. And so that's just something that you can get access to. Uh they're in the JSL files. You can write your own tool. You could use a tool like this. Um but the other thing is if it's not on by default, um you can turn on a setting and I'll have to look up exactly what it is or you can ask Claude to just turn it on. But it's this recap feature where after you come back and focus on a window again, Claude will attempt a recap and it'll just give you this like italicized summary of where you're
20:10 at right now and maybe what it's waiting on from you. And that's a great way to um to do that as well. And then part of that same session uh fleet thing, you can see that sessions up here. This is just telling me um sessions is done. Like it it finished whatever task it was doing. And I can click up here to go to it and that'll clear out of there. So it's like a notification system.
20:31 I just didn't want everything like yelling at me with like different notifications. I know people use like Zelda themes or like like sounds from video games to to do all that. When you have 12 agents running, it's way too much. So I just have like a visual indicator for that. Did you want to show folks where they can get fleet? >> Oh yeah. Uh just nicknfleet on uh GitHub.
20:51 >> Uh and then also if you've got less than 10 tabs or if you want the lower tech solution, you can also just rename your own sessions rename. You can also within your terminal rename your own tabs. So if you do both, right, you can get to a point where you're just managing that yourself. >> Yeah. >> Um did you and the sessions tools available through brew?
21:09 >> Yeah. >> Right. So you just brew brew install sessions if you're interested in that. We'll have links to all of these afterwards as well. >> Great question though. Thank you. >> Yeah. >> Cool. So, the next thing that we're going to do is uh there's another skill in the in the repo. So, if you type this or if you say this to Claude run my workshop check-in, this is the part that is opt-in only.
21:29 You don't feel obligated to do it. Only if you want to participate. But it'll ask you a few simple questions. I think it's four. And you can answer them all in one line if you want. And just hit enter. And then uh that'll get posted to our back end. And then a little bit later on we'll have a data visualization of where the room started and ended there.
21:46 >> We'll show that in a little bit. >> Y >> um let's move on to the next section. And in here we're going to talk about loops. So now we're kind of like cooking with our voice uh but we're still kind of babysitting everything. Um, so the next step of that is to not eliminate us completely, but eliminate the non-essential things so that we can focus on the things where we're good at and where we focus and or where we're we can be focused and contribute actively and not just babysit the agent as it's going.
22:12 And that unlocks you to let the agent kind of go for a bit while you go do something else. And that something else could be a whole another agent in another tab doing another thing like I had with 12 of them going at a time. Or it could be go take a walk in the woods. Y >> anything. Um so just to recap, uh an agent is uh like there's lots of definitions for this, but like one could be like it's a model in a loop.
22:34 It's a model with tools in a loop uh where you're asking it to do something and that will think about it and then it'll act and then you want it to observe something about that and then decide what to do next. And oftentimes that's go back through that loop again and go through it over and over and over. And there's several ways that you can do this.
22:52 You can build your own of course uh you can build the exact workflow that works exactly for you or for your team or for your company. Claude also ships and uh other ones do as well like codeex and and others. Uh they ship some built-in helpers to make that very easy. We'll talk about two in in depth here and that's goal and loop. Uh and there are ways to let claude or or your model go unattended for a bit, your agent go unattended for a bit uh and make decisions on its own.
23:21 >> Yep. So, I'll just say if you find yourself now uh commonly, you know, you're getting lots of good work done with cloud, but you notice that you're constantly having to go back and say, "Okay, no, do this again. Okay, that wasn't quite right. Okay, it's almost there." Uh, you should be thinking and reaching for goals and loops. >> Yep. And I'll also say that uh from everything that we read and ingest and from talking to folks as well that are you know constantly practicing all this it seems that loops and similar
23:43 are going to be one of the kind of key unlocks for just higher levels of productivity that that are also less demanding on you to monitor and babysit. So >> highly recommend looking into them. Um but quickly the difference is a goal is there is a clear definition that that I want you to meet and so I want you to work until that clear definition is met.
24:05 It could be refactor this simple file until these tests pass. It could be I want you to read this website repeatedly, you know, for a long time until an update is there and then let me know or this like, you know, this flight price changes, etc. Um, that that has a clear defined state. You're trying to get to it and it's going to terminate at that point.
24:23 >> A loop is I want you to repeat a task, you know, non-stop until I tell you to cancel it or until the timer that I set for it is expired. So, you know, different ways to leverage them, but you know, that's an important distinction. >> Uh, and it really just does that iteration in both of these uh until it passes. So, for goal, for example, you could say like the goal is I want you to do this until all of my tests are green.
24:48 Like I have some failing tests. The goal is to make them all passing. The goal is to get to 90% test coverage. Whatever your goal is, it's something that is tangible and that is measurable by the LLM uh by the model. Uh if it can do that, then it can work towards that goal. And once it reaches that goal, it will uh it will stop. But until then, it will try whatever attempts that it has to get to that goal.
25:09 It'll check whether it's gotten to that goal. And then it'll fix whatever it can and start the loop over and uh or fix its its approach to it and start over. So you don't have to reprompt. It's going to reprompt itself continuously until it gets there is the go the the entire goal of goal. Uh and so the the real thing to just kind of internalize here is that uh as we've all experienced, especially if you're working with cloud and other agent tools, you know, with any regularity, uh they will punch out way sooner than
25:40 the task is actually done done the level of done that you would be happy to sign your name to and say this work is complete and send it to your colleagues to be scrutinized. Right. And so that's the gap that we're trying to close with these primitives of loop and goal and eventually schedule as well. >> Yep. So, uh, goals, they're made up of these four boxes.
26:02 None of them are optional. Uh, it has to, um, for example, like reproduce a bug, uh, with a failing test, implement the smallest possible fix for that, lint, type check, make sure the tests are are running in green, uh, and then like afterwards, it can open a PR um, to submit that. And so, it's really bringing back like that idea of like red green refactor.
26:21 like it's a a really that's a really good loop for it to get into where it can write the failing test, make write the minimal amount of code to get it to pass and then verify it. And you can do a whole bunch of extra steps in there uh automatically as well. Uh things like like adversarially checking or diffing uh based on like another model looking at it and and giving it a review uh and then pulling that data in and uh going from there to update it.
26:49 So yeah, like we said, uh agents will punch out early. Um set up your loops and set up your actual uh re repeatable task that you're constantly doing, the code bases that you care about so that it is impossible for them to sort of lie to you and it's impossible for them to uh kind of bail out before the task is truly done. >> Yeah. A big thing that I had I was working on like this whole agent loop uh that was not part of goal.
27:11 It was like my own loop thing and I wanted it to like give me some verification that it was actually like doing the work and testing and everything. And so I had it just like make sure that you're running the test and it would like say like oh when you run the tests touch this file called like case tested the the file the the agent was called case. Uh and if that file was existed that means that it ran the tests.
27:34 Well Claude figured that out pretty early and was like all I have to do is touch this file. I don't have to run the tests actually. And so uh that was like it lying and it's it's a good programmer right it strives to be lazy like us and that's exactly what it did. So you have to make the lazy route the the correct route and so I had to add in all of this like check some verification in that file to make sure that it was actually running correctly.
27:58 >> Yep. And then so just again to the these visualizations just to drive it home. Uh so a goal is a bounded place we're trying to get to and as soon as you're done I want you to terminate. The loop is going in a loop. think in a in a sense like a loop is a little bit um more creative and and more more adaptable to like wider purposes. For example, like with as little as one loop command, you could set up an autocompleting dev implementation loop at you know for a given codebase where you say read this slack.
28:24 You have MCP access to Slack. See my my colleague complaining about bugs and asking for feature requests for this project. When you see them, add them to linear as subtasks. then take a linear subtask off the queue and implement it and merge through all the tests passing and then deploy it yourself, right? And then close it down. You could maybe split that into two loops.
28:44 It might be a little bit cleaner. The point is that um it's incredibly um just reusable and and adaptable to many different situations. You could also think about regular reports that need to happen and we'll we'll talk about when schedules a little bit more appropriate for that too. But a ton of different things you can do with loop if you haven't tried them yet.
29:01 >> Yeah, for sure. And this is where all of this is going too. Uh, one thing that just came out like last week and we've been playing with it but didn't have enough time to put it in the slides is uh, like claude tag for example. Yeah. >> And that is just a simplification of being able to run loops and have some kind of trigger that that triggers in Slack to get it to go.
29:17 It could be a message from a colleague or an update to a PR that triggers it and it does everything autonomously on a cloud environment. So, uh, it needs to be able to run in that independent loop without you so that it can continue working and be efficient. >> I've got Claude Tag set up in my personal Slack now and, uh, about seven or eight versell Eve bots, their new agentic framework.
29:40 And, uh, they're doing a ton of different, you know, independent work on my business for me. And then claude tag is the the main interface for that. So, >> yeah, >> we see this as all just kind of converging towards the same primitives that you need to be kind of very familiar with. So here are some examples of uh goals and loops uh in the wild. So like a goal for example would be I want this test suite to be green.
30:01 There's failing tests. Let's make the goal to have everything be green. That is something that's measurable. When there's no more failing tests, it's done. Uh another one is every single file in the source file uh source directory is under 200 lines. That's something that's very easily measurable by running various commands that that uh Claude can do.
30:18 So it'll go in and do it. uh a loop for example would be every minute I want to run the tests and make sure that they're all passing. Uh there's much more efficient ways to do this. So this is a little contrived but that is one way that you could do it. And another one is like check the deploy and ping me when it's healthy. So this could be like you have like a long CI or a long build pipeline that's getting your your code deployed and you want to be notified when it happens.
30:39 Again, there's easier ways to do this, but Claude could also do it and you could just have it loop and check every five minutes and then it it can act on that. It doesn't necessarily have to uh notify you. It could do run like an additional script or do something after that happens. >> Loop and tell me when my favorite video game is finally available for publishing so I get in early, right?
30:56 And text me. Um tons of stuff you can do with it. >> Oh yeah. >> Okay. And then uh we're going to introduce the concept of git work. So if you're a developer, you you've probably used them before. uh think of them as a way to give each independent agent a a kind of isolated and safe to work on concurrently copy of your uh git repo. And so there are issues with this.
31:21 We won't get too deep into it. Um there's experimental things coming out that are a little bit different, but uh we found success in using git work trees. And so you can run your agents in parallel. And that means that you can think of uh breaking down like a full afternoon or five or six hours worth of development, implementation, bug fixing into discrete blocks and assigning those discrete blocks out so that they run concurrently and you compress the time down to, you know, 45 minutes or an hour.
31:48 >> Yep. How many of you are using work trees today? >> Cool. Like half the room. >> How many are using cloud-workree to do that? >> D-Worktree. couple. Okay. >> Yeah, that is one of the easiest ways to get started with it. You can literally just be in your repo uh you know whatever that is and say claude-workree and it will just create a work tree right there and then start a cloud instance in that work tree so that you can immediately start parallelizing your work and going from there.
32:17 Uh and depending on your repo setup, that's um that's easy or difficult depending on how difficult it is for you to run concurrent models uh concurrent versions of your repository locally. One thing that we have is uh a a monor repo. I'm sure a lot of people work in monor repos and it can be very difficult to run multiple instances of that uh in different work trees because it's like gigabytes and gigabytes of node modules for example >> virtual machines and docker.
32:44 >> Yeah. Setting up the right build systems. Yeah. uh database all of that the right ports like being mapped correctly that's where like more cloud environments come in uh that can do all of this as well um we've been experimenting with cloud tag uh Devon and others uh but that's another way to parallelize this work so that agents can work independently in the same codebase without ever conflicting with each other and then open up independent separate PRs that never uh bleed code changes from one to the other and of
33:12 course this just loops all the way down Um like we think of these as like atoms. Um every block is just the same loop at a different higher altitude. Uh and you can think about that like with literally everything. Like we definitely consider ourselves like loop engineers now. Uh because that's what we're we're thinking about when we're thinking about like the atomic bits of work that we're actually trying to uh complete.
33:32 And we could have like a a loop that's like fixing this bug while another loop is fixing this and another is fixing uh around that. and think about like how we manage all of that like with with that fleet tool I showed you like it's thinking about all of those in a loop. I am literally looping between all of them going from one to the other to keep things moving along.
33:53 >> Yep. Absolutely. You ran a loop that was uh two hours long the other day. Right. >> Yeah. Regularly the the loops I do a lot of upfront planning and then I can just let Claude go and it will go for two three hours at a time without bothering me at for anything. And it has all of these checks built in to where it's like keeping things small, keeping things readable, reusing code that we have, uh running the tests, getting feedback from Gretile or Code Rabbit or whatever uh code tool we use.
34:21 We use them all honestly. Uh and then um implementing that back in and then once everybody's happy then it brings it to me and that's something that I can then look at and then approve for someone else, a human to look at. We still have human in the loop reviews uh on a lot of things. And so uh but by the time all of the dust has settled, it's working.
34:40 It's shown me that it works. The code looks clean. It looks like code that I would have written. And it's ready for a human to actually review it. And it's not just slop code that I'm throwing over the fence for somebody else to to worry about. So now we're going to have you try it. Um we're going to use loops and goals. Uh and you can hand over a job and walk away.
35:05 Uh so in that repo you can run slashgoal and then tell it um this bun command bun playground goals check.ts and the goal is that when you run that check it will show five out of five. So everything's passing everything's working. Uh and if that's not the case immediately claude will just spin and run until it is complete. >> Y give you a minute or two to run this and see what that output looks like.
35:30 >> Yeah question. is it? >> Oh, nice. >> Which program is that? Uh, >> oh, okay. >> 3.5 gigabytes. Uh, >> what you talking about? >> What you talking about in there? working on it all day. >> Nice. >> Cool. >> Nice. >> Yeah. >> Okay. >> All right. >> Do you just run compact often? Is that like how do you manage? >> Yeah. >> Uh yeah, when you're working with claude, um one thing I'll show in mine here in this one, um you can see down at the bottom I start.
36:32 >> Yeah. Okay. Yeah. >> Oh, >> yeah. >> Gotcha. Okay. Okay. >> Appreciate the bug report. Thank you. Yeah, we'll fix that. >> So, yeah, for the recording that was uh it's failing on a 3.5 gigabyte JSONL file. We didn't anticipate that. I'm sorry. >> Thank [laughter] you. >> Yeah. No, you're good. >> Oh, it's great. >> You're good. [laughter] We did not I have not seen that yet.
37:03 >> Cool. Yeah. Question. >> Yeah. Sorry. >> Yeah. So what's what's the So like obviously there's like usage limits. What is the impact of running things like continuously? >> That's a great question. The question is what's the the impact on usage limits for running things in a loop? And that is something that you definitely need to be cognizant of. >> Yeah.
37:24 I'm going to say this gentleman built and it's actually in this repo if you've cloned it. It's in the cloud settings. It's called the ideation plugin, the Nick Nissi ideation plugin. Uh, one of the best ways to handle it is to just do your planning up front and to get to deep and abiding clarity on exactly what you're trying to build, what finish looks like, what failure looks like, what you're not going to include in it, what's scoped into it, and then set up the loop that's going to be a higher token burn for sure.
37:51 >> Yeah. >> And the way that that plugin specifically, that's why I don't use like uh the built-in ones often. I use my own is because it's specifically trying to break things up into these manageable specs uh where it can run them in independent context windows and it will try and I'm mostly thinking about like context size in that case. Um we we're fortunate uh that we don't really have usage limits currently uh but I can see a a very near future where that is the case.
38:19 Um, but it's it's meant to uh do things in these like small atomic pieces that it can do in independent context windows and then not have a bunch of blow uh that's comes associated with that to keep it more manageable. But yes, running things in loops and there are things as well like um in in more recent versions of cloud there's this this thing called dynamic workflows uh where it will effectively like build like a small state machine to run a whole loop in uh in various different ways and has gates between them uh
38:49 and that can be very tokenheavy. I've seen it get up to four or five million tokens that it's used for a single task uh spread across several workflows like it'll spin up 35 agents and run them all simultaneously. Things like that. >> Yeah. Well, we're also going to talk about verification gates a little bit here in a bit. And I think that's the other, you know, solution to that problem, which is really the token bloat or token waste in that in the in the worst case when you're just looping um kind of comes from like
39:13 the unbounded exploration without anyone telling you you're off track and without the model being able to determine that it's wasting wasting term, you know. Yeah. And so having uh you know doing the the planning up front even burning extra tokens on getting that right using max thinking there and then setting up your loop so that there are hard verification gates so that it cannot just get lost is are two ways I'd think about it for sure.
39:36 >> Yes. >> Go ahead. We built a UI or something and we come in as it turns out. We don't want to completely Yeah. Yeah. Great question. >> So, the question is, uh, how do you not babysit? Like, what are the the ways to, uh, get it to work more independently without you having to babysit it along the way because it's making a lot of mistakes? And the the solution that I've come up with uh, for that is to let it make those mistakes, but record them and then uh, keep track of that so that it doesn't make those mistakes
40:26 again. for every loop that I run, it run I run like a little uh retrospective agent uh that will go over the performance of the loop at the end and it will see okay what did we do here? Why were we wasting a bunch of time here? I had to correct it a bunch here or here's where it was running a bunch of tools without like changing things and like figuring like running into messy loops like that.
40:49 >> Computer use. >> No, >> no, >> no. This is just at the end after it's uh I'm not using computer use. just at the end after running a loop where I have had to do a lot of babysitting uh we work through that I work through that with Claude so that it builds up this memory file that gets loaded back in uh at the next session that we do to tell it things and it it also does some curating of that.
41:10 This is like a separate memory system that I'm just kind of like toying with on my own and it's markdown files uh but it's it's something that gets curated as well. So it thinks about things in short and long-term memory. Uh so short-term things will get uh filtered out over time. longer term things are like the core tenants that we keep running into that I don't ever want you to forget like this is where our design system lives this is how you test this specific uh component for example uh things like that and then
41:37 that gets loaded in and it gets broken up uh using um progressive disclosure so that you know if I'm not talking about testing it's not going to load the memories about testing specifically or if I'm not talking about the react portion it's not going to bring up React specific memories it'll only bring those up when we actually touch React files >> I'll also say you know specific to front end um the way I might approach that problem too is cloud design first similar to the ideation the higher spend up in the beginning
42:01 if you haven't tried that yet that's like a a specific like instance of claude adapted for design and it is able to develop for you like a design system that then persists across sessions you can reuse it so like starting with the upfront planning first and then second and we'll touch on this a little bit later adversarial review um you know injecting either with hooks towards the end of an implementation loop and specifically loading skills that are designed to test UX and front-end uh best practices.
42:27 And you want that piece to be as negative and pnicity and nitpicking as possible and to fail those first 25 builds and say this is not good enough. Keep going. Fix these things. Right? That that's how like once you get those pieces in, you can engineer a loop that yeah, you go and you make your coffee and you come back maybe 40 minutes later and you've burned a bunch of tokens, but you're actually happy with it and it's pretty close to merging.
42:52 >> Yeah. We'll take one more question and then we'll move on. Yeah. room full of operators having something that >> that's a great question. So the question was uh have you experimented with any passive voice or the system listening to the ambient discussions of people in the room and other operators and then working on that? Uh kind of not yet. One thing I will share that has been insanely and this is a little bit surprising of an answer but has been insanely um uh impactful and so I think that that actually would
43:26 bear fruit and is worth exploring is uh we did build a system that allows uh everyone in the org to blog in a very uniform way with a common voice and it fixes you know AI slop and it uh allows the everyone to produce a draft of a blog that looks exactly like you know the expert at making blogs for work OS would do it and the highest leverage connector that we added to that was granola >> because as soon as you pipe in all the calls from people people are saying oh it's just this little idea that we just had while
43:56 talking to customer if we had a blog post on that it would explain it and that is all that needs to be said and you just pipe that's a single MCP call and the blog bots's off and running and what comes out is excellent so I think that what you mentioned the passive listening ambient listening and then also computer use is going to get pretty freaky at some point soon where you combine those two and it starts to say here was your year and these 16 specific speific actions or what actually drove your business.
44:19 So now I'm going to do those like at rapid speed or you know etc. Yeah. >> And the only other piece of that I would say is experimenting with Devon and Claude tag and them kind of ambiently popping up in cloud or in Slack messages. Uh that's more annoying right now at least than uh helpful. >> But um yeah that's >> oh we got to cook a little bit here.
44:39 Yeah, so there's test goals as well. And that the whole point is that take that repo with you if you want to continue practicing uh loops and goals and you can use the playground scripts there and then you can adapt them to your own programs. >> Yep. Never tell it to try harder. Uh it doesn't know what you mean by that. Uh always give it some kind of >> you can rage at it though.
44:55 It's it's it's good for the soul. It's cathartic. You should do >> we have a whole channel called bot hate where we're just screaming at the bots for doing different things. Um but yeah, f when you're working on these and it's failing, fix the system. don't don't uh fix the the code itself. Fix the system that generated the code so that next time uh it can be fixed automatically or it doesn't hit that in the first place.
45:15 And that's what my retro uh idea does. And speaking of that, like this is how you you prevent that or you how you prevent it from going off the rails. You give it ways for it to verify itself and you trust those gates and trust that they will prevent the model from moving on until those gates are met so that you have more confidence in letting it go on its own because it's going to understand what done means and done for me does not mean what done means for the bot.
45:43 So we need to to come to an agreement on what it actually means and that's what this verification is. >> Yep. So it it could change depending on any project. It could be if you've got a typescript it's going to be the obvious ones like lint and bill and you know did it compile actually safely. Uh when I'm writing for my own uh blog I have a system that the stuff that drives me bananas that are non-negotiables for a PR.
46:04 So like all the images must return 200 off the CDN. The open graph image must be formed perfectly. There's other checks that are in that script and a failing um you know status code from that bash script fails the entire build and then the system picks it up and then it tries again and it works on fixing all of that. But it it's those problems are solved before it even gets to me and the end result is still a working PR.
46:25 >> Yep. And this is what I meant by um you just make it make the the trusted route the way that it should work the lazy route. Uh and it just make sure that it has all of these gates. And there's various ways that you can do that. You can force it uh through things like hooks. Hooks will prevent it from being able to move on or do things uh or it will make it like run linting after it runs after it writes uh updated files, things like that.
46:49 And those are ways that it just can't avoid. It won't stop. It won't forget about that. One common thing that that we've seen like when we're doing this is like we'll tell it a whole bunch of steps. We'll be like, "Here's the seven things that you need to do every time we work together." And if our conversation goes for like 350 400,000 tokens, it will forget steps four and five because it just like that's so far back in the the context that it doesn't remember to do that.
47:10 Uh and it gets lost in the minutiae of what we're actually doing. And so we've had a lot of success in breaking that out into like a state machine in Typescript for example that forces it to go forward. We've been building on top of PI uh quite a bit to to do that. Um so that it's not something that can be avoided. It has to move from one step to the next and there's no way to get from step C uh from step A to step C without going through step B.
47:37 >> Y so we [clears throat] talked about some of these uh we can skip ahead. One of my favorite things to do that is uh bearing the most fruit I would say in the last like three four weeks since I started experimenting with it is add a hook to Claude. And the the fun thing about hooks too is that you can just tell Claude to add a hook, right? And it'll do it correctly.
47:53 So add a hook that every single time you're working on something that is either very sensitive or urgent because it's going to unblock some production issue or it is of a certain complexity and you can even be the judge of that. then you must fan out your code diff for an adversarial review from codeex 55 via the CLI. So you install Codex yourself or you get Claude's help with it.
48:14 You off into Codex and you log in with your OpenAI account. You do the Google, you know, the oath roundtrip and then from that point forward as Claude is building with its, you know, constant and ever cheerful self saying, "Oh, I nailed it. It's perfect. Don't worry. Don't look at it." Right? And then it's about to open the PR, but the hook says you must go and get an adversarial review.
48:33 Uh that is finding significant issues on the regular for me. And that which reduces the overall time because it's not kicking off expensive builds and reptile reviews we had to pay credit for yet. There's something finding it locally locally um before it even gets to that process. And so we have found even uh experimented with this couple weeks ago and it was finishing a pretty um major migration of a system and found that having four different agents from different companies review the same code all of them found
49:03 separate things and different things which is not surprising given that they've got slightly different training you know uh processes and weights and everything. And Codex does ship some skills. Uh they ship a review skill, an adversarial review, a rescue, uh all of these easy ways for Claude to call it, get advice from uh Codeex, and then bring that right back in so that Claude can act on it in different ways.
49:25 And this is a great way to get them to work together um when it and it's a a really good way to just get that multimodel uh perspective on the code that you're writing and uh how it's performing. >> Yep. Because you know the one model will say especially if it's still in the same session it's going to adapt back to the context it still has where it told you you were telling it I want you to complete this goal and it's going to try and continue to confabulate helpfully and pleasantly and tell you that it did that goal.
49:51 Uh whereas you know a model that doesn't necessarily have that same context but has a very uh succinct prompt about here's the inputs, here's the outputs, here's the goal can start finding issues in the code immediately. So uh yeah, basically the idea is don't um trust the model's confidence as much as you might enjoy or feel partial to Claude or any other agent.
50:11 Uh you trust the gates that you set up that make it impossible for the system to lie to you. >> And when you find that it did lie or figured out a way around that, fix it. Y >> don't fix the code. So we can try that. Um now with verification verification gates um we can make something done uh to it so that we prove that this actually works. Uh so in that repo you can just say like add a hook that runs lint uh type check and the test on every change and fix its uh flags.
50:40 Fix anything it flags. I can't talk. All right. You can also ask it to get a second opinion. So ask it to just fan out and give you an adversarial review and then fix anything that it finds. And this way you're having the two models work together. One of them is fixing it, one of them is reviewing and they can go around and around on that um and get the work done.
51:02 Yeah, you can also, by the way, for what it's worth, uh it doesn't always have to be adversarial review. You can say, >> you know, scope this workout um until it's it's very clear and then divvy up the pieces where it's like the clean seam so that you can both work on it uh more rapidly and and burn more of my money faster. >> Yeah. A big thing that I do is I we have access to all of these models uh and so I will often use this adversarial piece uh in the planning phase.
51:27 So, I'll have Claude and I make a plan uh using the various skills and and ways that I would make that plan. And then once I have those artifacts for it, then I'll pass it to Cla to Codeex uh and have it do like a review on it and say like what what makes sense, what doesn't. And it will often pair it down like a lot be like, "Well, this kind of doesn't make sense."
51:46 And it will send that right back to Claude and Claude will be like, "Yeah, Codex is right. I see what they're going for here." And we'll concede a lot of the times. It's very amunable to uh influence in that way. Uh but uh it in the end it works out to be a better plan and for me I'm always making it right like HTML plans and so it's something that I'm following along with too uh feeling like I'm like the dog at the computer like I'm contributing [laughter] but we do get things done.
52:14 >> Okay, scheduled tasks. So uh we have a few minutes left. We're going to kind of rip through this. Uh one of my favorite things to do when you're first setting up. So what are scheduled tasks and why do you care about them? Uh it's that thing that you still come into Slack every Monday dreading to do. It's like pulling together reports and I have to go dig through notion.
52:28 I have to go find that Slack about how they wanted the format done and I had to go download this CSV and I have to look through these like customers. That stuff can be uh you know first worked on in a session until you get it working. And then as soon as you have it working, you can say now schedule this for every Monday at 9:30 a.m. because I never want to do it again myself.
52:46 And that will actually kick off and save a durable, you know, cloud session in a VM that um is able to do that work for you on a repeatable basis whatever you schedule. There's utilities in the schedule command like list and then you can see what you have running and you can modify them or delete them. But think of it as anything that you need to do for work or for yourself that is tedious and you'd like to automate.
53:09 Schedule is the command that you want to be reaching for. >> Yeah. >> Go ahead. Go ahead. >> Oh yeah. So just quickly, since we've introduced a bunch of these primitives to you, just a a quick word on when to think about them. Hooks are things that are going to run automatically in certain cases. Maybe you could write a hook to say always run, but then only proceed if this case exists, right?
53:28 Because it's a bash script, for example. But it's things that you want Claude or any other agent to always remember and do and never forget and not be able to forget. Uh goal is if you have a clear goal in mind, like I did just need to get this PR over the line, so I need you to fix these 13 tests and then also refactor this off. TS and once that's done and you can prove it's done then I want you to stop and stop burning money.
53:49 Loop is uh I want you to go babysit this PR you know I want you to um constantly ping here and then tell me when uh this web page changes and this is the data that comes back. I want you to go through Slack and get all the complaints that are coming in from customers and file them all as linear tickets. So run until I tell you to stop or in a timer I've given you.
54:06 And then schedule is something that you want to be durable. So I want you to run every two weeks, every Monday, every Wednesday, every Friday, every day like a crown job, right? And I want you to do some amount of work and you may have secrets involved too, but these are the kind of building blocks that uh together you can kind of create like fully autonomous systems and and automate a lot of the TDM out of your workflow.
54:27 Mhm. >> So any of these are great candidates if you've done any of this dependency bumps, status reports, eval runs. Um you know if it recurs and you have access to clawed and scheduled tasks, you shouldn't have to keep doing it manually. >> Yep. The biggest one for me is always just what like what do I need to prepare for meetings? That's the biggest thing for me a lot of the time is just like getting things ready for that on a on a specific schedule.
54:53 There's other triggers as well. It doesn't necessarily have to be timebased. It could be um different things come in like new pull requests come in. Uh how do you handle that? It could automatically kick things off. It could be messages in Slack come in and it just takes it over. We have a like a whole nits channel for like our docs for example and anything that gets posted in there immediately gets a linear ticket created and then a bot automatically tries its best to fix that.
55:19 And like 90% of the time it can fix it on its own because they're just little nit thing nitpicky things. Uh and it'll go fix that automatically. Otherwise, those would just sit on a backlog that we never prioritize. Uh, but this way, as soon as they're filed, they are fixed, which is just such an amazing feeling. I also have this one that runs weekly uh on my um website, and it's just like a fun little visualization that just shows me what I did for the week, uh, specifically with like token usage, and it shows me like
55:48 the trends of how I'm using tokens over time, the code that I'm changing. >> It's not his personal spend only. >> It's not my personal spend. horrified. >> Uh, but it also like summarizes my week from PRs. So, these are all of the PRs that I opened. Uh, and this is a summary of like the work that we were specifically focused on for that week. And it just goes back uh I think the last four months is how far.
56:10 >> Yeah. >> Because I thought this was brilliant when Nick told me about this because it's like you don't want to scramble before your performance review and go figure this out and then, you know, mine it every single time, especially if you know it's coming twice a year. So, it's just easier to just have it set up once and then automated. >> For sure.
56:26 Okay, so uh just a quick tip that I found that works really well uh if you are trying to set up your first schedule task. One thing that works really well is to provide all the context that's necessary to claude. Hook up the MCP connectors. Uh add the skills, right? Get that first working session to the goal yourself just by chatting through and seeing it work.
56:43 And then as soon as it works, you say now I want you to slchu this for whatever and just keep doing it. Um, that way you've got confidence that all the the, you know, there's no missing secrets, there's no missing connectors, you have all the data you need. And this lets you really walk away like once you have it set up where you're setting things up that are loopable, you're thinking about how I can loop this, how I can make it think about the what it means to be done or what what uh phases it needs to go through to
57:12 give me the confidence that things are going to be done in the way that I expect them to. Then that lets you schedule it. You can have it scheduled and kick off like when a new bug comes in, for example, a new GitHub issue. Uh it's going to kick all of this off and it's not going to bother you. You can be out on a walk in the forest and uh only be notified when a fix is ready or only be notified when it actually needs your help on something that it can't figure out on its own because it's new and novel.
57:37 And then you can go in and fix that loop so that it doesn't have to ask you about that again. And then you can move on. >> Yep. Uh, and so as I mentioned before, there's a whole interface built around schedule. You can list out what you've got going and uh you can see stuff. You can modify the schedules or tune it as needed and then find them when they're running and then delete them.
57:58 >> Y >> and you can mix and match these too. You can say like schedule this goal every week and it will run that goal uh and and do that. >> So now you can try it in the last two minutes that we have. We'll kind of uh pass pass pass through this, but um you can try scheduling it with like different loops like every minute run this report. Uh schedule every weekday at 9:00 a.m.
58:21 you're going to run this report and summarize the list that you get from there and then list out all of the scheduled tasks that you have. Uh and then finally, you can close the loop that we had from the beginning of this. So if you just say run my closing check-in, uh we should be able to hopefully see a visualization. >> If we got Yeah, if we got enough submissions, then we'll still see it.
58:39 And while that's working, um, >> that's there. Yeah. >> Yep. We'll look up and see. Um, so you keep all of this this, uh, content. It's all out on GitHub. These slides are out there as well. Um, we'll also be downstairs at the Work OS booth uh throughout the week. So, if you have any questions or want to chat about this, we live and breathe to talk about this stuff.
58:59 So, we would be happy to, but also we'd be happy to take uh we have one minute left for any final questions. Yeah, >> sure. >> Yeah, we'll get you in in the read me at the top of the read me are the links to the slides, the link to that glossery that we'll keep up and that you can use and ask any questions to. Um, and then Yeah. So I I'm sorry I'm kind of deaf.
59:28 I I think the question was in my company the PRs especially agent and enabled um PRs in the volume is really difficult because you might get a stack of 30 PRs sometimes and then sometimes the code's bad or there's just an immense amount of like 50,000 files and in a single pull request and you have to review it. We we think of that in in terms of like applying the same kind of concepts.
59:47 So like now it's like a it's a large system complexity problem and uh some of the tools that we've found success with are you know Devon GPile getting like a first warning analysis. It doesn't mean that if G reptile because they're wrong sometimes as we know why like if they say say it's five out of five ready to merge that's just your first kind of signal and then um so you can fail out all the ones that are twos and ones and tell people this is not ready for me to be bothered by and kick it kick it back to them
01:00:16 honestly and then for the stuff that is like fours and fives and looks okay there'll be comments and findings on it I'll tell agents in a loop babysit this PR and fix these until there's no more comments and then when I'm starting to get towards the end of that loop, I will actually start going and using my human brain for the first time in three years and open up the diff and say like okay what are those files I remember used to be sensitive let's look at configuration let's look at the off like let's look at the the
01:00:41 routes that are published right and then just kind of spot check that too um that's the way that is is currently working for us as well but this guy deals with a significantly higher volume of inbound PRs than I do >> yeah I'll I'll just say since we're out of time uh come talk to us at the booth I have I have a tool called DAB that I'm building that tells you a story about your PRs uh so that you can stay in the loop and and keep up on all the context of it.
01:01:03 Yeah, >> I'd be happy to talk to you about it. >> Yeah, great question. >> All right, folks. Thank you everyone. >> We're out of time. Thanks so much. Appreciate it. [applause]