← All transcripts

Ex-Google Engineer's Agentic Engineering Workflow Transcript, AI Summary & Key Points

Maddy Zhang · 10 days ago · Education · 09:37 · EN

Watch on YouTube

AI Summary

Agentic engineering works best when agents receive persistent codebase context, tasks are planned before implementation, and workflows are designed around reusable loops or graphs rather than repeated manual prompting. Memory files capture global and repository-specific conventions, while skills hold detailed instructions for occasional tasks. Plan mode exposes expensive design mistakes before code is written. Terminal workflows, Git worktrees, isolated parallel agents, and agent harnesses support concurrent development. Loop engineering lets an agent write, test, fix, and repeat until it reaches a defined finish line; graph engineering combines separate loops for research, implementation, and review. AI-generated code still requires risk-based review, real-user validation, manual testing, staging checks, and careful inspection of high-impact changes.

Key Points

  • Maddie’s agentic workflow starts by giving agents persistent context about the codebase, conventions, and recurring corrections.
  • Memory files such as CLAUDE.md or agents.md can be read automatically at the start of every session; global rules belong in a universal file, repository-specific rules belong in a local file, and occasional procedures belong in reusable skills.
  • Planning before implementation prevents an agent from committing to many small decisions about schemas, file layout, naming, helpers, and database assumptions before those decisions have been reviewed.
  • Plan mode lets an agent inspect the code and propose an approach without modifying files, allowing expensive questions about integration points and scalability to be resolved first.
  • Meeting decisions and requirements can flow into coding through Granola’s transcription, summaries, speaker tags, chat, integrations, and MCP connector.
  • Terminal-based workflows allow the same setup to be resumed later, including from a phone; Maddie uses T-Max and works with Claude Code, Codex, and DeepSeek’s open-source agent.
  • Git worktrees provide separate directories for multiple branches while sharing Git history, commits, and remotes without duplicating disk space.
  • Parallel agents can run in isolated contexts, with a lead agent assigning tasks and collecting results; cross-session messaging can transfer messages between Claude Code sessions.

AI in practice

Used for

Agents

  • Get an entire test suite passing. 2 held 07:06
  • Claude Code — Distribute coding work across multiple parallel sub-agents. 1 held 10:18
  • Carry out research, implementation, and review as coordinated stages. 1 held 07:23

Tools & resources

4 items

CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

Mentioned in
74 videos
Kind
AI
DNo. 2126
AIAINotes.us AI product

DeepSeek

deepseek.com

DeepSeek (深度求索) is an AI research project and model provider that develops and open-sources large foundation models (notably named DeepSeek-V4 and DeepSeek-R1) and provides hosted inference via a chat interface (“DeepSeek AI”) and an API. It publishes open-source model releases and offers a hosted API for text-based prompt inference; video reports state DeepSeek models can be exposed as external backends and routed into the Codex model picker via Codex Router.

Mentioned in
5 videos
Kind
AI
ONo. 0214
AIAINotes.us AI product

OpenAI Codex

openai.com

OpenAI Codex is an AI coding agent from OpenAI available as a command-line tool (Codex CLI) that helps developers produce software. It can be used alongside Gemini for adversarial audits of software requirements and implementation plans, listed as a supported coding-agent or model option in several projects, and its logs can be joined with task and test evidence. The Codex CLI can also receive and answer requests from the Penako canvas.

Mentioned in
55 videos
Kind
AI
TNo. 2226
AIAINotes.us Tool

tmux

Open source · tmux/tmux

tmux is an open-source terminal multiplexer that creates, accesses, and controls multiple terminals from a single screen. A session can be detached while its terminals continue running in the background, then reattached later, allowing terminal-based processes such as coding agents to survive disconnection. It runs on OpenBSD, FreeBSD, NetBSD, Linux, macOS, and Solaris, and is built from source with dependencies including libevent and ncurses.

Mentioned in
4 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Ex-Google Engineer's Agentic Engineering Workflow — Maddy Zhang (09:37). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Maddy Zhang. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 Hi friends, I'm Maddie and in this video I'm going to walk you through my entire agentic AI workflow as a senior software engineer. I've previously worked at Google Internet other big tech companies like Amazon, IBM, and Microsoft where I worked on launching Google health search features to hundreds of thousands of daily users, generating dozens of billions of dollars on revenue per year in Google search ads and building Roslin analyzers for Visual Studio, a tool used by over 25 million developers every month.

00:28 Over the last few years, I've been working extensively with AI and software engineering, and I've developed a workflow that helps me use AI agents effectively in my day-to-day engineering work. We'll start by looking at how I set up my agents and give them the right context, including memory files and skills. Then, we'll talk about planning before building, how I use agent harnesses and different tools, and how I work with parallel agents.

00:50 We'll also get into loop engineering and graph engineering. How to move from manually prompting an agent to designing workflows where agents can work through tasks more independently. And finally, we'll look at a code review workflow and how I decide how deeply I need to review AI generated code depending on the risk of the change. And more importantly, I'll explain why I do things this way so you can take the ideas and build your own aentic workflow.

01:14 Let's get into it. So everything actually starts before you give the agent a single real task. Here's what convinced me to set up my agent properly. A fresh agent knows nothing about your own codebase or conventions. Early on, I catch my agents doing something that looked reasonable and wasn't. So I cracked it and it'd be fixed fine. Then 2 days later in a new session, it did the exact same thing.

01:37 I then started pasting in the [music] necessary context which did work but I would need to do this over and over again with every single new session. So the fix here is to add memory files claw.md or agents.mmd depending on your tool that the agent reads automatically at the start of every session. [music] Everything you've learned about how code should look in this repo written down once so you never have to reexplain it.

02:02 And you should be adding continuously to these memory files. [music] So if your agent reaches for the wrong library, if it writes a component in a style that you would reject in code review, if it helpfully adds error handling no one [music] really needs, go ahead and write the underlying rule into the file. There's also a difference between context that is helpful for all your code versus context that is specific to one repo or issue.

02:21 So I keep the universal stuff in a global memory file and the repo specific rules in a local one. And not every useful workflow belongs in these default contexts. a database migration, a release process, a particular debugging procedure. Those might be extremely detailed but only relevant once in a while. For those, you can use skills, reusable instructions the agent can pull in when it encounters that specific kind of task.

02:45 Otherwise, just adding it to the memory file means that it's loaded every single session even for irrelevant tasks. Okay. So now my agent understands the project. Next, I plan extensively before any code gets written. So, why should you do this? Well, an agent that's already written the feature has quietly committed to a hundred small decisions. For example, schema decisions, file layout, naming, which helper it reused, what assumed about your database, and unwinding that all after the fact can be slower than if it had

03:17 never started. So, I use plan mode for anything substantial. Plan mode is the mode where the agent reads the code and proposes an approach without touching a single file yet. It comes back with a plan, and I treat it like a design dock. I'd push on the things that would be expensive to get wrong. Where does it touch existing code? How does it hold up as your users scale?

03:36 I'd say doing this for half an hour, getting the requirements clear can save you hours of reviewing and redoing at the end. But that only works if the requirements are already right in the first place. And that's where I keep on running into the same problem, which is that the requirements almost never start in the terminal. They start in a meeting, the sprint planning call, the design review, the instant call, where someone explains why the old approach broke.

03:58 And by the time I sit down to plan with the agent, half the context I need is a decision that got made out loud 3 hours ago, and I'm reconstructing it from memory. Which brings us to the sponsor of today's video, Granola. Granola is an AI notepad for people who live in backtoback meetings. So, I've been working on a group scheduling app, and I have calls with my project mates where I used to be taking my usual messy shorthand notes.

04:20 Granola now transcribes the call's audio right there on my computer in the background, so there's no bot joining the calls attendees list. Afterward, it combines the shorthand along with the transcript into a clean summary with the actual decisions and action items pulled out. And with speaker tags on, it even knows who said what. So I can ask its chat, for example, who raised the time zone edge case and what were they worried about?

04:43 And get the exact answer attributed instead of scrolling through a wall of transcripts. Grdola also has recipes, pre-built prompts, and integrations to connect to other apps like Slack, Atio, or Zapier for automated follow-up emails, CRM updates, and team notifications. My favorite part of Granola is its MCP connector, which means that it can plug straight into a coding agent.

05:02 The meeting context then flows directly into planning. The decisions from the call become the requirements for the code without anyone retyping them. So that issue I mentioned where the context lives in a conversation, but the code lives somewhere else is completely nullified. If you want to try out Granola for yourself, I've put the link in the description, and new users get their first month completely free.

05:25 Now, let's talk about the architecture and tools of Aentic Workflows. I work a lot directly in my terminal. I use T-Max, but any multipplexer will do the job. This means that I can close my laptop and pick the exact same setup backup later, even from my phone. Also, let's talk about agent harnesses. I bounce between a few, but mostly stick to Claude Code and Codeex personally, even though I've been trying out DeepSeek's open source one on this PC I recently built.

05:50 Nothing I show you today depends on one specific harness, but I'll be demoing stuff on cloud code in this video. And something that I found to be super useful is Git work trees, which allow you to check out multiple branches of the same repo simultaneously in separate directories. So instead of relying on git stash or making messy commits just to switch branches for a quick bug fix, you can create a parallel workspace that shares the exact same git history, [music] commits, and remotes without duplicating disk space.

06:15 Finally, let's talk about how to manage parallel agents. If you're using cloud code, you can spawn sub agents that each run in their own isolated context and a lead agent will hand them tasks and collect their results. [music] You manage the whole fleet in one place, seeing which are still running and which are done. And cross- session messaging lets Claw deliver a message from one of your cla code sessions to another.

06:41 Now let's talk about loop engineering and graph engineering. Loop engineering is the idea that instead of prompting the agent, correcting it, prompting it again, you stop handwriting each prompt and you design the loop the agent runs on its own. So you give it a task, a way to check its own work and a clear definition of done. Then it runs the cycle itself.

07:01 So, write, test, fix, repeat until it hits that finish line. So, instead of write me a function, I say something like get this whole test suite passing and don't stop until it's green and I walk away. It writes an attempt, runs a test, sees what fails, fixes it, runs them again, around and around without me. Once you're comfortable running loops, graph engineering is the next [music] step.

07:23 So, while a loop is one agent running its cycle, a graph is when you write several [music] loops together. So, for example, one researching, one writing, a separate one reviewing, each with its own clean context. And now, last but not least, we're at the moment where the agent is [music] done and you need to figure out if the diff is good enough to send out for review.

07:43 You might be tempted to just submit it as is. And for a tiny low-risk change, sure, that's probably fine. But if you do this for more complex changes, you'll break prod. Plus, 2036 data shows engineers who lean hardest on AI without staying engaged measurably lose ground on their own debugging skills. So, how do you decide how long to spend on code reading?

08:05 [music] I divide code features into three categories based on the impact it would cause if it went wrong. So, if the impact is [music] basically nothing, a copy tweak, a log line, a test only change, I read the summary, glance at the diff, and move on. If the impact is a feature gets ugly but nothing catches fire, I skim the lines of code and make the agent prove it works.

08:24 That means it exercises the feature the way a real user actually would end to end and outputs a screenshot, a short screen recording the actual output. I'll also often go in and manually deploy and test out the feature myself [music] just to double check. And finally, if the impact is something like money will move, data will leak or someone can't log in, especially [music] anything touching payments off or permissions or the shape of the DB.

08:46 I read every single line carefully and slowly and I make sure to manually deploy and test out the feature. I'll also flagguard it so that after the code is submitted, I can still make sure it works in staging before deploying to prod. So in conclusion, that's the identic [music] workflow I use from start to finish. Set up and give your agent the proper context.

09:05 Plan before anything is built, set up the architecture [music] and tools, design loops or graphs, and then review the code that you generate. I want to mention that none of these concepts are unique to a specific tool or harness. Models and agents will change every few months, but the usefulness of these habits will persist. [music] And that's all I have for you in this video.

09:24 If this gave you a clear picture of what agentic engineering actually looks like day-to-day, [music] hit that like button, hype the video, and subscribe. Thanks for watching, and I'll see you in the next one.