← All transcripts

Your AI Agent Is Confidently Wrong About Production — Willem Pienaar, Cleric Transcript, AI Summary & Key Points

AI Engineer · 4 days ago · Science & Technology · 14:51 · EN

Watch on YouTube

AI Summary

Production debugging is harder for AI agents than coding because production lacks tests and linting, spans changing and unclear system boundaries, contains transient failures, and provides diffuse feedback. Agents tend to produce confident answers when effective debugging requires uncertainty, multiple theories, and evidence-based elimination. Common failures include mistaking symptoms for causes, lacking a baseline for normal behavior, losing information through sub-agent summaries, and anchoring on past incidents. Grounding techniques include generating parallel theories, computing baselines, using dependency graphs and temporal relationships, and propagating evidence. Verifying whether a fix actually restores services, metrics, logs, and state to normal enables confidence calibration and can bring production signals into development to predict failures before deployment.

Key Points

  • Production has no focused verification loop like tests and linting in development or CI, so agents receive diffuse signals from logs, noise, deployments, and changing system state.
  • Production is difficult for agents because its boundaries are unclear across teams and clusters, its state changes continuously, and many failures are transient and unreproducible.
  • Large language models are optimized to be decisive and confident, while production debugging requires doubt, uncertainty, multiple theories, and elimination. OpenAI system cards or reports showed confidence and accuracy deviating during RLHF.
  • Mistaking the symptom for the cause — agents often inspect the service showing the alert, find a nearby error log, and stop without searching deeper through the stack.
  • Not knowing what is normal — without a production baseline, agents can treat common recurring errors as anomalous and stop searching.
  • The abstraction problem — sub-agents processing large production environments may omit raw findings, leaving the main agent to decide from partial information and become overconfident.
  • Anchoring on past incidents — agents may assume a current problem matches a previous incident instead of investigating whether the underlying conditions are different.
  • Generate many theories — parallel hypotheses provide breadth and diversity, and help because agents are better at relative ranking than absolute scoring.

AI in practice

Used for

Agents

  • Diagnose production incidents, assess proposed fixes, and determine whether the system returned to normal. 2 held 11:30

Tools & resources

2 items

CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

TypeScript
Stars
★ 149,337
Forks
25,438
CNo. 5046
AIAINotes.us AI product

Cleric

cleric.ai/

Cleric is an AI agent service for production operations and incident response. It follows software changes after merge, checks their effects under real production traffic, maps services and dependencies, investigates alerts and regressions using production signals and prior investigations, and distinguishes confirmed issues from false positives. For confirmed problems, Cleric provides the cause and supporting evidence, can prepare a fix pull request for review, and monitors the deployed fix to verify that the issue has recovered. It is read-only by default, logs actions and investigations for auditing, encrypts data, and states that customer data is not used for training.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Your AI Agent Is Confidently Wrong About Production — Willem Pienaar, Cleric — AI Engineer (14:51). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:13 Hello everybody. Uh we're going to get started. Um my name is Wllm. I am the co-founder and uh CTO of Cleric. So we build agents that keep your systems online. Today we're going to be talking about how things fail in production. I think everyone here is really familiar with shipping code, but writing code is not the same as like running it in production.

00:35 Um, so let's dive into some of the ways things go wrong. So I think we're all familiar with uh like prod failures or issues in the production environment. And I think the modus that people get into today is like applying agents to solving these problems. You're running claude code to diagnose some alert. Um unlike with a coding use case, what we find is in a lot of cases, there's a lot of back and forth.

01:00 So if you deploy an application, you get maybe like a 500 um errors are spiking. Um the agent comes back very confidently with an answer. It says um you know there's a memory leak that you just deployed and it requires you as a human still to go and look and see did this actually happen. And often this results in this back and forth where you have to second guess the agent because unlike in the coding case, the production environment is a lot harder.

01:26 And so you go back and forth, it claims it's the readiness probe. You say, well, you look at the the event and you see that it's not the case. And eventually you realize or through iteration with the agent what the problem actually was. So this is a frustrating experience. Even though the agent's doing a lot of work for you, you still are in the hook on on the hook to iterate with it.

01:48 So we're going to talk about how do we improve this? How do we get this to a point where there's higher agency and autonomy with agents. So we'll talk about what are the failure classes that we see and how we as cleric have developed some approaches that you can apply on your own to improving the accuracy of your agent in the production environment as with coding.

02:04 Um and then how you measure that and and roll that out. So one of the key differences with the production environment is that there's no verification. So if you take dev or if you take even CI, there's tests, uh there's linting, there's very focused signal that comes back from all the changes that your agent is making that makes iterations very fast.

02:25 Um and the responses that the agent gives you are very accurate in most cases. Prod on the other hand is diffuse. You have a little bit of logs here. You have noise over there. People are deploying things. there's no focused signal that's coming back to you. So the real question is what do you have on the prod side of things where the the time you can sync there could be in the order of hours.

02:47 Um like you need some kind of signal coming back to your agent. So that's really what this talk is about. How do you build that signal? Let's talk about a little bit about why prod and large language models are really hard compared to code. In the case of code, you have a very bounded environment. It's a single codebase. There's a line around that. it's the repo.

03:07 But unlike code, a production environment is often an amorphous uh environment where the boundaries aren't clear. Sure, you have like one GCP or AWS account, but there are multiple teams, there are multiple clusters, and for an agent, it's unclear how wide or narrow to go. Additionally, prod is also moving. So, there's state that's constantly changing.

03:28 Um, people are deploying things. Systems are not sitting still unlike a codebase that's static. And I think one of the most pernicious parts is prod is also transient. A lot of the failures you get can't be reproduced in a codebase. If the agent runs into a problem, it can reset and try again. But in prod, sometimes a failure is only seen for a specific moment in time.

03:50 But that's not the only challenge. The other challenge is that models are not really trained for this. In fact, large language models are trained to almost be the opposite for what you need for a production debugger. So what you need for production is doubt and uncertainty and trying many different theories and eliminating those theories. But these large language models are really trained to be decisive uh confident and to give you an answer as quickly as possible.

04:14 And in fact like one of the uh system cards or the reports that OpenAI released showed that the confidence and accuracy deviated in the RLHF process and this is partly to do with the the post training being optimized for preferences of humans. So it's a part of the the models and it's something that you have to counteract and so how you build the counteraction to that is a part of the challenge here.

04:40 So let's talk about some of the failures that we see in production. We work with many different uh companies and enterprises. Uh we see how agents fail in debugging uh production systems and and the failures they um uh are trying to diagnose. The most common failure we see and this is for claude code for our product pretty much across the board is uh mistaking a cause and a symptom.

05:03 Often when you have an alert, the agent will run to that service where the fire or the smoke is and it'll find an error log right there and it'll come back very confidently and say this is it. I found the root cause instead of going deeper into the stack. This is a very hard one to debug because these agents or these models are very good at uh finding a needle in a hstack, but they don't know the breadth of the search space that they should be looking at.

05:31 The second one we see a lot is not knowing what's normal and what is anomalous. So if you don't have a baseline for production, what happens is that the first thing that they see, they will think it's anomalous and they won't won't try and continue to search for something else. The third one which is a new class of problem is the abstraction problem.

05:48 If you look at most modern production environments, they have a lot of information. So you need classes of agents to process that. But these agents now don't bubble up all of the uh raw findings to the master agent or the main agent that's making a decision. And that leads to overconfidence in the main agent making a decision based on partial information.

06:12 And the final one that I think is also very common now with memory systems and context being so important in diagnosing issues is using past context or state as part of the investigation often leads to agents anchoring on that information. and for example saying okay last week we had this issue this is probably the same thing here's your diagnosis so this is really hard because the agent is very eager to please and give you an answer and you kind of have to destabilize it a little bit to have it go further and deeper

06:40 so let's talk about some of the techniques we've applied and and how they've been effective for us at least so the thing that you need to do is replicate what works in coding in coding you have signal that comes back very quickly it's critical signal So the agent tries a code change and it comes back and says this didn't work, it didn't compile or here's an error message.

07:00 So you want something that the agent can ground uh ground on um as part of an investigation or also at the end and that's basically going to give you like a confidence score that the agent has as it's executing. So what we do is we generate many different theories. This is a very effective like primary thing that we do to give the agent uh kind of like a breadth a diversity of ideas that it's exploring simultaneously.

07:24 This is important for multiple reasons. One is the diver diversity but also secondly agents are really good at relative ranking and it's they're terrible at absolute scoring um of their theories or ideas. So generating multiple ideas for why a root why uh what the cause of a failure is is very important. Secondly, what we do is we compute a baseline.

07:47 This is also extremely important, especially in large production environments. Most prod environments are in a permanent state of kind of disarray and chaos. Engineers know to ignore certain error logs. They need know to some metrics are not even used, but an agent doesn't know that. An LLM doesn't know that. So what we do like an example of this is if we see logs or if we're uh presenting logs to an agent we'll cluster the logs and tell the agent which logs are um like reoccurring every day and very common.

08:18 So we give it the statistics on those logs not just the raw logs and that makes it a lot more uh effective but this applies as well to metrics or object state or many other things. What we also do is we use a topology or we build dependency graph of your production environment. This is very important because the agent will often claim there's some causal connection between two services.

08:40 So for example, it'll say like an all service is connected to a cache. But if you have a graph or if if you know the dependencies between services, you can output uh all the connections that exist and then have another agent critique that. And this is a very effective way to introduce doubt into the answer and tell the agent to go further if it doesn't uh map to an existing connection.

09:06 Another one that we do is or another technique we apply is using time. So you know that the cause is going to happen before the symptom temporally. But a lot of people don't even apply temporal um assessment into their uh root causing. Additionally, you also know that in most cases t uh temporally things will be closer together in time versus further apart.

09:28 So you can add a relevant score to any findings that you have um that are closer to um the alerting time stamp. And then the most obvious one is just propagating sources. So if you have sub agents or agents handing context over to each other, have them build a case. Don't just tell them to give you the answer. And this is also by the way what SRRES do.

09:50 They don't just say you know prod is up. They say I looked at this metric and it was at this point um at this point in time. They give hard evidence. So if you don't propagate the evidence um this loses a lot of the fidelity and causes hallucination and overconfidence. Um and then finally this is these are some of the results that we've had um in our simulations.

10:13 Grounding the agent has been extremely effective. So what we find is that for localized incidents, incidents where um effectively alerting service and the the root cause are in one area, there's a very small delta. It's almost the same as clawed code. But what we found is that for distributed problems where the cause and the symptom are not in the same service, uh claude code is not as effective, not even nearly as effective.

10:40 So of course you can't use these numbers as gospel because this is a simulation environment. Every company is different. If you're optimizing your stack to just be like one monolith, you'll have a good time. But most companies can't do that. So, it's much harder and the the problem space is much more diverse. The there's one downside to doing what we're doing and that is cost.

10:59 Our runs are very expensive compared to just running clawed code. So, if you're running claude code, it's very surgical. It's one thing, but you have to guide it and steer it or any agent, right? But if you're running five or six theories concurrently, um, you know, it's going to cost you a lot more money. you're going to burn tokens, but it's also going to keep the human out of the loop much longer.

11:21 But how do we know this actually worked? I mean, these are all just vibes. We're just adding more context into the agent. So, we actually need to do a few more things to know what actually happened. So, the way we look at the um production incident response flow with agents is there's a detection, a diagnosis, and a fix. And most folks use agents for diagnosis.

11:40 But what we do is we also look at verification and the outcome. So we triangulate across many different signals what happened when you apply a fix. So if we give you a suggestion, here's a diagnosis, we monitor you as an engineer and we see if you apply the change like actually merge code that maps to the diagnosis we gave you and then we see if that solved the problem.

12:03 So we looked for a return to normal um in the service the metrics the logs and all of the state that um should be there in the baseline. Now of course if you just mute an alert you're also going to see a return to normal and so we look at multiple signals and ensure that it is actually lining up with our expectations. Now this is very important because you know why even have all these machinery to track this outcome.

12:26 It's important because you want to calibrate your um evaluation at the end of the day. So what we had earlier earlier was grounding that we added into our agent basically a critic and what you see is that the blue line is the critic and the red line is when you don't have grounding and so when you verify the outcome what actually happened at the end of the day did the the service return to normal.

12:50 You can see that with grounding if you the 45° line is the um confidence versus accuracy. So you want to be on the 45 degree line. What happens if you aren't grounded is you become extremely overconfident. So if the agent is just using vibes effectively, it'll always give you a very confident answer. And so the effect of this is if you have the grounding and you track verif the the outcomes um you have a signal that's going to come back from your agent to whatever is going to apply the fix that you can trust.

13:24 So if the agent says I'm 90% sure that this is the answer you can trust that signal and so at the end of the day you can complete this loop um taking the signal all the way from prod back to your agent. And so this has been our focus for three years. We've been building effectively ways to extract signal from the production environment and bring it back to our agents.

13:49 So this is both for oper systems that are operational already as well as during your development loop. And where we're going next is we're bringing that signal to the developer environment. Instead of waiting for an application to be deployed before it fails or or fail after it's deployed, we bring the model, this baseline model, the system graph, the uh the ontology all the way to your dev environment so you can predict failure before you deploy or make a change.

14:16 Um yeah, so this is the direction we're going and I hope that was useful. This is some of the learnings that we've had along the way. If anybody has questions, feel free to shoot. But yep, that's a talk. >> [applause] [music] [music]