← All transcripts

From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI Transcript, AI Summary & Key Points

AI Engineer · 3 days ago · Science & Technology · 01:51:27 · EN

Watch on YouTube

AI Summary

Reliable AI agents require traces, evaluators, datasets, and controlled iteration rather than subjective spot-checking. Traces expose every agent turn, tool call, model invocation, input, output, timing, token count, and metadata. Code evaluators provide fast deterministic checks, LLM judges handle semantic criteria, and human annotations calibrate the judges. A financial analysis agent produced 13 reports: a correctness evaluator scored all 13 as zero because it lacked the live web research context and only had information through January, while a faithfulness evaluator supplied with the agent's research context found six faithful and seven unfaithful reports. Reading traces exposed failed file writes, missing report bodies, excessive web searches, loops, and unsupported financial information. Success criteria should be defined before automation, evaluators should focus on one dimension each, and failures should become datasets for prompt revisions and controlled experiments. Production quality requires online evaluations and monitoring against new traffic.

Key Points

  • Evals test AI outputs while traces log agent execution, including agent calls, tool calls, LLM invocations, inputs, outputs, timing, token counts, and metadata.
  • Spot-checking a few queries creates a vibes-based workflow that misses untested inputs, does not scale through human review, cannot reliably catch regressions, and does not run in CI.
  • Code evaluators are deterministic, fast, inexpensive, reproducible checks for conditions such as valid JSON, length limits, forbidden phrases, required fields, or ticker presence.
  • LLM judges compare outputs with a rubric for semantic properties such as correctness, faithfulness, tone, and completeness, but they cost money, are nondeterministic, and require calibration against human judgment.
  • Human evaluation provides domain expertise but is slow and expensive; human annotators can miss up to 50% of defects because of fatigue.
  • Agent evaluation must inspect intermediate decisions and tool calls, including tool selection, arguments, stopping behavior, loops, and handoffs, because errors can cascade across multiple steps.
  • Capability evals measure whether an agent can perform a task and are expected to fail while the capability improves; regression evals cover behavior the agent should pass consistently after the capability has matured.
  • The financial analyst uses two turns: a research turn that searches the web for current ticker information and a writing turn that compiles the research into a report.

AI in practice

Used for

Agents

  • Financial analysis chatbot — Research a stock ticker and focus area on the live web, then write a current financial report. 2 held
  • Claude Code — Improve the financial analyst's prompts using failures found by evaluation. 2 held 01:32:49

Business ideas

A user provides a stock ticker and a focus area. The agent searches the web for current financial information, gathers research, and writes a report containing financial performance, growth drivers, risk factors, analyst price targets, conclusions, and actionable recommendations.

For
People who need current, ticker-specific financial analysis and recommendations.
Solves
Manually gathering current financial information and turning it into a readable, actionable report is time-consuming; a probabilistic agent can perform the research and reporting workflow.
  • Tesla financial performance and growth outlook: the agent searched the web and produced a report with a financial performance table, growth drivers, risk factors, analyst price targets, and a conclusion, but it also attempted an unnecessary file write.
🔒  Build steps and tools for 1 idea. Unlock

Tools & resources

7 items

ANo. 0159
AIAINotes.us AI product

Anthropic API

anthropic.com

Anthropic API is the cloud API provided by Anthropic, the AI research company, offering programmatic access to its Claude family of large language models for text generation, chat, and reasoning. The API is intended for developers to integrate Anthropic's models into applications and services.

Mentioned in
4 videos
Kind
AI
ANo. 5070
AIAINotes.us AI product

Arize AX

arize.com/products/ax/

Arize AX is an AI engineering platform from Arize AI for observing and evaluating AI agents, including voice agents. It provides a unified session and trace view that combines audio, transcripts, latency metrics, tool calls, and evaluation results, allowing failures such as dead air, interruptions, transcription errors, and incorrect tool actions to be investigated together. The platform supports tracing real-time audio with OpenInference semantic conventions, audio-native evaluations for factors such as sentiment, latency, interruptions, and task success, and agent experiments that test fixes against failed traces. It is available through the Arize AX product interface and a terminal-based setup.

Mentioned in
3 videos
Kind
AI
ANo. 5081
AIAINotes.us AI product

Arize Skills

Open source · Arize-ai/arize-skills

Arize Skills is a set of installable skills for AI coding agents, maintained by Arize, that guides agents through adding observability, creating datasets, running experiments and evaluations, and optimizing prompts for LLM applications. The skills use the ax CLI to interact with Arize and include workflows for tracing and debugging applications, OpenInference-based instrumentation, dataset management, experiments, LLM-as-judge evaluators, annotations, prompt optimization, provider credentials, and production monitoring. They can be installed into agents such as Cursor, Claude Code, Codex, GitHub Copilot, and Windsurf through npx, the repository's installers, the ax CLI, or a Claude Code plugin. The repository is licensed under MIT.

Mentioned in
1 video
Kind
AI
CNo. 4923
AIAINotes.us AI product

Claude Agent SDK

In the AINotes directory

An SDK for building AI agents that can operate beyond a standard chatbot. The video describes it as providing filesystem access, tool calling, code writing, and debugging capabilities; AssemblyAI used it to build the Joey support agent with local Markdown documentation, embeddings, agentic file search, and a large CLAUDE.md configuration file.

Mentioned in
2 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

TypeScript
Stars
★ 149,337
Forks
25,438
ONo. 5071
AIAINotes.us AI product

OpenInference

Open source · Arize-ai/openinference

OpenInference is an open-source set of OpenTelemetry-complementary semantic conventions, instrumentation libraries, and plugins for tracing AI applications. Its specification models LLM invocations and surrounding context such as vector-store retrieval and external tool use in a transport- and file-format-agnostic schema, while its instrumentations normalize traces from model, agent, and application frameworks across Python, JavaScript, Java, and Go. The project also defines conventions for real-time audio and voice-agent traces, allowing audio, transcripts, and tool calls to be represented in a queryable session trace. It is natively supported by Arize Phoenix and Arize AX and can send data to any OpenTelemetry-compatible backend.

Mentioned in
3 videos
Kind
AI
ONo. 0760
AIAINotes.us Tool

OpenTelemetry

Open source · open-telemetry

OpenTelemetry is an open-source, vendor-neutral observability framework for cloud-native software. It provides APIs, SDKs, automatic-instrumentation agents, and collector services for capturing distributed traces, metrics, logs, and contextual metadata such as baggage. Its context-propagation mechanism correlates telemetry across service boundaries and carries trace identity through an application’s request path, including standard HTTP communication. The OpenTelemetry Collector receives, processes, filters, and routes telemetry to observability backends such as Jaeger, Prometheus, commercial systems, or custom solutions. It supports native SDKs for more than 12 languages and is hosted as a graduated project by the Cloud Native Computing Foundation.

Mentioned in
6 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 — Laurie Voss, Arize AI — AI Engineer (01:51:27). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Hello everyone. Uh I have a packed room for a two-hour workshop which is amazing. Welcome to AI Engineer. I know no one's officially welcomed you but uh welcome to AI Engineer. Uh this venue is amazing. This crowd is really quite surprising and I have uh a ton of content for you today. Uh my name is Lori Voss. I'm head of developer relations at AriseAI.

00:37 Uh some of you may remember me from when I used to co-found npm Inc. Uh so you may remember me from like JavaScript days. Uh but these days I talk about AI uh and how to test it uh and how to make it work. So what are we going to be covering today? Uh we've got a really good stretch of time together. Uh so we're going to cover a lot of ground and I'm going to start with fundamentals.

01:02 That's why this is called a 101 session. Uh we're going to start with what eval are, why you need them, uh and why agents make evaluation harder than simple LLM applications. Uh then we're going to set up tracing in Arise AX, uh which is how you capture the raw data you need to run evals in the first place. Uh, we're going to build a simple AI agent uh with the claw agent SDK.

01:26 We're going to run it and we're going to look at uh the traces it produces. Can I get a show of hands if you've already built an agent? All right, great. I'm going to spend no time at all on how the agent works because I assume everybody knows how an agent works and you're here to learn how to evaluate them. Uh once we've got some data, uh we're going to do something that a lot of tutorials skip, which is we're going to actually look at the data.

01:50 Uh we're going to read our traces. We're going to categorize what went wrong. Uh and we're going to figure out what to measure before we write a single eval. Then we're going to write three kinds of evals. There are code evals, which are simple deterministic checks, uh very similar to unit tests. Then we're going to use built-in LLM evals uh like faithfulness uh where a second LLM judges the uh output of your first one.

02:16 Uh and we're going to do a custom eval uh LLM as a judge from scratch. Uh and we're also going to test whether our judges are judging correctly, a process called meta evaluation. Then we're going to finish with data sets and experiments, which is how you iterate on your agent uh and actually measure whether it's getting better. uh including we're going to use cloud code to automatically improve our app in response to feedback from our evals.

02:43 So, first things first, uh get your notebook. Uh I messed up earlier and didn't have permissions set correctly. So, if you scanned this earlier and it didn't work, it works now. Uh grab this notebook and make a copy for yourself. That is where we're going to be living all day. And I'm going to wait for that to happen because otherwise I'm going to have to keep swapping back to this slide.

03:11 All right, most of those are down. No, no, I'm still seeing people taking pictures. All right, cool. We're also going to need a certain amount of AI infrastructure. I'm going to be using anthropic and cloud code today. Uh, Arise AX works with OpenAI. It works with every other uh, lab provider, but I had to pick one in order to do a workshop about it.

03:36 So, I've used Claude. So, this is going, this notebook is going to expect you to have uh, an anthropic API key, which you can get from console.anthropic.com. If not, you can, if you are very clever, do a fast swap to OpenAI. Um, you're also going to need an Arise account. Go to arise.com and start a free trial. Um, from there you will need an Arise Space ID and an Arise API key, both of which you can get from the settings page.

04:02 Uh, my lovely assistant Dat is here. Uh, that's him in the corner. If you run into any problems getting your API keys or anything like that, uh, just flag him down and he will run over to you and help you figure it out, which is, uh, incredibly nice of him. Thank you, Dad. Um, yes. So, uh, who who had all that stuff already? Anybody already got an Aris X account?

04:38 All right, I'm going to let you hang out for this. I'm going to let you set that one up for a little while. The way you do it is you go to arise.com and you hit Oops, it's sent me to the docs because I always live in the docs. You go to arise.com. Thank you. Uh and you hit the big pink get started button and then you sign up from here. Do not accidentally sign up for Phoenix.

05:18 Phoenix is our lovely open- source project which also has a hosted service underneath Phoenix. Uh sometimes when you Google Arise, you get sent to Phoenix instead. Don't sign up for Phoenix. You're doing AX today. All right. While you are signing up, uh I'm going to say a little bit more bit more about what AX is. It is an observability and evaluations platform.

05:42 It captures traces from your AI application. It stores and runs your evaluations and it watches your app in production and alerts you when something goes wrong. Uh by default, we host it for you. That is why we're using it today because you don't have to like run some software to make it run. Uh so there's no infrastructure to manage. Uh though if you are from an enterprise that is super precious about your data and you're like no, we need to run it ourselves, you can absolutely do that if you want to.

06:09 Um, and like I said, uh, if you've used Arise Phoenix, which is our open source tool, a lot of this will feel familiar. Uh, they have the same core ideas. One is, uh, built for open source accessibility, and the other one is built for, uh, enterprise with production and teams and things like that. So, how are we doing on signing up for keys right now?

06:35 >> Wi-Fi sucks. >> The Wi-Fi sucks. That is predictable really, isn't it? All right. Is anybody like desperately in need of assistance right now? >> All right. Okay. Well, the advice is to switch to your phones as a hot spot. I apologize for the Wi-Fi. I'm really worried about my 2PM workshop where I have a Git repo that I know is 200 megabytes large and I'm like, no one is going to get the repo because it's too big.

07:29 >> Yeah, start cloning that repo now. Anyway, uh let's talk about these these fundamentals. There's a lot of theory in this workshop which gives you time to uh catch up. Um here's the mental model. Evals are testing for AI and traces are logs for EI AI. Why we had to come up with two completely new names for those things? I don't really know. Uh you write tests for your code.

07:52 Given this input, I expect this output. Uh eval do the same thing, but for AI outputs where the outputs are never exactly the same twice. Just as logs record what your server did at runtime, uh, traces record what your AI did. So, every agent call, every tool call, every LLM invocation with the inputs and outputs at each step. The building blocks of a trace are called spans.

08:17 Uh, in this little screenshot, you can see a trace in Arise AX. It's made up of a bunch of spans. It's got this tree structure. Uh, and each span represents one step in the execution. An LLM call is a span. A tool call is a span. And a full agent turn is a span that contains other spans inside of it. Each span records its input, its output, uh timing, token counts, and a whole bunch of metadata.

08:39 Uh everything that you need to understand what happened at that step. Traces tell you what happened. Evals tell you whether it was any good. That is the model for mental model for today. This is not rocket science. But let me talk about why we even need evals. Uh and that is the vibes problem. Uh AI features usually get shipped by building an AI feature, running a few queries, saying does this look right?

09:07 Yeah, it looks good. And then you ship it. That is by far the most popular method of shipping AI applications today. Uh and the problem is that then it fails on inputs that you didn't test. Uh and the thing is it probably looks good on your test queries, but the problem is that you ran it three times. Three times is not a test suite. Uh the usual fix for uh flaky software is unit tests and they don't work here because the output is non-deterministic.

09:33 Uh the same prompt produces different text every run. So two different responses might both be correct. So there's no expected string that you can assert against. So teams fall back on human review. They watch it run. It looks fine and they ship it. Uh that doesn't scale. It doesn't catch regressions. And most importantly, it doesn't run in CI. Here's some things you can't do without evals.

09:55 Uh if you change your system prompt to fix a tone issue, then the tone would get better. Uh but now your uh bot is hallucinating product features. Without eval catch that uh until a user reports it. With a faithfulness eval, you would have seen the spike before it shipped. Uh and every prompt potentially affects every kind of input that users send. So you can be changing the prompt for one thing and have a completely unexpected effect at some totally different part of your application because the prompt is just a huge

10:24 block of text that you're sending to your AI. Uh you can't manually test every combination. So without eval you're playing whack-a-ole. You fix one problem and you break something somewhere else. Evals instead give you a number that you can track and compare and act on. You also can't switch models without eval. This matters more and more because this field moves so fast.

10:45 Uh new models drop every few months. Uh without eval switching models means weeks of manual testing. Is it better? Is it worse? Did something get did something break? You don't know. With evals, you run the suite. You compare the scores uh and you know within hours. And that is the difference between flying by the seat of your pants and flying by instruments.

11:08 And this is not theoretical. uh real teams, Dcript, Bolt, Anthropic's own clawed code all follow the same arc of shipping on vibes and then discovering that vibes don't scale uh and then uh building evals and that is the same arc that we're going to follow in today's workshop. As I mentioned earlier, there are two broad types of evals. Code evals are deterministic functions.

11:31 They run in milliseconds and they cost nothing and they give you an immediate unambiguous answer like did the output parses JSON? Is the output less than 500 tokens? Did it mention the thing that I asked about? The big advantage of Kodi valves is that they are fast and cheap and totally reproducible. They get the same they get the same input and they give the same answer every time.

11:51 That makes them very similar to old style unit tests. The downside is that like old style unit tests, they can be brittle. Uh we have some strategies that can help avoid that and I'm going to be covering those later. um because two perfectly correct answers might be worded completely differently and you have to adjust your expectations of unit tests to be able to handle that.

12:14 The second type of eval is an LLM as a judge eval. These use a second LLM to grade your outputs against a rubric. Rubric being yet another word that we made up uh that just means rules. It is your set of rules that you put in your prompt about what counts as good. Uh LLM evals are much more flexible. uh they can handle the kinds of questions that Kodi valves can't answer.

12:34 Uh is this response uh factually accurate? Did it stay faithful to the source material? Is the tone right? Uh for a customer service context, the strength of LLM judges uh is that they understand meaning uh not just the strings that are in your output. But LLM judges have trade-offs. The weakness is that they cost money to run and they are nondeterministic themselves.

12:57 Uh so they can be wrong which means that you need to calibrate them against human judgment and we're going to do that later today. Um there's also a third type of eval which is human evaluation. Uh a domain expert can read your output and grade it. Uh and that is the gold standard for quality. Uh but it is slow and expensive and humans get tired which is why we can't use it all the time.

13:20 Uh fun fact that's not very fun. Human annotators miss up to 50% of defects due to fatigue. So they're not perfect either. So when you build an LLM as a judge and you're like, well, it doesn't always catch my errors. Remember that if you were judging with a human, the human would also not always catch your errors. Uh but these are complimentary approaches.

13:40 Most real applications are going to use a a collection of Kodi valves, LLM as a judge, and human evaluation. So when do you use which? Use Kodi valves when the answer is deterministic. So, format validation, length limits, forbidden phrases, uh required fields, use LLM judges when you need semantic understanding, a correctness, eval accurately, a faithfulness eval asks, did it stick to the source documents without hallucinating?

14:09 Uh, and you should keep humans in the loop for failure modes that you haven't seen before. Uh, and to verify that your LLM judges are actually judging correctly. LLM judges can be wrong. uh judging them against human judgment is how you know whether or not you can trust them. All of that is true of any application that uses LLMs. Uh but agents make things even harder in some ways that are uh worth calling out.

14:34 You can think of it as a ladder of complexity. A single LM call is relatively simple. You get an input, you get an output. When you're talking about an agent, an agent is making a series of tool calls and a series of decisions, which means that at every single step in the uh chain, it is the chances that it has made a mistake are getting higher and so the chances that it will go off the rails uh are going to get uh greater.

14:57 So you have to judge the intermediate steps as well. You have to uh check whether your agent is using the right tool. Did it pass the right arguments? Did it know when to stop? getting stuck in loops is a real failure case for agents. Uh and then multi- aent systems if you are building those add yet another level of complexity. They add handoffs. One agent decides to pass work to another.

15:20 Did the triage agent route correctly? Did the specialist agent handle the the handoff gracefully. Each layer adds new ways that things go wrong. Uh and non-determinism means that you need to have more evals because errors can cascade. Uh here's what an example of an cascading failure looks like. If you're you ask a query about Tesla, meaning the car company, your agent searches the web, finds a whole bunch of information about an 18th century inventor, passes that back to your agent, which writes a beautiful, well-

15:52 formatted report about an 18th century inventor, and passes that to your boss. Nothing that it did there was wrong. It you told it to search the web for Tesla. It searched the web for Tesla. It got information about Tesla, and it wrote a report about Tesla. where where did it go wrong? Uh this is worse than an obvious failure because you have to be reading your inputs and outputs for that to work.

16:13 You can't there was no unit test that would have caught that. Uh but agents can also do the opposite. They can get things right in a way that your tests weren't expecting and get graded wrong as a result. This happened to Anthropic when they ran a benchmark called Tao Squared Bench uh which simulates multi-turn customer to service tasks. uh they asked the agent to book flights uh and asked it to do something that they thought was impossible which was reschedule an economycl class flight.

16:44 Uh but it turned out that the policies that they gave the agent allowed the agent uh to first upgrade the flight to first class. First class flights can be rescheduled which uh economy flights cannot. So the agent found a way to reschedule an economy flight uh and failed the test because it was supposed to be not able to do that and it was able to do that.

17:05 Uh so this is an example of how agents being flexible agents finding things that you didn't think of uh needs to be accounted for in your tests. Uh there is a difference between creatively correct and wrong and we're going to come back to that when we talk about grading your agents later or rather grading your evals. Um there's also a second way of categorizing your eval into two categories uh which is uh capability evals and regression evals.

17:32 A capability eval asks uh can my agent do this thing at all? Um capability evals are expected to mostly fail. Your agent is expected to get a very low score on them. Uh because they give you a hill to climb. It gives you something for your agent to try and get better at. Uh once you've climbed that hill, your capability eval turns into a regression eval.

17:53 A regression eval is an eval that you expect your agent to pass completely or nearly completely every single time. So as your agent matures, you'll be giving it capability eval after capability eval and slowly turning them into a suite of regression evals that you run all the time to make sure that uh previous behavior has not regressed. This is what an eval result looks like.

18:20 Uh every eval produces a score and a label. In this case, it's a numeric score just often just zero or one. Binary scores are very easy for uh LLMs to do. Uh and it comes with a human readable label like correct or incorrect or valid or invalid. Um LLM judges add a third thing which is the explanation. Um code eval produce explanations because they are just code.

18:41 Um but LLM, an LLM as a judge will say not just that something is incorrect, but it will say why it decided that something is correct. And this is incredibly valuable uh because if you are living in a world of coding agents, you can take a whole bunch of explanations from a bunch of evals and pass them back to your coding agent and say here is all the reasons that you messed up in the last test.

19:05 What could you do that would make you better at this? And your coding agent will just go cool thank you for the feedback. Improve your application uh and it will get better at it next time. This is something I'm going to show you uh in today's workshop. It is how you turn a failing eval into a prompt improvement. This is a real judge explanation. Um, in this example, it it's a travel planning agent.

19:28 Um, and you're running a correctness eval. The judge doesn't just say incorrect. It tells you exactly what's missing. You asked for a budget flight. It didn't give a budget breakdown. It says the agent gave destin destination info and recommendations. Uh, but the user asked about budget travel and there were no cost estimates. That explanation makes the eval actionable uh because you now have a concrete failure.

19:49 You know what to fix in the prompt. Uh and the explanation is what makes evals into a useful debugging tool and not just a scoreboard. Uh and if you're seeing the same explanation across 50 different traces, uh you know you have a systematic problem and not a one-off edge case, which makes it the first thing you should fix. So you have to take your eval explanations uh and categorize them into types of failures uh and count up the categories.

20:13 This is called coding. Again, we can make up words from the L from the ML word uh from the ML world. Um but categorizing things is known as coding. Um and of course each explanation is natural language. Uh so it is hard to do this coding uh unless you use yet another LLM to do the categorization for you. We'll be seeing how that works uh towards the end of today.

20:37 Um, and like I said, you can hand a whole pile of these explanations to a coding agent and let the app fix it for you. Uh, this is the full loop that we're going to build today. We're going to, uh, instrument, trace, eval, and iterate. Each step feeds the next. You define what you want. Uh, you build it, you measure how well it works. You ship it, you monitor it in production, and you iterate based on what you see.

21:01 Every traditional software product goes through a loop like this. Uh but for AI the measurement step is where most teams fall down and evals are how you measure. Evals are the connective tissue uh across this whole life cycle. Uh so let's get started building it. Step one is setting up tracing with Arise EX AX before we can run evals. We need something uh some data to run our evals on.

21:26 Uh and you can't evaluate what you can't observe. So observation comes first. So hopefully you've conquered the Wi-Fi uh and you have our notebook at this point and that is running around to make sure uh that you do. Uh let's go to our very first code cell which is where I install our dependencies. Uh claude agent SDK is self-explanatory. That is the claud agent SDK.

21:56 uh open inference open inference instrumentation claude agent SDK is the auto instrumentation package for the claude agent SDK. This is how it knows to capture traces from cloud agents uh without you changing your application code. Uh as I mentioned earlier, you can use open AI instead if you want to uh or some framework. Um we have packages for all of those as well.

22:19 You just need to swap them in. Um, Arise and Arise Otel are the packages that send uh your traces to AX and let you read them back. Uh, and there's a Phoenix package in there that we use as utility package. Don't worry about it. Um, the anthropic package lets you use Claude both to power our agent and to judge its output later. Uh, so uh, go ahead and run this cell if you haven't already.

22:45 And while it's installing, I'm going to talk about what it is that we're going to actually build. Um the claude agent SDK if you haven't already used it is Anthropic's framework for building agents. So Anthropic's answer to lang chain and things like that. Uh it can use tools, it can search the web, it can maintain conversation context across turns.

23:02 Uh open AI has their own uh agent SDK and of course there are whole agent frameworks like crewa align chain and mastra and uh AX is compatible with all of those. So no matter what framework you've built uh your agent in uh it is already instrumented and you can just turn on logging and all of your traces will light up. Um hopefully your install is done by now which is uh optimistic timing on my part.

23:31 Um next set your keys. Uh in this case I have stolen my keys from uh collab. If you are feeling naughty, you can just paste them in directly to your cell and put them there. Um, three things go in here. Like I said earlier, you need an anthropic API key, uh, which is going to power our agents LLM calls, and it's also going to power the judge later. You need an Arise API key, which you get from your settings, which authenticates your traffic to the AX service, and you need an Arise space ID, which tells AX which uh,

24:01 workspace to put all of your data in. Uh if your keys aren't working, the usual culprit is a copype error. Uh and of course that is still running around to help you if uh your keys aren't working. That is has been described as a walking security hole. Uh so he will definitely give you a key if you can't figure out how to get your keys to work. So now we scroll to the register section.

24:28 Uh I'm going to bump up my fonts here. One more time. >> The QR up one more time. Sure. Probably. >> Sorry. >> Uh, yeah. I didn't put my mine lives in Collab. So, Collab's secrets feature pulls it out. All right, I'm going to go back to where I was. Hopefully, I can do that. All right. Uh so we've installed dependencies, we've talked about the SDK, we've talked about the secrets.

25:31 Uh and now we're at the register. So this is the magic. Uh anybody who works for Arise will tell you that you only have to add two lines of code to your application in order to instrument your call. And this is these are the two lines of code. Uh what we're doing here is we're passing in our arise space ID, our API key, uh and a name for it to trace everything to.

25:54 Um the register function is setting up open telemetry. Open telemetry is the industry standard for application observability. Uh and it has a layer on top called open inference which adds LLM specific attributes. Things like prompt text, completion text, token counts, which model was called, which tools were invoked. um we point it at ax uh and we give it a project name and that is how your traces get grouped in the UI.

26:19 Uh in this case we're also passing batch equals false because uh when you run in a notebook uh you want to send your telemetry as soon as it happens and not batch stuff up which it does for efficiency in production. Um that one line at the bottom, the claude agent SDK instrument tells the cloud SDK to send a span to ax every time it makes an LLM call or invokes a tool.

26:41 Uh and that is now instrumented. The reason this works with so little code is because open telemetry is an open standard that nearly everybody uses. So the cloud the cloud SDK authors, the open AAI SDK authors, the lang chain authors, all of them have already written the code inside of their SDKs that calls open that calls open inference and sends data back.

27:04 So all you have to do is say, "Hey, I'm an open inference collector. I live here. Send it to me." One more bit of setup is that the register call sends data to AX. Uh and the Arise client reads data back out of AX. Uh so we'll use it to pull our spans back into the notebook. Uh that is what this line is about. So you need to make sure that you've put in the same keys there or uh copied them across.

27:40 >> Oh, you need to make a you need to make a copy of the notebook from the file menu to be able to edit it. Cool. Um, all right. So, if you've run that run those cells, then we are ready to build. Uh, before we build, one thing worth knowing about, uh, everything we're going to do by hand today, you can also get your coding agent to do for you. The reason we are doing it by hand is so that you have a fundamental understanding of what it is that you're doing and you're not just vibing your way to success.

28:15 Um but we publish a set of skills called the Arise skills uh which know how to do all of this for you at the skill level. So you can install your Arise skills uh and um run evaluations, create data sets, uh instrument your app in the first place. Um they all use uh AX's recommended patterns and you install them once with a single npx command. They work in cloud code, they work in cursor, they work in codeex and dozens of other coding agents.

28:43 Um, so we're going to do stuff manually today, but in production in a real life environment, we are using our own skills every day to do this stuff. Uh, so now let's build our agents. Uh, when I asked if everybody had built an agent before, everybody put their hand up. So I'm not going to spend a lot of time explaining how this agent works. I'm going to assume that you already have an agent somewhere and you're just trying to instrument it and make it work for you.

29:07 Um, the fake agent that I'm using today is a financial analysis chatbot uh using the cloud agent SDK. You give it a stock ticker and a focus area and it searches the web for the latest information about that stock ticker. Uh, researches real current financial data and writes a report for you. Uh, this is a a real use case that whole startups are built around.

29:29 Although uh my agent version is extremely simple. Um, I'm assuming uh that you built an agent already. Our agent works in two turns. Um there's first a research turn that uses tools to search the web and gather data. Then there is a writing turn that compiles the research into a readable report. Uh so let me walk you through that. Uh here is my research prompt and my write prompt.

29:53 These are deliberately extremely simple. They are going to cause errors later and those are the things that we are going to use our evals to debug. Uh and then I'm setting up my claude agent options. I'm using uh Claude Haiku uh as the agent underneath because Haiku is you know capable but it will make mistakes and we want some mistakes so that we can fix them.

30:12 Uh and it's also fast and cheap so I don't burn a lot of money uh of your money when as we do this. Um the permission mode controls uh what the agent is allowed to do autonomously. There is a catch hiding in the in the specific choices that I have made and the specific allowed tools. uh that I am giving this agent right now as we're going to find out later.

30:38 Uh and now let's look at our two turns. Um turn number one is research. There is a bunch of boiler plate in here that I'm not going to explain which is just about outputting the output so that we can see what it's doing as it's going. Uh turn two is writing the report with again a pile of boiler plate. Um the prompts are critical. So this is where I want you to start doing your own changes in your own notebook.

31:02 I have given it extremely basic prompts at the top. If you can see if you can think of a better prompt just off the top of your head uh that is going to do a better job than the prompts that I put in uh at solving this problem. What would be better than what I put in to uh do the research in the first place? What would be better than what I put in uh to do the report writing?

31:24 Um, and if you oneshot it, then you're going to have uh less stuff to do later on. Um, one of the things we do here is we wrap the whole thing in an open telemetry span. Uh, this is because if I don't do that, then claude inter Claude sends this as two the two turns as two separate spans and we wanted them to be grouped together in the UI. Uh, so I've wrapped them in a span together so that they come out as a single agent turn.

31:55 Um, this is just saving us some time later. I wanted to be clear about why I'm doing this uh bit of code. Um, and now we can run our agents. Uh, I have run my agent in advance because I wanted to know exactly what it was going to do. Um, but now is the time to kick off your own agent. Um, I've asked it to analyze Tesla, specifically its financial performance and growth outlook.

32:19 Um, you should go ahead and run this cell. It'll take a minute or two because the agent will actually search the web. Uh, and the LLM is doing multiple rounds of reasoning. Um, the agent receives our research prompt. It decides it needs to search the web, which is a tool call. It reads the search results. It decides if it has enough information and maybe searches again.

32:39 That is the agentic part. You didn't say do a web search, take the web information from the web search and write a report about it. The agent is deciding is this web search enough information? And if not, I will write I will run more web searches. It will run an in indeterminate number of web searches to get enough information until it has decided what enough means.

32:59 Uh and then it will write the report. This execution path is non-deterministic by design and that is the key here to understand. Uh if you run this again with the same input, your agent might search for different things. It might find different results. It might write a different report. Uh and both reports might be good. Um or one might be good and one might be bad.

33:18 And that is exactly why we need evals. We can't predict the output from the input alone. Uh, and every single one of those decisions is being captured as a trace. Uh, so I'm just going to pop open my last 12 hours or my last 24 hours, I guess. Cool. This is a pile of traces. You're going to understand what all of this is later, but here you can see the Tesla one that I ran earlier.

33:43 This is what it looks like when I go through. You can see every single turn. Uh it's calling a skill. It's doing a tool search. It's doing a web search. Uh you can see every single thing that it does. You can see that it ran four web searches before finally deciding that was enough information. Then it did the second uh call in the step and it wrote the report.

34:04 Uh one of the things that it does uh that we're going to see later is uh I told you there was a catch in the tools that I gave it. Uh here's this step that was unexpected when we were putting this demo together. It tried to write the report. I told it to write a report. It tried to write the report as a markdown file to disk. It is living in a collab notebook.

34:27 So there is no uh there's no file system for it to write to. Uh so the write step would fail. Uh and this is something that you need evals for. You need evals to detect when your agent is doing something helpful but incorrect. Uh because this fails silently. It gives you output anyway, but it's doing this unnecessary extra step of trying to write to disk uh without you knowing that it was there.

34:52 Um, so back to the collab, this is a pretty good report. Um, it's a financial performance table with real numbers, growth drivers, risk factors, analyst price targets, and a conclusion. Um, it's not bad for a first pass, but it not bad as a vibe. And we are here to replace vibes with actual numbers. Um, so like I said, uh, you can see all of that stuff in the uh, traces.

35:25 We didn't have to do anything to get all of that information. We didn't have to like instrument every single line to get the LLM call and get the tool call and all of that stuff. All of that is built into cloud agent SDK. just it is like it's built into open AIS SDK. Uh and each one of these rows in this table is a span. Um traces reveal every decision the agent made.

35:44 Without them, all you see is I gave it a prompt and I got a report. With traces, you see every step. Uh if you click into any span, like I said, you'll see exactly what the model received as input and exactly what it returned. Um, and that is what observability means in practice. Not just did it work, uh, but how did it work and where exactly did it go wrong?

36:08 Uh, when we run our eval in a few minutes, uh, we're going to pull the input and attribute output attributes out of these spans and feed them to our evaluators. The input becomes the evaluator's input. The output uh, becomes what gets graded. Uh now we need to generate some test data to run meaningful evals. Uh we need more than the one trace that you've generated so far.

36:32 We need a body of data. Uh in the notebook there are 12 test queries covering different tickers and analysis types. Uh you should kick those off now. They take about 5 to 10 minutes to run assuming that you've got Wi-Fi and everything's working for you. Um I'm going to show you what they look like. Uh and oops this is them here. So the key thing here is uh diversity.

37:04 I've given it a bunch of obvious queries that should work. I've also given it some edge cases. For instance, I've given it one where I gave it two tickers, Apple and Microsoft, and told it to run a comparison. Uh which is something it could theoretically do. uh I've also given it the same uh tickers multiple times uh in places so that we can see uh the non-deterministic output uh when we say uh Amazon profitability trends and outlook versus Amazon AWS performance and profitability um different tickers and different

37:38 questions different levels of complexity are important uh to generate test data that's going to cover the range of things that users will actually ask although spoiler alert in reality your users are going to ask really weird things that you're not going to be able to predict in advance and that is another one of the reasons that eval are important.

37:54 Um the uh Rivian query rivn uh asks about a company with much less public data than Apple or Microsoft and the Coca-Cola dividend yield query is a very different kind of analysis from a growth growth stock query. Um again this is diversity that is important. Um, diversity in your test set is how you catch uh things that your agent is going to be not good at.

38:22 Uh, so while yours are running, mine have already loaded so I can show you all of my data. Uh, these are all of my spans from all of those runs. Um, you can see uh at the bottom there are 13 uh the which is the initial Tesla run and 12 test queries. Each one is a complete execution of the cloud uh financial analyst. So now we have the data but before we write any evals uh we need to actually uh look at the data.

38:51 Um this is error analysis. Uh and the whole step is one instruction. Read your traces before you write eval. Before you write a single evaluator start with your data. Uh read or a dozen or more of your traces end to end. This is the most important practice in this entire workshop. uh focus on the ones where something went wrong. What was the input? What was the output?

39:14 What specifically is broken? This sounds extremely old-fashioned. Just read the input. Uh but it is one of the highest value activities in agent development. Uh Anthropic, for instance, invested in tooling specifically for viewing eval transcripts and their team spends a whole lot of time just looking at eval transcripts every day. Um, a trace tells you whether the agent made a genuine mistake or whether your graders rejected a valid solution.

39:38 Uh, if you automate before you understand your failures, then you're going to create an eval that measures what's easy to measure instead of what actually matters. You need requirements first before you can categorize failures. You need to know what success looks like. You can't say it doesn't work if you haven't defined what it works means. Uh for our financial analyst, what does a good report look like?

40:06 Uh it should reference the correct ticker. It should include some re real recent actionable financial data uh and include actionable recommendations. It should distinguish between forward-looking analysis and historical summary. Those are our success criteria and they seem obvious uh when you write them down, but lots of teams never do. Um they ship an agent and then react to complaints instead of defining the bar up front.

40:28 Uh, writing requirements down turns vague disappointment into specific testable criteria and those criteria are what turn into your evals. Here's the thing about defining success that it's important to note though, which is that it is not something that your engineers can do or not something your engineers can do alone. The definition of good usually lives in the domain knowledge and the domain knowledge lives in the people that your company will often laughably refer to as non-technical.

40:56 So your product managers, your QA, your support team, those are the people who know what the definition of good really is. They know what the failures are going to look like. Those are the people who should be helping you write your evals. Uh OpenAI put it really nicely. They said that people management skills are AI skills. So clear goals, uh direct feedback, knowing what your value proposition is, uh those skills matter more than ever when the system is probabilistic.

41:23 Uh and this is one of the things that AX is built around. Anyone on your team can read traces or add annotations and contribute to your criteria. This is not just an engineering surface. You're expected to have other members of your team in here. A quick note on where to get this test data. Uh we have these traces because we already ran the agent ourselves.

41:43 But what if you're building something new and you don't have real traffic yet? Uh you can do what we just did, which is you can use synthetic data. You can have an LLM generate diverse queries across your expected categories. So, uh, research Tesla financial rep performance is one phrasing. What's going on with Tesla stock is another query that's asking the same thing.

42:02 Yo, is Tesla a buy right now is a third way of expressing the same query. Uh, they're all the same intent, but they look very different. So, vary the phrasing, the complexity, the level of specificity. Uh, a domain expert should review your set of queries. Uh, because LLM generated queries tend to cluster around obvious phrasings and miss the weird stuff that real users write.

42:25 You also need to include edge cases in your test data. So things like non-existent tickers, multi-part questions, jailbreak attempts. Uh those might be 1% of your traffic, but they are the 1% of your traffic that ends up, you know, being a PR disaster in the press. Your test data should look like production data, not what you wish production looked like.

42:42 Uh so synthetic data can get you started, but as soon as you have production data, production data should be what you're running on. So let's actually examine our traces. Uh like I said, the first 13 traces are the first 13 traces that I wrote. Um most of them produced long structured reports in line. So Tesla, Apple, Nvidia, stuff like that. Uh 3 to 7,000 characters uh of report right there in the trace output with executive summaries, valuation tables, explicit buy or hold recommendations.

43:16 On the surface, most of these look fine, but like I said, three of the 13 were doing something weird, which was uh they were uh the rep output in them is much shorter because it just says, "Oh, I wrote this to disk for you." Uh and the output went to disk. Uh so instead of having a full report, uh it has a summary of the report and the the actual body of the report went to the non-existent disk.

43:44 This is what the eval is designed to find. Um the agent decided to write a file with no permission. The write silently failed and it told us the report was saved. Uh and that is the kind of pattern that you only see if you actually read the trace. Um so this is I believe one of the three sources. Nope, that one had worked. Oh no, the Wi-Fi. Here we go.

44:24 So, this one's showing another kind of failure where it just got stuck in a loop forever. It was trying to find information about who was it trying to find information about? Microsoft. Uh, and it just did an endless series of web searches and took forever. Um, this is what it tried to write to disk. Uh, and this is what it actually wrote to output.

44:45 Uh, right here we go. This is the report is saved as Microsoft cloud segment financial report.mmd. This report is useless because it's referring to a report that doesn't exist. Um, there's also uh an example of a confidently wrong. If we go to the Rivian report uh which is RIBN, where did it go? There we go. Uh, Rivian has the same problem that the Microsoft one does, but even inside of the full report that it tried to write to disk, there's a whole bunch of really confident information about a private company that you

45:21 can't possibly verify. Uh, these might be hallucinations, they might not be hallucinations, but we haven't done anything to check whether or not uh this report is really grounded in reality. Uh, so we're going to do an eval about that later. There is a structured way to do this reading uh which as I mentioned earlier is called coding. There's open coding and aial coding uh and it comes from qualitative research.

45:47 Open coding means you read each trace and you write down what you see. So stuff like vague recommendation made up a number talked about the wrong ticker. Uh you're not trying to be neat. As a programmer, it is very uh tempting to try and come up with categories in advance and say, "Oh, this is a you know, retrieval failure or whatever." Uh don't do that when you're doing open coding.

46:09 Just write down what the problem was in as natural language as possible. Uh and uh then when you find the failure, sorry, and then uh you want to do aial coding afterwards. AIAL coding is when you look at uh all of the open codes that you've put together and you then you do the categorization. You say okay now there are five things that look similar.

46:31 I'm going to give them all a single category and uh turn them into categories. Um the second pass the the uh aial coding pass is structural. Uh and like I said the tend the tendency the temptation is to try and go straight to a coding and you should resist that. When you find a failure, you should ask why it failed. So the response was wrong is a symptom.

46:56 Did it get bad search results? That is a retrieval failure. Did it get the right data uh but the wrong conclusion? Then that is a reasoning error. Did it make up a stock price? That's obviously hallucination. Did it answer outside of its domain? That is a scope violation. Did it do something that you didn't tell it that it should be able to do. Each root cause, each axial code uh points to a different fix.

47:17 If you don't know the right call, if you don't know the cause, you can't pick the right remedy. Once you've categorized uh which I did here, uh you can sum them up into uh a table. So, uh confession time, I didn't actually read all of those traces and code them. I just gave them random axial codes. Uh but what you end up with is this table of root cause frequency things like looks good possible hallucination reasoning gap unverifiable data missing recommendation.

47:55 It's tempting to look at the top one the possible hoodation and decide that that is the most important one to fix. But in reality uh you're going to want to uh balance between uh frequency and severity. So if one time in a hundred instead of doing a financial report, it gives the user instructions on how to make a bomb, you fix that one first because that is the most severe possible failure.

48:22 Uh so you have to uh multiply the severity by the frequency to get to which ones you decide to fix first. Uh and that is why you look at the data before you write your evals. Uh one more concept before we get started on that is the Swiss cheese model. um which is a concept from safety engineering. Imagine each layer of defense as a slice of Swiss cheese.

48:44 Each slice has holes uh gaps where problems can slip through. Uh but if you stack enough slices, the holes don't line up. So what gets through one layer gets caught by the next. Uh your code eval catches format issues but misses semantic problems. Your LLM judge catches reasoning gaps but misses subtle hallucinations. And your human review catches the subtle stuff but it can't scale to every trace.

49:07 No single eval in that set uh catches everything and the combination is what gives you coverage. Uh so stack the cheese. Um now that we know what's wrong, let's automate the checking. Um the next step is code evals. This is no model configuration, no API calls, uh just Python. We're going to write the simplest useful eval that I can think of. Uh, our agent is supposed to analyze a specific stock ticker.

49:36 What I'm going to do is write a Kodi valve that checks whether that stock ticker actually showed up in the report. Uh, which is a test that my initial demos absolutely failed a number of times. Uh, an LLM judge would be overkill for this. You're looking for a specific string. It's somewhere in the report. You don't need an LLM to judge whether or not you mentioned the ticker.

49:58 So, there's two ways to do this. There's two ways to do everything in AX. one is programmatically uh and one is via the UI. Uh in the uh notebook you can see the programmatic way. What I'm going to show you on screen is how you do it in the UI. So this big button in the top right is what you want. You want to add an evaluator. In this case, I've already added my mentions evaluator, but you will want to add an evaluator which will give you this screen.

50:31 Uh here you're going to see basically some Python. Um in the notebook I've given you the Python that you can use uh to either do your uh programmatic version or your version in the UI. Uh this is a very simple Python function which returns uh an evaluation result which is looks exactly like those evaluation results that I showed you earlier. Uh it has a label, it has a score.

51:00 Um so uh once you've written what your code evaluator does, you're going to have uh parameters to your code evaluator. In this case, I've called them query and report. And you're going to have to tell AX uh which variables in your trace are those variables. So there's this extremely fun UI, the single trace UI, uh where you look at a single trace, and you can literally just click uh the specific thing in your output that you want to be that variable, and you tell it, okay, that one should be the query, and that one

51:41 should be the report, and then it's mapped. Um, you should also uh be changing your evaluator to uh work on uh a trace as opposed to a span. That is what gives you that UI. Um, once you've done that, you need to save your code evaluator to our evaluator hub. Uh, and then you hit save to save the evaluation. So assuming you've done that. >> Yeah. >> Sure.

52:30 Uh so he asked for more detail on how code evaluators work. Uh so code evaluators are they're completely deterministic functions. You're just taking input and output and saying was this output good according to some definition of good. So in this code uh I'm doing very very simple Python. I'm just saying uh look for all of the words that are in capital letters.

52:51 Uh strip out the ones that are obvious acronyms that I already know and of the remaining words are any of them the uh the acronym which is the stock ticker that I'm looking for? That's all it's doing. Uh lots of code evaluators are even shorter than that. They're something like is this output 500 characters long or less? Uh is this output parsible JSON?

53:11 You're doing something very simple, very deterministic uh for the purposes of having that first line of defense in your Swiss cheese. Um so I talked about how to add stuff already. Uh in the notebook you can run the eval programmatically. Uh this is how to do this is these are the instructions on how to do it via the UI if you followed that. The alternative is to do it in the code where you can just run it directly which is uh easier to do if you're following along in a notebook.

53:50 Um, a code evaluator returns a label and a score. Like I said, there's no LLM needed. There's no API call. It's instant. Uh, and uh if you're wondering why we're wrapping it in suppressed tracing, it's because uh the evaluation is itself uh a call to the uh agent SDK which would then get stored as traces and I didn't want traces of my tracing happening because that gets too meta.

54:15 Uh so I told it to suppress tracing when I'm running an evaluation. Uh if you look at this programmatic run uh you will see that 12 passed and one failed. Uh there was a missing ticker for Amazon. Uh the specific trace that went wrong uh is when I asked you'll note you'll notice I told you earlier that I made two queries to Amazon. One where I asked about its profitability and one I asked were about uh AWS's profitability.

54:43 Uh if you click into the traces uh which I could attempt now um you'll see that the report that asked about AWS uh wrote a report that was entirely about AWS only about the AWS branch of Amazon rather than uh all of Amazon. Uh and so it failed to mention the Amazon stock ticker entirely. It talked only about AWS. So it's a really subtle failure because it did a bunch of research.

55:12 It found out a whole bunch of stuff about what was happening at Amazon. It just didn't think about what was happening about Amazon as a company. Uh only about AWS because I gave it that extra prompt like I also want to know about AWS profitability. Um so uh looking at those results back in AX um they show up as annotations next to our traces. So you can click into any one of these traces uh and you'll see evaluations.

55:42 I've run a lot more evaluations uh than you have so far, but you can see my mention sticker evaluation has run with the span. It's run at the span level. It has hit a pass uh and it's given a score of one. Um the notebook has a little helper called log eval to ax which is doing that work for you. takes uh the spans that it ran offline and the tests that it ran offline and pushes them back to AX for you.

56:07 Um so now every score now every eval has uh a mention sticker evaluation uh that you can look at and filter by uh and look for the failing failing ones. So why does this matter? Because it is your first line of defense. Uh if your financial analyst agent writes a beautiful report about m Microsoft when the user asked about Tesla uh which is a thing that actually happened to me when I was building this uh then no amount of answer of eloquence from your agent uh matters because the answer is wrong.

56:38 Uh so you can catch this with five lines of Python. Um the ticker check catches uh really basic failures but it catches them really cheaply. That is why it is important. Um, Kodi valves aren't just toy examples. Often, you're going to want to know that your output is JSON, uh, or that it has a certain length, or you're going to want to avoid forbidden phrases like as an AI language model.

57:03 Uh, and those are production critical checks. A code eval doesn't have to be a simple string operation. Kodval means the grading logic is deterministic. Uh, but it can do complicated things. It can query a database for you. It could hit an API and get the actual stock price to make sure that it wasn't hallucinating a stock price. uh anything with a grading always gives the same answer for the same input can be a code evaluator.

57:29 One important principle is to grade what the agent produced and not the path that it took. Uh I referred to this earlier. There's a common instinct to check that the agent followed a specific sequence of steps. Did it call this tool first and then that tool in that order? Uh and uh practitioners have found this to be in practice too rigid. uh agents regularly find valid approaches uh that the eval designer didn't anticipate.

57:52 Uh so uh if you grade the path um you punish creativity. Instead grade what the agent produced. Did the output have the right information? Did the final state match what you wanted. Uh our ticker check doesn't care how the agent found the data. It just checks is the ticker in the output. Uh that is the right level of abstraction for most Kodi valves.

58:17 So now let's get to built-in evals. LLM evals. This is where it begins to get more interesting. Uh code can check whether the ticker appears but it can't check whether the analysis is good. Is the financial data accurate? Is the report complete? Are the recommendations well reasoned? Uh these are semantic questions and for semantic questions you need a judge.

58:35 Every LLM as a judge has three parts. Uh it has a judge model which is the LLM that is doing the grading for you. It has a prompt template which as I mentioned is also called a rubric. Uh which is the criteria that the judge applies the rules of what defines good. Uh and it has the data which is the examples being evaluated. AX keeps these three things separate which means that you can mix and match.

58:59 You can try the same criteria with different judge models. The same judge model with different criteria. Uh it is modular by design. The good news is that AX ships with a B with a bunch of built-in eval so you don't have to write your own. There are a bunch of types of evals that are always uh that are often going to apply to uh most AI applications.

59:20 So we've already written the prompts for those and you don't have to write them yourself. Uh correctness, for instance, checks whether a response is factually accurate. Faithfulness checks whether the response stays grounded in the source material. Uh and there are evals for tool selection. Did the agent pick the right tool? Uh and tool invocation. Did it pass the right arguments to that tool?

59:40 stuff like document relevance, refusal detection, many of the things that you'd want to check in an agentic application uh come out of the box. Uh so the first eval that we're going to run uh is a correctness eval. Um and spoiler alert, it's not going to work. Um setting it up is very easy. We give it an LLM. In this case, uh I've chosen uh the anthropic LLM.

01:00:10 uh sonnet 4.6 because it is better than haiku which is the thing that is doing the work. You often want to pick a bigger slower model to be your judge. Uh and then I've just uh given correctness evaluated the LLM uh and uh run it right here. Um like I said uh once we've run the evaluation with evaluate data frame uh we've also uh turned on suppressed tracing here so that our evaluation doesn't send a bunch of extra traces to uh AX.

01:00:46 Um it gives us this output which is a bunch of evaluation rows. Um these are not very readable. This is why you need AX to look at them. Um the judge reads the input and output from each span and grades it against the rubric. Um and we can see that every score is zero. If I scroll across, status completed, name correctness, score zero. The reason it's doing this is because correctness is judging against the uh LLM's definition of of what is correct and what is true.

01:01:27 Uh which means that it is relying on the LLM's training data. What we're doing in this agent is a whole bunch of live web search to get current stock prices and up-to-date financial information, which means the agent is talking about things that happened in 2026, but the LLM that's doing the judging only has information that stops in January. So, if you read the explanations of our correctness eval inside uh inside of AX, you can scroll to evaluations, you can look at correctness, you can say that it is incorrect.

01:02:01 uh I will save you a bunch of reading of a very densely packed column. What it's doing is complaining that there's a bunch of stuff that comes from the future that it couldn't possibly know. Uh so our correctness eval here has completely failed. This is not what you wanted to do. So instead we are going to use a faithfulness eval. Uh faithfulness is a better eval for our use case because uh the judge actually has context.

01:02:24 In faithfulness, you give the uh LLM as a judge the information that it is being based on uh and then say based on the same information that the jud that my agent collected, did it do a good job of writing the report? This is this way they are on a level playing field. They both have the same amount of information and your slow complicated judge agent uh can judge whether your uh quick fast uh worker agent uh was uh doing a good job.

01:02:50 Um so let's look at the faithfulness cells. Uh just like everything else, these can be done online as well as offline. So uh there is a faithfulness built-in evaluator uh which isn't in this view. There is a faithfulness built-in evaluator which you can pull up uh and you can give your input, output and your context to it. Uh again, you can do this in code as well.

01:03:25 And if you're using the skills, this is probably what's going to happen. Your skills will say, "Oh, I'm not going to use the UI. I'm going to use the code." Uh which is why I show it here. Um so the first thing we do is we extract the research context from the traces. As you'll remember, we do uh a two-step process. So I take the output from the first step where it does the research and I attach that to all of my spans uh as uh the context for it the judge to work on.

01:03:54 Uh and then I run the faithfulness evaluator. So uh here you can see spans with context. These are the same spans that I just pulled out of AX. I attached at the column level all of my context. Uh and then uh I called evaluate dataf frame with those spans and I gave it my faithfulness evaluator which I instantiated here just like I instantiated the correctness evaluator.

01:04:21 This gives us uh much more interesting data. This gives us unfaithful seven and faithful six. Uh so uh roughly half the time my very simple report is hallucinating. It is coming up with things that were not in the source data according to sonnet. Um and that is this is a useful eval. This is an eval where we're failing half the time. This is a capability eval where we can climb the hill, modify our prompt and get it better at writing a grounded report.

01:04:56 This is exactly the kind of eval that you want. Um, so that's two built-in evals and two very different signals. And there's a really important lesson here, which is correctness gave us zero out of 13 and faithfulness uh told us that reports aren't grounded in their sources roughly half the time. The difference isn't that one eval is better than the other.

01:05:16 The difference is uh the use case. There are lots of use cases where that correctness eval that I showed you earlier works just fine for some use cases. Uh it's just not in this specific use case where we were doing live real-time research. Um so it's important to know the question that you're actually asking. Built-in eval are your starting point. They give you an immediate signal without any prompt engineering.

01:05:41 Uh but even with a good built-in like faithfulness, uh there are a bunch of things that it can't check. uh our financial analyst uh should produce actionable recommendations. It shouldn't just summarize data. Um so not just you know give you a financial report but tell you what to do with the financial report. Should I buy, sell or hold? Uh no built-in template checks for that.

01:06:07 This AX is does not cover every single use case. So you have to build a specific eval that checks for actionability. You have to build an custom eval that is about your domain. If you are writing real eval, you're absolutely going to have to write a custom eval with a definition of good that we didn't think of. So that brings us to the next step, which is writing a custom eval rubric.

01:06:28 This is where you make eval truly specific to your application. And writing a good eval is complicated. So I'm going to spend a little time talking about what makes a good eval rubric. Uh we recommend four parts of a custom rubric. The first is to define the judge's role. Um, you should spell out ex explicit pass and fail criteria. You should label the data with XML tags.

01:06:50 Uh, and you should define the output choices outside the prompt itself. Uh, I'm going to add one more thing on top of those which is actually kind of controversial amongst practitioners, which is I tend to include examples of good and bad output. Uh, some practitioners believe uh that including examples will make your uh, eval overfit. Uh but let's walk through all of those.

01:07:14 So the first part is defining the role. Um you'll notice this uh what is by now a tired trope which is you are an expert financial analyst evaluator. That is not the important part. Tests show it makes some very small difference uh to whether or not uh your eval works. The important thing is to give it the context of what it is looking at. You are looking at a financial report and what I am expecting is financial analysis.

01:07:40 Those are the important pieces of context that you are giving the LLM there. Um you're also telling it that it should be that what you want is something actionable. So that brings us to part two. So uh this is where most people underinvest. Don't say a good response because that is uh an aspiration. That is not a criterion. A good response is helpful and accurate sounds like you're adding more detail, but you're still being just as vague because helpful and accurate don't have any meaning as far as the LLM is concerned.

01:08:11 Uh instead, you should be listing exactly what makes a report actionable. So list exactly what makes it actionable and not actionable. Every criterion should be specific and observable. Something that if you were a human looking at this report, you'd be able to know whether or not it was there and you'd be able to say yes or no. Uh so contains specific recommendations.

01:08:32 That's something you can definitely say yes or no. Includes forward-looking analysis, not just historical data. That's a clear distinction that you can say yes or no to. Uh and the important thing here is that each of these criteria maps to something we saw in our analysis. If we when we were reading the reports uh we saw we actually observed when we read our traces or at least when I read our traces uh that uh these are ways that these reports are going wrong specifically that is not a coincidence.

01:09:00 Your error analysis is what tells you what criteria to write. You don't come up with them from whole cloth because it's not going to occur to you. Where you get that data from is from reading your traces and going oh it didn't do that. My axial coding said it does this specific wrong thing. Uh so specific criteria produce consistent judgments uh while vague criteria produce coin flips.

01:09:22 The third thing you should do in your rubric is you should label the data with XML tags especially if you are using anthropic. Anthropic loves XML and finds it very easy to understand. Um your rubric has variables the input and the output and you want to make sure that it can tell the difference between your instructions and the input and the output and XML tags are an excellent way to do that.

01:09:42 Um, so put a user query tag around the input, a financial report tag around the output. Uh, clear boundaries reduce the chance that the judge confuses the question with the answer. It is especially important if you're including examples uh to label the examples with XML tags because otherwise it will go, "Oh, this is exactly what I'm supposed to write and just spit out the example back to you."

01:10:04 Um, which brings us to the examples. This is the part a lot of people skip. Uh, in my opinion, it is the part that most improves quality. LLMs are okay at following instructions. They are increasingly good at following instructions. But what they are really good at is copying an example that you've already given them and replacing the things that are relevant to the current situation.

01:10:25 So if you give it an example of exactly the kind of report that you're looking for, it's going to follow that format exactly. That might or might not depending on your domain be a good idea. If you if your example is too rigid, then the LLM will always give the same report and you have reduced its scope for creativity. Uh but on the other hand, if it is being too nondeterministic and your report is being crazy, uh then giving it an example will ground it in what exactly you are expecting.

01:10:55 Uh here's what an actional example looks like in the template. Uh it has everything. It has uh percentage revenue growth, a PE ratio. It identifies a concrete risk with a number to back it up and it gives a specific recommendation. Uh that is what actionable looks like. Here is a non-actionable result. Uh it's not wrong. Uh everything it says is correct.

01:11:15 Uh but it doesn't say anything useful. It doesn't have any specific data. It doesn't have any specific risks. It just says investors should consider various factors. Uh that's not a recommendation. Uh so the judge now has a concrete reference for what you actually mean by the labels that you were giving it. Uh and this can dramatically improve consistency especially on the edge cases where the report uh would have otherwise been somewhere in between.

01:11:40 Uh and this is uh an important part. It is very tempting if you're writing this rubric to say and now spit out the word correct or incorrect because that is what you wanted to do. You do not need to do that anymore. Uh the way that it actually works under the hood is we have given the LLM a tool that says you call this tool when you are done and it will call that tool with correct or incorrect when you're done.

01:12:03 So you don't need to tell it what output and in fact telling it to spit out a specific string as output is more likely to confuse it and make your eval fail. So just in the UI uh set your uh choices uh or when you are setting the when you're creating it programmatically set the choices uh and let the LLM do its thing. Um when you are coming up with these choices it is best to keep those choices simple.

01:12:27 Uh binary is best two clear labels. Uh numeric scales like a scale of 1 to five. Uh they are tempting but the LLMs are bad at them. what is the difference between a three and a four? You would have to write a whole bunch of rules for the LLM to know that. Uh it doesn't know that automatically. Uh whereas you can give it a pass fail and it will be pretty good at deciding what is a pass and what is a fail.

01:12:52 Uh so if you really need more categories, you can do like a partial result. You can say a 0.5 score is where it did a a half-ass job. Uh but by default, go for a zero or one. Um, one more practical tip is chain of thought for judges. Ask the judge to reason before it scores and have it explain its thinking first. Uh, this measurably improves uh, accuracy.

01:13:17 Uh, when the judge has to articulate why something is actionable or not or why it is correct or not. Uh, it makes fewer mistakes than when it just picks a label uh, without thinking first. Uh, this explanation is also very helpful for you. Uh, but it also helps the judge think better. Just one of those weird things about LLMs. Uh so if you are coding along, now is the time to find your actionability template.

01:13:44 Uh this contains all of the example all of the instructions that I uh was just talking about. Uh can you do better than me? Can you oneshot your way to a perfect eval uh by modifying this eval template? Um, as with uh all things in AX, you can do this via the UI as well. So, you can create an online evaluator uh and you can create a new LLM as a judge.

01:14:08 In this case, I've already got it, my actionability evaluator uh which has literally exactly the same text uh that I just showed you in the notebook. Um, and as before, I had to pick a trace uh scope. Uh, and because it's an LLM as a judge, I had to set up a template. Uh, sorry, I had to set up an LLM for it to use in the UI. Uh, which means that you need to give it your anthropic key.

01:14:36 Again, uh, and I picked my inputs and outputs via the single trace method. Uh, and then I saved it. One of the things that you can do with online evals uh is you can set them to run continuously against new incoming traces. This is something you'll almost certainly do in production. Uh and if your evals are expensive or you have a whole lot of traffic, you won't want to do it on all of your traffic.

01:15:00 You can uh set it to serve to uh rate only a percentage. That's the sampling rate. So you can say, you know, just look at 1% of my traffic and that will still give you uh a whole bunch of eval data without breaking your budget. Uh so back to the code version, wiring it up is just like wiring up the other action the other uh evaluators that we've given.

01:15:29 Uh we've created a classification evaluator. We've given it a name actionability. We've given it an LLM. Uh we've given it that template that I just showed you and we've given it those choices. This is us programmatically doing what we would have done in the UI where we said your choices are correct or incorrect. Uh actionable or not actionable scores of one or zero.

01:15:47 Uh and then uh we ran evaluate data frame. That gives us uh the data that I was skipping over because I was already run earlier. So you can see actionability actionable. If we look at our evaluations tab, we can scroll down to our actionability and we get this really great explanation about why the report is actionable or not actionable. Uh lots of our reports scored actionable uh but several are not which is awesome because it means that we have again we have a capability eval uh and a hill that we can climb.

01:16:30 Um, so in the notebook, uh, this is where you log your results back to AX or if you've run it online in the UI, your evals, your eval results are already in the UI as you've seen. Um, so you can see now in in AX that every span has multiple eval scores. It has mention sticker, it has correctness, it has faithfulness, it has actionability. uh and you can filter and select uh for specific uh eval results.

01:17:02 So, eval.actionability.label equals not actionable. that finds you the failing set. This is the set that you're like, okay, my eva my eval says that my code messed up here. I can turn this into a data set and I can start running uh new prompts against this. I can run experiments against this as we're going to see. Um before we move on, I want to flag a few a few common antiatterns uh that turn evals from a useful tool into noise.

01:17:46 Uh the first one is to treat your eval prompts like code. And I mean that literally. They're very sensitive to war to wording. Small shifts in your eval prompt uh can lead to dramatic changes in your eval results. Uh so you should version them. You should test them on examples where you know the right answer. If the judge if the LLM as a judge disagrees with your human labels uh on 40% of examples, then your eval is probably wrong.

01:18:12 It's not that the your results are wrong. It is that you have written the eval incorrectly. uh AX has a prompt playground where you can iterate on uh rubrics without touching your code and it's very useful for doing this kind of iteration of your eval. Uh the second antiattern is the god evaluator. Um it is tempting to build one big mega evaluator that checks everything at once.

01:18:36 It'll look for accuracy and tone and completeness and policy compliance and formatting all in a single prompt. Don't do this because it is a nightmare to calibrate. uh when your god evaluator says fail, which of those five dimensions failed? You don't know. So suddenly you're like, "Oh, well, I'll have it spit out labels." And then you're like, "Oh, it's spitting out 15 different labels."

01:18:54 Suddenly it's way too complicated. Uh you can't tell, you can't judge by it, you can't score by it. It is much much simpler and easier to split this up into uh one evaluator per dimension. I've already done this today. I have a ticker check. I have faithfulness. I have actionability. Three separate focused evals. and each one tells you something specific.

01:19:14 Uh the combination gives you a richer picture than any single eval ever could. If you find yourself writing a rubric that lists six different things that the response should do, then stop. That is six evaluators, not one. Uh the third one is know your guardrails and your northstar metrics. Some evals are guard rails. They're ship blockers. If an agent hallucinates a stock price, that is a hard fail.

01:19:38 Don't ship on that. Uh others are northstar metrics. they are aspirational. So always recommend complimentary investments is a nice to have in this case. It is not a dealbreaker. Uh so know which of your eval is which. It changes how you act on the results and later it changes how you set up your production monitors. A guardrail uh regression should page someone in the middle of the night.

01:19:59 A northstar dip is something that you look at in a weekly report. Which brings us to the next section which is can you trust your judges? This is a question we've been dancing around. I've mentioned it a couple of times. Uh it's time to evaluate your evaluator. We built an LLM judge and it told us that some reports are not actionable. Uh but can we trust it?

01:20:21 How do we know that the judge is correct? Uh the key mental model here is that your judge is a classifier. So you can treat it as a classifier. It takes an input in this case a financial report and it makes a prediction actionable or not actionable. That prediction can be right or wrong. So just like any classifier, you can measure its performance by comparing its predictions against ground truth.

01:20:42 In this case, a human's judgment. Uh you can check the judge's homework. Um the thing about applying judgment is that it is a ton of work. You have to read actually every single output. Uh you have to think hard and apply a label. Uh AX does what it can to help by letting you define annotations and attaching them to spans uh right in the UI. So, let me show you what that looks like.

01:21:11 When you're looking at a span, you can click annotate span and you can add and remove annotation configs. In this case, I've already created one called human actionable. And you can set this span as being actionable or not actionable. And then you can just use these little arrows to flip through all of your traces and set them as actionable or not actionable, which I have conveniently already done.

01:21:30 Uh this is the kind of work that you can give to a nontechnical member of staff. That is why the UI looks like this. They can just click in, read the report and click a checkbox. They don't have to write any code. They don't need to know anything other than their domain expertise of whether or not this report should be considered actionable or not actionable.

01:21:52 What we're doing here when we add human annotations is building a golden data set. Building golden data sets is incredibly helpful uh because it's how you measure uh whether your evaluator is actually doing its job. Uh the way to do it is the same thing uh the LLM did. Give yourself real concrete criteria and stick to them. So don't say this was good, this was bad.

01:22:11 Uh be specific the same way that you've told the LLM to be specific. You failed at this specific thing. You asked I this thing was supposed to be present and it wasn't. uh and eliminate the chance to get lazy uh when you get tired of labeling. Um as I said uh the labeling is extremely uh timeconuming so I did it already. Um once I've labeled some spans in the UI uh I can reexport the spans into the notebook.

01:22:41 Um, and this is a nice thing about AX is that the human annotations that I added by hand come back as columns uh in the export. Um, so we filter down to just the rows that I labeled here. Uh, and now we're going to compare uh what I said was actionable and not actionable to what the judge says is actionable or not actionable. Um, before we talk before we do that though, uh, a little bit more about golden data sets, uh, keep your tasks unambiguous.

01:23:25 If your agent scores 0% consistently, that's almost always a broken task. That's almost always, uh, you've written your rubric incorrectly. uh if you've made your task something that only a human can do or worse that somebody something nobody could really do uh then you're going to get a zero score. Um so for each task create a reference solution a known working output that passes all of your graders.

01:23:48 This proves that the task is solvable and you should also test in both directions. You should have cases where a behavior should occur and cases where it shouldn't. So if you test for instance does it search the web uh then you will end up with a uh an agent that always searches the web even when it doesn't need to. Whereas you should also have a test that says did it not search the web for this very obvious answer where it didn't need to do that.

01:24:16 Uh and if you're doing this for real you should split your label data that you just created uh into a development set and a test set. Uh so use maybe 75% of your labels to it to to iterate on your eval uh tweaking criteria, adjusting examples until the judge agrees with you. Then hold out the remaining 25% to run the judge and see that it actually gets the scores that you are expecting.

01:24:40 Uh that's how you know that the judge is actually generalizing to new examples instead of just fitting to the specific data set that you gave it. Uh this is the same principle as train tests, splits in machine learning. Don't overfit your evaluator to your golden data set. Now let's look at our actionability judge. I have run the judge on some examples and it has come up with uh a list of where they agree and disagree.

01:25:06 Uh spoiler alert I just hit actionable and not actionable at random uh so that I would get lots of disagreement because actually Sonnet is pretty good at this stuff. Um that gives us uh an agreement rate of 46% six out of 13 times. Um in real life uh this is exactly where you dig in. This is where uh you would ask is my rubric wrong or were my human labels wrong.

01:25:36 When the judge disagrees with you, you should read its explanation and decide whether or not the judge was right or you were right. Uh that is the most valuable output of meta evaluation. It usually reveals an ambiguity in the rules that you set down. To fix a rubric, you should read the explanations on the disagreements, find the ambiguity, uh, and tighten the criteria.

01:25:57 So, instead of just includes forward-looking analysis, I'd write includes forward-looking analysis with specific recommendations or guidance. That's more precise. Uh, it rules out the case where the report talks about the future, but never tells the judge what to do about it. Uh then I'd rerun the judge on the same examples and see if the disagreement goes goes away.

01:26:14 Uh this is rubric iteration. Your eval prompt is an application. Uh it is another LLM application. Uh so it needs testing and iteration just like your agent does. Uh and that is what I meant when I said that you should treat your evals like code. So now I'm going to talk very briefly about an ML concept called precision and recall. Uh, precision asks, when the judge says not actionable, how often is it really not actionable?

01:26:44 Uh, and recall asks, of all the reports that are actually not actionable, how many did the judge catch? Uh, with a tiny sample like uh 12 examples, these numbers are going to judge jump around a lot. Uh, but even with the 12 that I have, uh, you can see whether the judge is in the right ballpark or completely off. In most scenarios, you want to uh you want to prioritize recall.

01:27:09 These two measurements are at odds with each other. If you are if you prioritize recall, then you're going to get false positives where the judge has said something is wrong when nothing is wrong. That is much better than accidentally letting through things that are wrong. Uh however, some in some cases when you are tuning, you will say actually I want precision.

01:27:31 I want it to be absolutely right when it's right and I don't care if it misses some things. There are domains where that is the output that you want, but most of the time uh you want to prioritize recall. Um a few known pitfalls with using an LLM as a judge. Um one is position bias. If you present it with two options to judge between, uh the judge tend to favor tends to favor whichever comes first.

01:27:56 Uh this depends on the model. Sometimes the model depends uh prefers whichever comes last but it always has some kind of preference to position. Um there's also length bias. Uh LLMs like longer responses. They tend they tend to score those higher and that say that they are better even if the length extra length is just filler. Uh and the last one I love which is confidence bias.

01:28:19 The judge gets fooled uh by a response that sounds confident just like humans do. If your LLM says things that are wrong, but it says them in a really confident tone, uh, your LLM judge is more likely to believe them. Um, there's also self-preference bias. If you use the same model to do the generation as you use to do the judging, they tend to like their own output.

01:28:41 So, if possible, use both say OpenAI and Anthropic. Use one to do your agent and one to do the judging and they are less likely to agree with each other. uh self-preference bias is why I used sonnet as a judge uh for output generated by haiku. Uh but it works even better if you use a different uh labs model entirely. And remember that the benchmark here is human performance.

01:29:08 Humans are not going to be perfect at judging whether things are or are not actionable. They are going to get it wrong quite a lot of the time. What you're trying to do is come up with an eval that is reasonably close to human judgment. Uh human judgment also varies from human to human. Human inter interrator reliability, so two humans judging the same thing uh is often as low as 02 or.3.

01:29:31 So two experts with the same output and the same rules disagree a surprising amount of the time. So if your judge, if your LLM judge achieves that uh higher consistency than humans, then it is doing better than humans. the judge uh sometimes disagrees with me is not by itself a reason to distrust it. You what you're looking for is a judge that disagrees with you all of the time.

01:29:54 That is the danger signal. A judge that disagrees with you some of the time is a judge uh correctly handling a genuinely ambiguous task. Uh and also failures should seem fair. That is a principle from Anthropic Eval teams that I want to leave you with. When a task fails, uh, it should be clear what the agent got wrong and why. If you look at a failing trace and you think that answer looks fine to me, the problem is probably the eval and not the agent.

01:30:26 So, uh, now we've built evals, we've tested them, uh, let's use them to actually improve our agent. This is how you close the loop. Uh, here's the problem with one-off fixes. Um, if you found some failures, you've read some explanations, you know what to improve, you change the prompt, and then what? How do you know uh whether the fix actually worked?

01:30:49 How do you know if it didn't break something that was working before? If you just run the agent again on a couple of examples and eyeball it, then you're back to vibes. Uh, you need a structured way to compare before and after. And as I mentioned earlier, this is what experiments are for. As I showed you earlier, uh you can filter your traces down to just the ones that are failing and you can turn them into a data set like this.

01:31:14 So you select all of your traces, you click add to data set, and you create a new data set. In this case, uh I've called it AIEWF financial demo fails. Uh and I've already created my failure data set. But this gives you a smaller set of places where you know that your eval is failing so that you can then run uh your changed prompt against just them.

01:31:38 This is faster and cheaper. Um this is a regression test set. uh sorry you should also save a regression test set rather you should get the ones where it is not failing and make sure that those are also a data set that you can check less frequently to make sure that when you've changed your model sorry when you've changed your rubric uh that you haven't accidentally made it fail in places that it was succeeding before um and your data sets are not static um pre-production when you don't have users yet your data sets

01:32:13 are probably going to be synthetic queries you generated yourself or queries that you got an LLM to generate for you. That is fine. It gets you off the ground. As soon as real traffic starts coming in, you should turn your uh that into your eval sets. Uh real failures should replace imagined ones. The union of your failure cases and your passing cases is what people call your golden data set.

01:32:34 Uh the curated labeled set you trust as ground truth for measuring quality. Uh now we're going to improve our agent. Uh but we are not going to handw write the fix. Uh as I mentioned, you can do all of this with skills. You can do all of this with a coding agent. I am assuming it being 2026 that absolutely everyone is using a coding agent to do all of their coding.

01:32:58 Uh so why would you modify your prompts by yourself when you can get cloud code to pull down your failing traces uh modify your code for you uh and improve your application for you? That is what we're going to do here. Um Claude will pull down all of the failing traces. It will read all of the explanations. It will turn the explanations into action action that it takes on your codebase.

01:33:24 Um and here's the important thing. Nothing here is guessed. Claude isn't inventing improvements out of thin air. You haven't just told it get better at this. You haven't just told it make no mistakes. Uh every change it proposes is traces back to a specific failure explanation. Uh so if the judge said that it lacks specific recommendations, uh Claude will read that and rewrite the prompt so that it does.

01:33:47 Uh if the judge said that it presents risks without supporting evidence, then it will rewrite the prompt so that it fixes that too. This is datadriven prompt engineering. uh and now we can automate it. Uh so notice what we've passed in. Uh we've passed in our requirements, what we were trying to get done. Uh and we've passed in uh an improvement prompt telling it uh what you should be working on and how to get better at it.

01:34:16 This is yet another LLM application and yet another thing that is nondeterministic and yet another way that it can go wrong. You could, if we had more time, meta-evaluate your meta evaluator of your meta evaluator and make sure that your meta evaluator is doing the right thing. But we're not doing that here because it gets too complicated and also we only have so much time.

01:34:35 Um, but uh good requirements keep claud. So you should be spending time making sure that this uh evaluator uh is as accurate as you can make it. Uh so in this case we are taking this prompt and we are wiring them into uh our two pass uh we're taking the recommendations that we got from Claude based on pulling the uh failing traces uh and we are passing them back to anthropic and saying based on these explanations write a better prompt write a better research prompt write a better write prompt and then we're going to

01:35:16 feed them back into our agent. Um, and now we're going to run an experiment. So, uh, over in the UI, we have data sets and experiments. We have my demo financial fails. Uh, we can, assuming I can scroll, wire up an improved agent, uh, using those new improved prompts. So, you can see the improved research prompt and the improved write prompt. this is otherwise exactly the same agent that it was before.

01:35:49 Uh and you can create an evaluator uh with an experiment. Uh so what we're doing here is we're giving it a task. A task can be anything. So a task could be some subcomponent of what your agent did. So if we were running an evaluator that was about uh tool calls then your task could be just call this tool and see what the output of the tool is and then you would evaluate that.

01:36:15 In this case my task is run the entire agent again from start to finish but it doesn't need to be uh and then we are running the same evaluation results uh against our uh against this new running of the task. So, we've changed the agent. We've rerun the agent, and we're rerunning the eval. And we're turning that into an experiment. This is what the experiment looks like.

01:36:41 Experiments.run. We've given it a name. We've given it the data set that it should work on. Uh, we've given it the task. We've given it the evaluator. And we've told it to run. Uh, this gives us the results which are here. So for every single run uh we've got the output that it got and the new eval score. So uh as you can see this was a data set if you recall where previously all of them were not actionable.

01:37:14 Uh the eval has run and you can see that some of them are now actionable. So we have improved and if we look at the overall score here actually I've nailed it because I had you know a lot of chances to get this right. I have oneshotted my prompt. uh from getting being wrong 50% of the time to being right right 100% of the time. Uh real eval won't work like that.

01:37:38 Uh real iteration won't won't run so smoothly. That is why this is a graph. What you're expecting to do is run experiment after experiment and slowly watch the graph climb as your evals get more and more capable. That is the hill that I'm talking about that you're trying to climb. Uh so the key abstraction behind experiments is that a task is just a Python function.

01:38:02 It takes an example from the data set. It runs whatever you run want and it returns an output. So AX doesn't care uh what the task does internally. So it could run the full agent or just one component like I said or it could call an API as long as it takes an input and returns an output. You can run experiments on it. The power of experiments is controlled comparison.

01:38:23 So you take the same inputs, the same evaluators and the only thing that changed is the agent's prompts. Uh that means any differences in scores is attributable to your change. You're not wondering did it score higher because of my prompt change? Uh or did it score higher because the web search happened to return better results this time. Well, actually you're still wondering that a little bit because the agent is nondeterministic.

01:38:44 So it could have done 10 web searches instead of five web searches. But you've eliminated a major source of variation which is the test cases. So like I said, uh my agent has one shot at this uh which is great, but in reality uh it's not going to be that simple. Um the eval iterate cycle that we're talking about here where you go through uh run an experiment, change your prompt, run the experiment again over and over and over.

01:39:14 That is where the real value lies. Uh not in the score itself, but in the cycle. You run the evals, you look at the failures, you read the explanations, uh, you identify a pattern, and you adjust your prompt or as in this case, you get Claude to adjust your prompt for you. Um, did it get better is a question that can only be answered by evals. You have taken the definition of good and you have codified it in a way that is incredibly useful.

01:39:40 Uh, so this allows you to move from I think it's working to I can prove it's working. A practical question that comes up fast uh is how many samples do you need? Um for workshop scale experiments like today, 12 to 20 examples gives you uh directional signal. For actual shipping decisions, you're going to want uh more. You're going to want something like 200 to 400 is a good target.

01:40:02 But note that 200 to 400 is still like a humansized number that you could possibly accumulate over a couple of weeks. Uh it's not tens of thousands or millions like you would use if you were training an ML model. um there are diminishing returns. Uh to have your your margin of error, you have to quadruple your sample size. Uh so at some level, at some point, you have enough signal to be able to get along uh to get along.

01:40:30 Um when you're iterating, where do you invest? Uh there's a hierarchy. Um data quality fixes have the highest impact. If your agent is searching the wrong sources or your knowledge base has stale information, then no amount of prompt iteration is going to help. you should fix the data first. Uh then prompting improvements are next. Few shot examples, explicit instructions, the kind of stuff that I mentioned today, uh constraints on what the agent should and shouldn't do.

01:40:55 These are often the highest ROI change uh to your agent. Um model selection comes third. Sometimes a more capable model solves problems that prompting can't. Uh but it will also cost more and be slower. So that is a trade-off that you have to make. Uh the thing to not spend a lot of time on is hyperparameter tuning. Things like temperature and top p.

01:41:15 It's very tempting to just go for them because they're easy to tweak. Uh they very seldom make a high lever change like prompting would or improving the quality of your data would so you can try them but they should be the last thing you try. A practice worth knowing is eval driven development. Write the eval before you build the feature. Uh if you want your agent to always verify customer identity before it does something before processing a refund for instance then write the eval that checks for that first.

01:41:44 That gives you a clear measurable definition of what good looks like of what done looks like. Then you can build a feature uh until the eval passes. This is the same philosophy as test-driven development. Uh and something I love about eval driven development is that the people closest to product requirements are best positioned to define success. So product managers, customer service uh people, even salespeople can contribute to eval tasks.

01:42:13 They can tell you what looks good and what doesn't look good in practice. Uh they don't need to write code. They just need to describe in natural language what good looks like to them. And you can boil that down into your Eval rubric. Uh, and so far, uh, I've shown you how to do things both offline and online. Uh, we've been doing everything in a notebook.

01:42:38 Uh, but you're not really done when your dev experiment tra pass passes. Uh, you're done when your agent uh, keeps performing in production on traffic that you've never seen. Um and it's worth spending a couple of minutes talking about that even though uh we're going to stick in the notebook uh for this workbook for this work session. Um the first thing is online eval uh like I said uh and like I showed you earlier online evals are exactly what they sound like.

01:43:05 They take the evaluators that you wrote today uh the ticker check the faithfulness the actionability uh and run them automatically on incoming production traces. The same a val running on new data. Uh I showed you how to run this on a trace. AX lets you do this at three scopes. You can run your evals at the span level. So you can say this particular LLM call, this particular tool call, run a trace on it every single time.

01:43:29 Sorry, run a eval on it every single time. Or you can run it at the trace level where it has the full tree to look at uh and it can evaluate everything from start to finish. Um you pick the scope that matches the question that you're asking. Um, and you don't need to run evals on every single prediction trace I uh production trace rather. Um, LLM judges cost money and at production uh scale those costs can add up.

01:43:55 So you sample 10% of your traffic or 1% of your traffic. Uh, and that gives you a directional signal without you having to run an expensive LLM as a judge on every single iteration, every single uh, instance of your application. Um if you don't want to write your evals by hand uh we have good news for you which is that uh we have built an agent into AX itself.

01:44:18 It is called Alex. Um, you can describe your eval in plain English or you can get your non-technical uh colleagues to describe a pro describe the definition of good in plain English and Alex will actually manipulate the UI for you uh create an eval write the rubric set all the scores correctly get all the configuration correct and then just tell you that it's done.

01:44:40 Uh this is especially useful for non-engineers. Um and the full loop in production looks like this. So you go from application uh which produces traces to online evals which grade those traces and produce labels and explanations. The eval labels feed into monitors which we didn't cover today. Um the monitors can alert your team when something goes wrong.

01:45:03 Um the team can investigate, find the failing traces, save them as a regression data set, um improve the agent, run an experiment to verify and ship the fix. And then the loop can run again. Um, as new production traffic accumulates, you will accumulate new failure modes, you will accumulate new edge cases. Uh, you're always going to be uh iterating on your evals and iterating on your production monitoring.

01:45:32 But it is a payoff that compounds. Every failure that your online evals catch becomes a new test case. You save it to a data set and now it's a regression test for your next round of changes. Over time, this creates a data set that is unique to your application. Uh so your specific failure modes, your specific users, your specific quality bar, nobody else has that data.

01:45:50 Uh and that is a competitive advantage that grows every single time your agent runs. There's one last pattern that I want to leave you with uh which is the most powerful one I know. We did the simplest version today. We fed the judges explanations back to Claude and had it rewrite two prompts. This works beautifully for two prompts. Uh but a real system isn't two prompts.

01:46:12 prompts scattered across many files, retrieval settings, tool definitions, all tangled together. Rewriting that isn't a single API call. It's a job for a coding agent uh that can see your entire repository. Um so this is the pattern. You want to export your failing traces from ax along with their explanations uh and hand the whole batch to cloud code or cursor or another coding agent as context.

01:46:35 Uh this is where the Arise skills plugin that I mentioned at the beginning uh works really well because you don't have because you don't have to know how to call our API. You don't have to know how to do any of that stuff. You give it the skills and then tell cloud code, hey, pull down some traces, see what's wrong, and make make my application better and it works.

01:46:56 Uh there should be two guardrails on that approach first. However, first you should feed it your requirements too and not just the failure explanations. So the goal isn't make the eval pass because otherwise it will cheat. It will include all of your test data into its eval uh and automatically pass all of your evals. So you have to make uh um your you have to make sure that your requirements to your cloud code are clear that you don't want it just to pass the evals.

01:47:24 You want it to do this specific thing which then makes the evals pass. Um and second, you should tell it to find themes and not chase individual failures. It's very tempting for claude code to look at one particular failed trace and write a whole bunch of code to fix that particular failed trace when there are actually 10 failing traces that are all the same co cause that it should be focused on instead.

01:47:44 Uh I'm not going to do all of that time all of that because I only have 12 minutes left. Uh but um once you have eval this is one of the highest leverage things that you can do with them. So let's step back and look at the whole shape of it. Uh production produces traces. Online eval them and produce explanations. A coding agent reads the explanations and proposes fixes.

01:48:09 And experiments verify that the fixes work and then you ship. The production traffic that comes back gets graded by the same online evals that started the loop. Uh and this is the software development life cycle closing in on itself. uh your AI software is helping you improve your AI software and AX is the substrate uh that ties it all together. Your traces, the evals, the explanations, the data sets, the experiments, the online evals and the monitors.

01:48:39 So just to recap, we instrumented an agent with two lines of code. We traced it and we read our data. We wrote a code eval and built-in LLM evals. We wrote a custom rubric from scratch. We validated our judge against human labels. Uh we saved failures at the data set. We had Claude improve the prompts. We ran an experiment to prove that it worked. Uh and we saw how AX takes all of that into production uh with online evals.

01:49:05 Uh this is the full loop and now you know all of it. However, you don't have to do all of it at once. This has been two hours of extremely dense information and I'm aware of that. You can start very very small. Uh you can start by reading your traces. Just add those two lines of code and start looking at your traces and that is already that is immediately going to give you information about what your agent is up to that you didn't know.

01:49:27 Uh 15 minutes of reading real outputs will teach you more about your application than an hour of building dashboards. Then write one code eval. Code evals are easier than LME evals and they're faster and they're cheaper. Uh and make it for the thing that matters most. So correctness, faithfulness, whatever fits your use case. Run them, look at the results, see what patterns emerge and build from there.

01:49:51 Evals are infrastructure. They are not an afterthought. Some teams create evals at the very start of development. Add others add them later once they are at scale. Uh but either way uh the important thing is that you treat evals as a core part of your system. They should be as routine as unit tests and the value compounds but only if you keep investing.

01:50:10 Each time a regression shows up before it show reaches your users instead of after you'll understand why all of this work is worth doing. Don't hope for great. You can get to great systematically by specifying it, measuring it, and improving towards it. So, now is the time to try this for real on an app of your own. Uh, you've already got an AX account, I hope, by the time the Wi-Fi kicked in.

01:50:33 Uh, the docs at arise.com uh have companion notebooks with runnable code for everything that we covered today. Uh, and if you want your coding agent to do the heavy lifting, you can install the Arise plugins, uh, the Arise skills plugin with one npx command. Uh, and that is in that link is in today's notebook. Uh, if you've made it all the way to the end of this, you can also get a free year of Arise Pro uh, using that code that's at the on the screen right now. And that is it. Thank you all for your time and attention.