← All transcripts

What Is MLflow? Tracing AI Agents & LLM Workflows Transcript, AI Summary & Key Points

IBM Technology · 14 days ago · Education · 09:38 · EN

Watch on YouTube

Answer

MLflow is an OpenTelemetry-compatible platform for tracing, evaluating, and monitoring multi-agent systems and LLM workflows.

AI Summary

MLflow provides OpenTelemetry-compatible observability for multi-agent applications and LLM workflows. It records traces containing inputs, outputs, metadata, token counts, latency, tool parameters, prompt-completion pairs, and database results across LLM calls, tool invocations, and database queries. MLflow supports deterministic evaluation scores, LLM-as-a-judge evaluations, prompt versioning, asynchronous trace logging, sampling, CI quality gates, and exports to other OpenTelemetry backends. Production deployments should use a real database, object storage, authentication through an auth proxy, asynchronous logging, deliberate sampling, an appropriate judge model, and automated evaluations.

Key Points

  • Traditional request monitoring can show a 200 response and a 2.3-second duration even when a user receives a wrong answer.
  • Multi-agent observability must expose wrong tool calls, empty MCP server results, unnecessary token usage, cascading latency, context overflow, and nondeterministic outputs.
  • A trace is a complete record of one request, composed of spans for LLM calls, tool invocations, and database queries in a parent-child tree.
  • Traces include inputs, outputs, execution time, token counts, per-step latency, tool parameters, prompt-completion pairs, database results, user IDs, and session IDs.
  • Deterministic evaluations can use regex checks, exact matches, and latency thresholds.
  • LLM judges can evaluate tool-call correctness, call efficiency, relevance, safety, and natural-language compliance guidelines.
  • Mortgage-assistant guidelines can require helpful, professional responses that never promise a rate.
  • The prompt registry provides version control for system prompts.

AI in practice

Used for

Agents

  • Determine which mortgage loans a user qualifies for by coordinating agents with different roles and personas. 1 held 02:05

Tools & resources

9 items

FNo. 4047
AIAINotes.us Tool

FastAPI

Open source · fastapi/fastapi

FastAPI is a Python web framework for building APIs using standard Python type hints. It is built on Starlette for web functionality and Pydantic for data handling, providing validation and conversion for request data from JSON, path and query parameters, cookies, headers, forms, and files. Type declarations also generate OpenAPI and JSON Schema documentation, including interactive Swagger UI and ReDoc interfaces. The video describes FastAPI routing requests from a mortgage application to five specialized agents.

Mentioned in
1 video
Kind
Other
JNo. 4049
AIAINotes.us Tool

Jaeger

jaegertracing.io

Jaeger is an open-source distributed tracing platform for monitoring and troubleshooting workflows in complex distributed systems. It maps requests and data as they move across services, helping identify performance bottlenecks, trace errors to their root causes, and analyze service dependencies. Jaeger can receive OpenTelemetry trace data, including traces exported from MLflow's OpenTelemetry-compatible tracing setup.

Mentioned in
1 video
Kind
Other
KNo. 0464
AIAINotes.us Tool

Kubernetes

Open source · kubernetes/kubernetes

Kubernetes (K8s) is an open-source container orchestration system hosted by the Cloud Native Computing Foundation. It manages containerized applications across multiple hosts, providing mechanisms for deploying, maintaining, scaling, and scheduling workloads. Applications can be configured with YAML and deployed across hybrid-cloud environments; the platform is also used for production operations involving storage, databases, networking, and security hardening.

Mentioned in
9 videos
Kind
Other
LNo. 0163
AIAINotes.us AI product

LangChain

Open source · langchain-ai/langchain

LangChain is an open-source framework and agent engineering platform for building agents and applications powered by large language models. It chains interoperable components and third-party integrations, providing standard interfaces for models, embeddings, vector stores, retrievers, tools, and other data sources. It supports sequential agentic workflows, context and tool-call orchestration, model substitution, and application development primarily through Python; the project also provides a separate JavaScript/TypeScript library. LangChain can be used standalone or with related tools for agent orchestration, evaluation, observability, debugging, and deployment.

Mentioned in
5 videos
Kind
AI
MNo. 0773
AIAINotes.us AI product

MLflow

Open source · mlflow/mlflow

MLflow is an open-source AI engineering platform for agents, large language models, and machine-learning models, developed in the MLflow repository. It provides experiment tracking for models, parameters, metrics, and evaluation results, along with production observability, evaluation, prompt versioning and optimization, and model lifecycle tools. Its agent and LLM observability captures application traces using OpenTelemetry and supports different LLM providers and agent frameworks; its AI Gateway provides an OpenAI-compatible interface for routing requests, managing rate limits and credentials, handling fallbacks, controlling costs, and applying guardrails. MLflow can be run as a server and accessed through Python, TypeScript/JavaScript, Java, and other programming languages, with integrations including OpenTelemetry and MCP.

Mentioned in
2 videos
Kind
AI
MNo. 2247
AIAINotes.us Tool

MySQL

Open source · mysql/mysql-server

MySQL is an SQL database server developed by the MySQL team at Oracle and used as a supported database backend for the Magic application builder. The MySQL distribution also includes MySQL Cluster, an open-source transactional database for real-time workloads. It is distributed under the licensing terms described in its repository.

Mentioned in
2 videos
Kind
Other
ONo. 0760
AIAINotes.us Tool

OpenTelemetry

Open source · open-telemetry

OpenTelemetry is an open-source, vendor-neutral observability framework for cloud-native software. It provides APIs, SDKs, automatic-instrumentation agents, and collector services for capturing distributed traces, metrics, logs, and contextual metadata such as baggage. Its context-propagation mechanism correlates telemetry across service boundaries and carries trace identity through an application’s request path, including standard HTTP communication. The OpenTelemetry Collector receives, processes, filters, and routes telemetry to observability backends such as Jaeger, Prometheus, commercial systems, or custom solutions. It supports native SDKs for more than 12 languages and is hosted as a graduated project by the Cloud Native Computing Foundation.

Mentioned in
5 videos
Kind
Other
PNo. 3295
AIAINotes.us Tool

Postgres

postgresql.org

Postgres is an open-source relational database created by Mike Stonebraker and subsequently maintained and enhanced by a volunteer community and companies. It is presented as a widely used database whose open ownership contributed to its adoption and continued development.

Mentioned in
3 videos
Kind
Other
RNo. 4048
AIAINotes.us Tool

Red Hat OpenShift

openshift.com

Red Hat OpenShift is an application innovation platform for building, modernizing, and deploying applications at scale. It provides tools for organizations to manage cloud-native applications and can serve as a production hosting environment for an MLflow deployment behind an authentication proxy.

Mentioned in
1 video
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of What Is MLflow? Tracing AI Agents & LLM Workflows — IBM Technology (09:38). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 You're running a multi-agent application with agents taking on different tasks and passing information to one another. And something breaks. Maybe your user gets a wrong answer or no answer at all. And so you do what you've always done, and you open up your monitoring dashboard. Everything looks fine. HTTP request, error rates, response times. You probably see something like this.

00:33 We're going to name this agent our CEO agent. And it says 200 OK, and then it took 2.3 seconds. But your user got the wrong answer. And so just knowing that a request happened here is not useful enough. With a multi-agent system, you need to be able to see if your agent maybe called the wrong tool or if an MCP server returned empty data, along with many other concerns new to this AI era.

01:07 Like, did three out of my five LLM calls, burned tokens on a prompt that nobody's looked at in months. So today we're going to go beyond request level monitoring and into proper LLM observability. For multi-agent systems. And to do that, we're going to use MLflow. So MLflow is an open telemetry compatible platform that captures inputs, outputs, and metadata from every step of a request that occurs after our user here.

01:49 Sends in the request. By the end of this video, you'll know how tracing works, how to evaluate agent quality using LLM judges. And in the end, I'll share a few configuration choices that separate a demo from a production level deployment. Let's ground this in an example. You've built a mortgage lending application. Users can chat about their needs and the system tells them what loans they qualify for.

02:18 And under the hood, it's a fast API. Routing to five different agents that all have different personas. A prospector, a borrower, a loan officer, an underwriter, and a CEO assistant for internal analytics. These agents can call on different tools like databases. LLMs, MCP servers, and all of this happens before our user gets a response back. Let's talk about what hides in the gaps here that we're trying to catch.

03:02 First up, silent tool failures, where an MCP tool returns empty data and the agent proceeds anyways and confidently answers from incomplete information. Yikes. Next. Cascading latency, where a slow database query delays an LLM call, which delays the response, and you can't tell which link in the chain is the bottleneck. Next, context overflow. The agent stuffs too much into a prompt.

03:29 The call fails with a cryptic error, and the user sees a generic timeout. And then finally, the big one, nondeterminism. The same question produces different results run to run, so the bug won't reproduce itself on your laptop. Nondeterminism means no quality baseline, proving that your agents give consistent and accurate answers. And in a regulated domain like lending, that's a compliance problem.

03:56 MLflow's core observability primitive is the trace. The trace is a complete record of one request. Now the trace is composed of individual spans. Now, these spans... Are a single LLM call, a tool invocation, a database query organized into a parent-child tree that shows what called what with what inputs, with what outputs, and how long each step took.

04:23 Plus, you'll get token counts, per-step latency, tool parameters, prompt completion pairs, and database results. You can tag these traces with the individual user and session IDs. So you can pull up individual users' sessions and replay an agent's decision path. But before you even get to that point, during the inner loop of AI development, you can work in a notebook that allows you to tweak prompts, run evals, etc, with fast feedback before you actually get to users hitting your system.

04:56 So far, we've talked about MLflow's deterministic scores. Some are regex checks, exact match, latency thresholds, but another modern evaluation framework for the non-deterministic era is LLM as a judge scores. Now, you can use a second model with MLflow to grade your agent's output against defined criteria. This criteria could be tool call correctness.

05:26 Did the agent pick the right tools for the question? To call efficiency. Did the agent get there in a minimal number of steps or in calls, or did it wander? Relevance, did the response actually address what the user asked? Safety, is the response appropriate? And then finally, guidelines. You can write your compliance requirements as natural language instructions via the guidelines.

05:57 For our mortgage assistant, that might say, be helpful, never promise a rate. Keep it professional. Another cool feature is the prompt registry. So if you find a system prompt that's worth an A+, you can register it with version control for these system prompts that you can think of like Git. Now that we have a strong conceptual understanding, let's talk about four rapid fire tips that separate a notebook demo from a production deployment.

06:24 First. Give the tracking server a real database. The default file-based backend is fine on your laptop and falls over with concurrent writers. Point --backend-store-URI at Postgres or MySQL before your team shares an instance and put artifacts on object storage. If you're on Kubernetes or OpenShift, run it behind an auth proxy. MLflow ships with no authentication story of its own.

06:57 Turn on async trace logging. In production, you don't want trace exports sitting in your request path. MLflow supports asynchronous logging, so spans ship in the background instead of adding latency to every user response. Under heavy traffic, add sampling. You rarely need 100% of traces once the system is stable, but you always want 100% errors. Third.

07:21 Thank you. Pick your judge model deliberately. LLM judges default to a hosted OpenAI model, but in enterprise or air-gapped environments, that can often be a non-starter. So point the judge at your own endpoint instead. And remember, every judge call is an LLN call. So if you score a 500 example dataset with five judges, you've made 2,500 inference requests.

07:45 Budget for it, and lean on deterministic scores wherever a regex would do. Fourth Evaluate and CI, not just in notebooks. The inner loop is where you iterate, but the payoff comes when the same scores run automatically on every prompt or agent change. A quality gate, exactly like unit test. The prompt registry plus versioned eval runs gives the audit trail.

08:11 Wiring mlflow.genai.evaluate parentheses into your pipeline gives you the enforcement. MLflow tracing is built on open telemetry, the industry standard. It plays very nice with others. Your traces aren't locked in. You can dual export to Jaeger, Grafana Tempo, or whatever backend your platform team already runs. MLflow integrates with dozens of LLM providers and agent frameworks as well.

08:36 And for supported frameworks, it's one line. Mlflow.langchain.autolog Call that once it's start up, and every lane chain and lane graph operation, LLM calls, tool invocations, agent decisions, gets captured automatically. No changes to your agent's logic. So if you've got custom Python glue code between frameworks, decorate it with at mlflow.trace. And it lands in the same tree.

09:16 Hope this was helpful. If so, like and subscribe. And I wanna know, what's the weirdest silent failure an agent has ever gotten past your monitoring? Drop it in the comments. Thanks.